SAM Architecture

Deep dive into how SAM works under the hood.

Thinking about contributing to SAM? Integrating it into your product? Or just curious how a conversational AI with memory, RAG, and tool execution works? This guide explains SAM's architecture from the ground up.

What you'll learn: - High-level system architecture and design philosophy - How the memory system stores and retrieves context - Tool system implementation (MCP framework) - API server architecture and endpoints - Data flow through the entire stack - Technology choices and why they were made

Who this is for: - Developers contributing to SAM - Engineers integrating SAM into products - Architects evaluating SAM for their stack - Anyone who wants to understand the internals

Prerequisites: - Familiarity with Swift and SwiftUI - Understanding of REST APIs and SSE - Basic knowledge of vector databases and embeddings (for memory system)

Let's explore how SAM's components work together to create an intelligent, capable AI assistant.


Table of Contents

  1. System Overview
  2. Core Components
  3. Memory & Intelligence Layer
  4. Tool System Architecture
  5. API Server Architecture
  6. Data Flow
  7. Technology Stack

System Overview

SAM is built as a native macOS application using SwiftUI with a modular, layered architecture.

Key Principles: - Modularity: Clear separation of concerns between components - Extensibility: Easy to add new providers, tools, and features - Privacy: Local-first design with optional cloud integration - Performance: Native Swift with Metal acceleration for local models

High-Level Architecture:

graph TB
    subgraph "User Interface Layer"
        UI[SwiftUI User Interface
ChatWidget, Preferences, Help] end subgraph "Conversation Management" CM[ConversationManager
Lifecycle, Persistence, State] Topics[Shared Topics
Cross-Conversation Memory] end subgraph "Voice Framework" VoiceMgr[VoiceManager
TTS, Wake Word, Speech Input] AudioMgr[AudioDeviceManager
Device Selection] end subgraph "API Framework" EO[EndpointManager
RESTful API Server] AO[AgentOrchestrator
Request Processing] end subgraph "Provider Layer" OpenAI[OpenAI Provider
GPT-4, GPT-3.5] Copilot[GitHub Copilot
GPT-4, Claude, o1] Gemini[Google Gemini
Gemini 2.5 Pro, Flash] Local[Local Models
MLX, GGUF] ALICE[ALICE Provider
Remote Stable Diffusion] end subgraph "MCP Tool Execution" Think[think] FileOps[file_operations] Terminal[terminal_operations] Memory[memory_operations] Web[web_operations] Docs[document_operations] Build[version_control] Subagent[agent_operations] end subgraph "Memory & Intelligence" VectorRAG[Vector RAG
Semantic Search] YaRN[YaRN Context Processor
Dynamic Scaling] MemDB[(SQLite Memory
Conversation/Topic Scoped)] end UI -->|User Input| CM UI -->|Voice Input| VoiceMgr VoiceMgr -->|Transcribed Text| CM AO -->|Stream Response| VoiceMgr VoiceMgr -->|Use Settings| AudioMgr CM -->|Manage State| Topics CM -->|Process| AO AO -->|Select Provider| OpenAI AO -->|Select Provider| Copilot AO -->|Select Provider| Gemini AO -->|Select Provider| Local OpenAI -->|Request Tools| Think OpenAI -->|Request Tools| FileOps OpenAI -->|Request Tools| Terminal OpenAI -->|Request Tools| Memory OpenAI -->|Request Tools| Web OpenAI -->|Request Tools| Docs OpenAI -->|Request Tools| Build OpenAI -->|Request Tools| Subagent Memory -->|Store/Retrieve| VectorRAG VectorRAG -->|Persist| MemDB AO -->|Context Processing| YaRN YaRN -->|Enhanced Context| VectorRAG Topics -->|Shared Memory| MemDB Subagent -->|Create New| AO AO -->|Stream Response| CM CM -->|Display| UI EO -->|HTTP API| AO

Core Components

1. ConversationManager

Location: Sources/ConversationEngine/ConversationManager.swift

Responsibilities: - Manages conversation lifecycle (create, load, save, delete) - Handles message persistence and retrieval - Integrates with the memory system - Coordinates with AI providers

Key Classes: - ConversationManager: Main orchestrator - ConversationModel: Conversation data model - EnhancedMessage: Message with metadata - MessageBus: Single source of truth for messages

State Management:

@Published public var conversations: [ConversationModel]
@Published public var activeConversation: ConversationModel?
public let memoryManager = MemoryManager()
public let vectorRAGService: VectorRAGService
public let yarnContextProcessor: YaRNContextProcessor

Per-Conversation Storage Architecture:

~/Library/Application Support/SAM/conversations/
├── {UUID}/
│   ├── conversation.json    # Single conversation data
│   ├── tasks.json           # Agent todo list
│   └── .vectorrag/          # Conversation-scoped RAG
├── active-conversation.json
└── backups/

Each conversation is stored in its own directory, providing: - O(1) save time per conversation (vs O(n) with monolithic file) - Backward-compatible migration from legacy conversations.json - Automatic cleanup when conversations are deleted

2. MessageBus Architecture

Location: Sources/ConversationEngine/MessageBus.swift

The MessageBus implements the Single Source of Truth pattern for all message operations.

Architecture:

classDiagram
    class MessageBus {
        +messages: [EnhancedMessage]
        -messageCache: [UUID: Int]
        +addUserMessage() UUID
        +addAssistantMessage() UUID
        +updateStreamingMessage(id, content)
        +completeStreamingMessage(id)
        -scheduleSave()
        -notifyConversationOfChanges()
    }

    class ConversationModel {
        +messages: [EnhancedMessage]
        +messageBus: MessageBus?
        +syncMessagesFromMessageBus()
    }

    class ChatWidget {
        +observes messages
        +displays real-time updates
    }

    MessageBus --> ConversationModel : syncs to
    ConversationModel --> ChatWidget : observes

Key Principles: - All message creation/updates go through MessageBus - ConversationModel is read-only mirror (updated by MessageBus) - ChatWidget observes ConversationModel, never modifies directly - O(1) message lookup via messageCache

Performance: - 30 FPS streaming throttle for UI updates - Delta sync to ConversationModel (no array copy) - Debounced persistence (500ms)

3. AgentOrchestrator

Location: Sources/APIFramework/AgentOrchestrator.swift

Responsibilities: - Executes autonomous workflows - Implements tool calling loop (VS Code Copilot pattern) - Manages iteration budget and limits - Handles context pruning with YaRN integration

Key Features: - Sequential thinking architecture - Loop detection and prevention - Dynamic iteration adjustment - Subagent spawning

Workflow Loop:

1. Receive user message
2. Add to context
3. Call LLM with available tools
4. Parse tool calls
5. Execute tools
6. Inject results
7. Continue until completion or iteration limit

4. MemoryManager

Location: Sources/ConversationEngine/MemoryManager.swift

Responsibilities: - Stores and retrieves memories - Enforces conversation/topic scoping - Manages database lifecycle

Storage: - SQLite database per conversation - Lazy loading on-demand - Automatic cleanup

Operations:

func storeMemory(content: String, conversationId: UUID, ...)
func retrieveRelevantMemories(for query: String, ...)
func searchAllConversations(query: String, ...)

Memory & Intelligence Layer

Vector RAG Service

Location: Sources/ConversationEngine/VectorRAGService.swift

Architecture:

Document Input
    ↓
DocumentChunker (semantic chunking)
    ↓
EmbeddingGenerator (512-d vectors via Apple NaturalLanguage)
    ↓
MemoryManager (SQLite storage with vectors)
    ↓
Semantic Search (cosine similarity)
    ↓
Ranked Results

Components: - DocumentChunker: Intelligent content segmentation - EmbeddingGenerator: Apple NaturalLanguage integration - ProcessedChunk: Container for chunks + embeddings - SemanticSearchResult: Search results with scoring

Key Algorithms:

func chunkDocument(_ document: RAGDocument) -> [DocumentChunk]
func generateEmbedding(for text: String) -> [Double]
func semanticSearch(query: String, threshold: Double) -> [Result]

YaRN Context Processor

Location: Sources/ConversationEngine/YaRNContextProcessor.swift

Purpose: Dynamic context window management with intelligent compression

Profiles:

public static let `default` = YaRNConfig(
    baseContextLength: 8192,
    extendedContextLength: 32768,
    scalingFactor: 4.0,
    attentionFactor: 0.1,
    compressionThreshold: 0.8
)

public static let mega = YaRNConfig(
    baseContextLength: 65536,
    extendedContextLength: 134217728, // 128M
    scalingFactor: 2048.0,
    attentionFactor: 0.001,
    compressionThreshold: 0.95
)

Processing Pipeline:

1. Analyze message importance
2. Identify preservation candidates
3. Apply semantic clustering
4. Calculate compression targets
5. Generate compressed context
6. Archive rolled-off messages to ContextArchiveManager
7. Cache result

Context Archive System

Location: Sources/ConversationEngine/ContextArchiveManager.swift

When YaRN compresses conversations, older messages are archived rather than discarded.

Architecture:

Active Context (Fits model window)
    ↓ YaRN Compression
Rolled-Off Messages
    ↓
ContextArchiveManager (SQLite)
    ↓
memory_operations (recall_history operation)
    ↓
Agent retrieves archived context on demand

Key Features: - SQLite-backed storage for archived context - Topic-wide search capability for shared topics - Preserves rolled-off messages with summaries and key topics - On-demand retrieval via recall_history operation (memory_operations)

Shared Topics

Location: Sources/SharedData/SharedTopicManager.swift

Architecture:

Topic Storage (SQLite shared-data.db)
    ↓
Topic Metadata (id, name, description)
    ↓
Workspace Directory (~/SAM/{topic-name}/)
    ↓
Effective Scope Resolution
    ↓
Memory/File/Terminal Operations

Effective Scope Pattern:

let effectiveScopeId = if sharedTopicEnabled {
    topic.id
} else {
    conversation.id
}

Tool System Architecture

MCP Framework

Location: Sources/MCPFramework/

Tool Interface:

public protocol MCPTool {
    var name: String { get }
    var description: String { get }
    var parameters: [String: MCPToolParameter] { get }

    func execute(parameters: [String: Any], context: MCPExecutionContext) async -> MCPToolResult
}

Consolidated Tools:

public protocol ConsolidatedMCP: MCPTool {
    var supportedOperations: [String] { get }
    func route(operation: String, params: [String: Any], context: MCPExecutionContext) async -> MCPToolResult
}

Core MCP Tools: 1. think - Planning and analysis 2. increase_max_iterations - Dynamic iteration management 3. read_tool_result - Large result retrieval 4. user_collaboration - User input requests 5. file_operations - File operations (read, search, write) 6. terminal_operations - Execute shell commands 7. memory_operations - Memory operations (store, search, recall, LTM) 8. web_operations - Web research, search, fetch 9. document_operations - Import, create, and manage documents 10. calendar_operations - macOS Calendar integration 11. contacts_operations - macOS Contacts integration 12. notes_operations - Apple Notes integration 13. spotlight_search - macOS Spotlight search 14. weather_operations - Current weather and forecast 15. image_generation - Generate images via remote ALICE 16. math_operations - Calculate, convert, and run formula math 17. code_intelligence - Find symbol usages and search commit history 18. version_control - Git version control operations 19. remote_execution - Run tasks on remote systems over SSH 20. apply_patch - Apply multi-file patches 21. agent_operations - Spawn and coordinate sub-agents

Tool Execution Flow

sequenceDiagram
    participant User
    participant ChatWidget
    participant ConversationManager
    participant AgentOrchestrator
    participant LLM as AI Provider
    participant MCPManager
    participant Tool as MCP Tool
    participant System

    User->>ChatWidget: Send message
    ChatWidget->>ConversationManager: Process message
    ConversationManager->>AgentOrchestrator: Orchestrate request
    AgentOrchestrator->>LLM: Send with tool schemas

    loop Tool Execution Loop
        LLM->>MCPManager: Request tool call
        MCPManager->>Tool: Route to specific tool
        Tool->>System: Execute operation
        System->>Tool: Return result
        Tool->>MCPManager: MCPToolResult
        MCPManager->>LLM: Inject result into context

        alt More iterations needed
            LLM->>MCPManager: Request another tool
        else Task complete
            LLM->>AgentOrchestrator: Final response
        end
    end

    AgentOrchestrator->>ConversationManager: Stream response
    ConversationManager->>ChatWidget: Update UI
    ChatWidget->>User: Display response

API Server Architecture

EndpointManager

Location: Sources/APIFramework/EndpointManager.swift

OpenAI-Compatible Endpoints:

POST /api/chat/completions
POST /api/chat/autonomous
GET  /api/models
GET  /v1/conversations
POST /v1/conversations
GET  /v1/conversations/{id}

SSE Streaming:

func streamResponse() -> AsyncThrowingStream {
    AsyncThrowingStream { continuation in
        // Stream chunks as they arrive
        continuation.yield(chunk)
        continuation.finish()
    }
}

Architecture:

graph TB
    Client[API Client
HTTP POST] Router[EndpointManager
Route Handler] Validate[Request Validation
Check conversationId] Check{Has conversationId?} LoadConv[Load Existing
Conversation + Memory] NewConv[Create New
Conversation] Orchestrator[AgentOrchestrator
Process Request] Tools[Tool Execution Loop
MCP Tools] StreamCheck{Streaming?} SSE[Server-Sent Events
data: {...}] JSON[JSON Response
{...}] Client -->|POST /api/chat/completions| Router Router --> Validate Validate --> Check Check -->|Yes| LoadConv Check -->|No| NewConv LoadConv -->|With context| Orchestrator NewConv -->|No context| Orchestrator Orchestrator -->|Execute| Tools Tools -->|Results| Orchestrator Orchestrator --> StreamCheck StreamCheck -->|stream=true| SSE StreamCheck -->|stream=false| JSON SSE -->|Stream chunks| Client JSON -->|Complete response| Client

Data Flow

Message Flow

User Input
    ↓
ChatWidget (UI)
    ↓
ConversationManager.addMessage()
    ↓
Persistence (debounced)
    ↓
EndpointManager.processMessage()
    ↓
AgentOrchestrator.executeAutonomousWorkflow()
    ↓
LLM API Call (streaming)
    ↓
Tool Execution (if needed)
    ↓
Response Assembly
    ↓
UI Update

Memory Storage Flow

Content to Store
    ↓
MemoryManager.storeMemory()
    ↓
EmbeddingGenerator.generate()
    ↓
SQLite INSERT with vector
    ↓
Tags + Metadata
    ↓
Success

Memory Retrieval Flow

Search Query
    ↓
EmbeddingGenerator.generate()
    ↓
SQLite Vector Search (cosine similarity)
    ↓
Filter by threshold
    ↓
Rank by similarity
    ↓
Return top N results

Document Import Flow

File Drop/Selection
    ↓
DocumentImportSystem
    ↓
Format Detection (PDF, DOCX, etc.)
    ↓
Text Extraction
    ↓
Page-Aware Chunking
    ↓
VectorRAGService.ingestDocument()
    ↓
EmbeddingGenerator (per chunk)
    ↓
MemoryManager.storeMemory()
    ↓
Indexed and searchable

Technology Stack

Frontend

  • SwiftUI: Modern declarative UI framework
  • Combine: Reactive programming
  • AppKit Integration: Terminal, file pickers

Backend

  • Swift 5.9+: Type-safe, performant
  • Vapor (API Server): HTTP routing, SSE streaming
  • Logging: Swift Logging framework

Storage

  • SQLite: Conversation and memory storage
  • FileManager: Workspace file operations

AI Integration

  • OpenAI SDK: GPT-4 and compatible models
  • GitHub Copilot: OAuth + API integration
  • MLX: Apple Silicon local inference
  • llama.cpp: Cross-platform GGUF models
  • NaturalLanguage: Vector embeddings

Voice & Sound

  • NSSpeechSynthesizer: Native macOS text-to-speech
  • CoreAudio: Audio device enumeration and selection
  • AVAudioEngine: Voice input processing
  • Streaming TTS: Sentence-by-sentence speech during LLM streaming

Terminal

  • PTY (pseudo-terminal): Full terminal emulation
  • Process: Subprocess management

Design Patterns

Dependency Injection

public init(
    endpointManager: EndpointManager,
    conversationService: SharedConversationService,
    conversationManager: ConversationManager,
    maxIterations: Int = 300
) {
    self.endpointManager = endpointManager
    self.conversationService = conversationService
    self.conversationManager = conversationManager
    self.maxIterations = maxIterations
}

Protocol-Oriented Design

public protocol AIProviderProtocol: AnyObject {
    func processStreamingChatCompletion(...) async throws -> AsyncThrowingStream<...>
}

Observable Objects

@MainActor
public class ConversationManager: ObservableObject {
    @Published public var conversations: [ConversationModel]
    @Published public var activeConversation: ConversationModel?
}

Async/Await

public func executeAutonomousWorkflow(...) async throws -> String {
    let response = try await provider.processChatCompletion(...)
    return response
}

Security Architecture

Authorization: - Working directory: Auto-approved - Outside working directory: Requires user confirmation - Path normalization prevents ../ attacks

API Keys: - Never logged or transmitted (except to provider) - User-managed configuration

HTTPS Enforcement: - All web operations require HTTPS - HTTP URLs auto-converted to HTTPS

Sandboxing: - File operations respect authorization - Terminal operations require approval - Tool permissions configurable


Performance Considerations

Lazy Loading: - Memory databases loaded on-demand - Vector embeddings generated as needed - Conversation history paginated

Context Caching: - YaRN caches processed context per conversation - Token counting cached - Embedding vectors cached

Debounced Persistence: - Conversations saved in batches (500ms debounce) - Prevents excessive disk I/O

Streaming: - SSE streaming for real-time responses - Reduces perceived latency - Better UX for long responses


Extension Points

Adding New Providers: 1. Implement AIProviderProtocol 2. Add to EndpointManager.providers 3. Add UI in Preferences

Adding New Tools: 1. Create class implementing MCPTool or ConsolidatedMCP 2. Register in MCPManager.initializeBuiltinTools() 3. Add to tool catalog

Custom System Prompts: 1. Create in Preferences → System Prompts 2. Variables: {{working_directory}}, {{date}}, etc. 3. Available in all conversations


Further Reading


Understand SAM's architecture to extend and customize effectively!