Personal AI Assistant
An assistant that does things, not just answers.
A custom AI assistant with a reactive, systems-style interface. It routes intent, calls tools, remembers context and automates the small workflows that eat a day.
- Design and engineering — architecture, orchestration, interface
- Desktop & web
- Python, LLM tool calling, RAG, Speech I/O, WebSocket, Web UI
- Private project — demo on request
Overview
A personal assistant inspired by the intelligent-assistant systems of film, built to be genuinely useful rather than theatrical. The interface communicates what the system is doing at every moment; the backend turns language into actions through a controlled set of tools.
Problem
General chat assistants answer well but act poorly: they don't know your context, can't touch your tools, and give no sense of what they're doing while they work.
Goals
- Turn natural language into real actions safely
- Remember relevant context across sessions
- Make latency feel like progress, not waiting
- Keep the architecture modular so new tools are cheap to add
My role
Design and engineering — architecture, orchestration, interface.
- Assistant architecture and orchestration loop
- Tool design, permissioning and error handling
- Memory and retrieval pipeline
- Interface design and realtime state visualisation
Technology
- Python / Async orchestration
- LLM APIs / Tool calling / RAG / Embeddings
- Web UI / WebSocket streaming / Speech-to-text / Text-to-speech
Architecture
A small orchestrator sits between the interface and the model. The model proposes; the orchestrator decides what actually runs.
Interface
Voice or text in; streamed tokens and state changes out over WebSocket
Orchestrator
Conversation state, intent routing, retries and timeouts
LLM + tools
Model selects from an allowlisted tool registry with typed arguments
Memory
Rolling summary for short-term context, retrieval for long-term knowledge
Automations
Scheduled and on-demand workflows reuse the same tool layer
Challenges & solutions
Latency is the user experience
Multi-step tool use can take seconds. A silent screen makes a capable system feel broken.
Stream everything: tokens, tool calls and state transitions. The interface maps each state to a distinct visual so waiting reads as work happening.
Letting a model act safely
An LLM with unrestricted tools is one misread instruction away from doing damage.
Typed, allowlisted tools with validation, and an explicit confirmation step for anything destructive or outward-facing.
Context that scales
Stuffing full history into every request is slow, expensive and eventually impossible.
A rolling conversation summary plus retrieval over personal notes, so each request carries only what's relevant.
Screens
- local / hud◇ ATLASModelonlineMemory1,284 notesTools12 enabledLatency240 msAutomations3 scheduledStateTHINKING3/4steps“Plan my afternoon around the 3 pm review.”Ask or command…⌘ KActivitycalendar.read✓today · 5 events120msnotes.search✓“review prep” · 4 hits88msplanner.compose✓3 focus blocks1.4stasks.create◌2 reminders—Needs confirmationCreate 2 reminders for 14:30 and 16:45?
System view — state, tools and memory in one frame - local / chat◇ ATLASModelonlineMemory1,284 notesTools12 enabledLatency240 msAutomations3 scheduledPlan my afternoon around the 3 pm review. Keep an hour for the XyloFit release.↳calendar.read(today)✓↳notes.search("review prep")✓↳planner.compose(blocks=3)✓Here’s a plan that protects your focus time:13:00Release checklist — XyloFit 1.414:30Review prep · 4 notes attached15:00Design review16:45Inbox & follow-upsSet reminders for 14:30 and 16:45?ConfirmEditReply…↵
Conversation with visible tool calls
Results
- A working assistant used for real daily workflows
- A modular tool layer where adding a capability is a small, isolated change
- An interface that makes the system's state legible at a glance
Lessons learned
“The orchestrator matters more than the prompt.”
“Visible state builds more trust than faster answers.”
“Design tools for the model like you design APIs for people: small, typed, forgiving.”