Local Vocal Agent — a privacy-first local AI voice assistant
A full-stack voice assistant that runs locally: FastAPI and React, an Ollama-hosted LLM, speech-to-text and text-to-speech, and SQLite plus Chroma memory.

The problem
Most voice assistants send audio and conversation history to a cloud service. Local Vocal Agent keeps the language model, speech pipeline and memory on the user's own machine.
Architecture and approach
- FastAPI backend exposing versioned endpoints for text chat, voice chat, sessions, user profile and system status/metrics.
- React + Vite frontend for the chat and voice workspace.
- LLM inference through Ollama; speech-to-text with Faster Whisper and a text-to-speech stage for spoken replies.
- Two memory layers: SQLite for sessions and messages, and Chroma as a vector store for retrieval.
- Internet search support with a configurable provider (Google News RSS by default, DuckDuckGo as an alternative) and an automatic fallback when a provider returns nothing.
Key features
- Streaming chat and voice-chat endpoints.
- Session history and per-session message retrieval.
- System metrics and status endpoints for the running stack.
- A CI check that keeps the CI requirements file in sync with the full requirements.
Engineering decisions
- Local-first by design: models run under Ollama on the host instead of a hosted API.
- Search providers are pluggable through configuration, so the assistant degrades gracefully when one source fails.
Technology stack
- AI
- Ollama
- Faster Whisper (STT)
- TTS
- ChromaDB (vector memory)
- RAG
- Backend
- Python
- FastAPI
- SQLite
- Frontend
- React
- Vite