Local Vocal Agent — a privacy-first local AI voice assistant

A full-stack voice assistant that runs locally: FastAPI and React, an Ollama-hosted LLM, speech-to-text and text-to-speech, and SQLite plus Chroma memory.

Local Vocal Agent — project screenshot

The problem

Most voice assistants send audio and conversation history to a cloud service. Local Vocal Agent keeps the language model, speech pipeline and memory on the user's own machine.

Architecture and approach

  • FastAPI backend exposing versioned endpoints for text chat, voice chat, sessions, user profile and system status/metrics.
  • React + Vite frontend for the chat and voice workspace.
  • LLM inference through Ollama; speech-to-text with Faster Whisper and a text-to-speech stage for spoken replies.
  • Two memory layers: SQLite for sessions and messages, and Chroma as a vector store for retrieval.
  • Internet search support with a configurable provider (Google News RSS by default, DuckDuckGo as an alternative) and an automatic fallback when a provider returns nothing.

Key features

  • Streaming chat and voice-chat endpoints.
  • Session history and per-session message retrieval.
  • System metrics and status endpoints for the running stack.
  • A CI check that keeps the CI requirements file in sync with the full requirements.

Engineering decisions

  • Local-first by design: models run under Ollama on the host instead of a hosted API.
  • Search providers are pluggable through configuration, so the assistant degrades gracefully when one source fails.

Technology stack

AI
  • Ollama
  • Faster Whisper (STT)
  • TTS
  • ChromaDB (vector memory)
  • RAG
Backend
  • Python
  • FastAPI
  • SQLite
Frontend
  • React
  • Vite