Real-time streaming VTT + AI Agent — invisible on any call app. Local. Free.
Sits in your calls, transcribes with speaker labels, speaks only when directly addressed. Invisible. Local. Free. No bot joins. The host app never knows it's there. Works on Teams, WhatsApp, Signal, Messenger, Webex, WeChat, Zoom, Discord — anything that pipes audio through your sound card.
Every "AI meeting assistant" on the market does the same thing: a bot joins your call, captures every word, ships it to a third-party cloud, and bills you per-seat per-month. Your conversation leaves your machine. Your data trains someone else's model. The vendor can read every confidential word, every private moment, every trade secret.
And the bot — it shows up. In the meeting. As a participant. "Hi, I'm Otter, can I record?" It changes the room. People self-censor. Off-the-record moments get recorded anyway. The whole dynamic shifts when a corporate eavesdropper has a seat at the table.
None of that is necessary. The microphone already hears the meeting. The audio already plays through your speakers. The LLM can already run on your GPU. The whole stack can be local, free, and invisible — if you build it that way.
Agent In The Armchair doesn't ask Teams for permission. It listens to the audio your computer already plays — through a virtual audio cable (VB-Cable) — and mixes in your microphone. Teams never sees a bot. The meeting never knows. The transcript never leaves your machine.
A streaming pipeline picks up the mixed audio, runs Silero VAD to detect speech (silence is free), feeds it to faster-whisper for incremental transcription with word timestamps, then runs pyannote-audio every 10 seconds on a 16-second rolling buffer to attach speaker labels. The result: a live, labeled transcript on a local dashboard.
In Talk mode, a local LLM (Ollama) decides whether the agent is being directly addressed — not just mentioned, not just present. If yes, the TTS layer synthesises a response and plays it back into the meeting through CABLE-A. The TTS engine is pluggable — by default Piper (any of the free open-source voices, e.g. Alan / British RP, Aru, Northern English, US, etc.), or swap in Kokoro-82M for 54 voices across 8 languages with sub-real-time generation. Or fire up Chatterbox by Resemble AI for zero-shot voice cloning from a 10–25 second reference clip — your own voice, anyone’s voice, with paralinguistic tags like [laugh]. All three hot-swap mid-call from the dashboard — engine, voice, and reference WAV all switch on the next utterance; the worker pool keeps used engines loaded so switching back is instant. Hit the dashboard’s ⚡ Activate button to pre-warm an engine before a call (Chatterbox cold-start ~8s for the first utterance, ~2.4s warm after). The agent can also carry a persona brain — point the dashboard at any folder of .md files and the agent calls with that persona’s memory loaded into the system prompt, not just its voice. Three engines, one pluggable interface, near-unlimited voice options. If no, the agent stays silent. LLM gate, not keyword matching.
VB-Cable on Windows. ffmpeg amix meeting audio + local mic into one 16kHz mono PCM stream. No software hooks into Teams. No API access needed. The audio already plays through your speakers — Armchair just listens to the same stream.
Silero VAD → faster-whisper (CUDA) → pyannote speaker diarization (CUDA) → Ollama LLM gate → pluggable local TTS (Piper, Kokoro, or any compatible engine). Every component runs on local hardware. NVIDIA GPU strongly recommended. No third-party APIs.
Silero VAD detects speech and skips silence (nearly free on GPU). faster-whisper with incremental transcription and word timestamps. ~0.3s end-to-end latency in Listen mode.
pyannote-audio on a 16s rolling buffer (runs every 10s). Per-segment speaker labels matched to Whisper word timestamps. Handles mid-utterance speaker switches. Name speakers in the dashboard — labels apply retroactively.
🔇 Listen: pure VTT with speaker labels. 🎤 Talk: VTT + LLM + TTS — agent responds when directly addressed. The LLM decides "direct address vs mention" and returns [SILENCE] for mentions. No keyword matching.
Default: Piper — pick any of the free open-source voices (Alan / British RP, Aru, Northern English, US, etc.). Or swap in Kokoro-82M for 54 voices across 8 languages with sub-real-time generation. Or fire up Chatterbox for zero-shot voice cloning from a 10–25 second reference clip — your voice, anyone's voice, with paralinguistic tags. All three hot-swap mid-call from the dashboard (engine, voice, reference WAV); the worker pool keeps used engines loaded. Hit ⚡ Activate to pre-warm before a call (Chatterbox ~8s cold, ~2.4s warm). All local, all free, no API keys.
Local browser dashboard at http://localhost:8765. Live labeled transcript, speaker naming, agent name config, mode toggle (Listen / Talk), voice selection. No cloud, no auth, no telemetry.
Each session archived to its own folder: transcript.txt, speaker_names.json, detected_speakers.json, audio.raw. Clean start each time — no bleed between sessions. Ctrl+C stops and saves.
Double-click start_armchair.bat — audio capture, dashboard, browser, and pipeline all start. The launcher self-cleans stale processes, health-checks the dashboard port, and opens the Piper voices catalog on first run if no models are found. No manual orchestration. Windows-native.
No bot joins the call. No API access. No participant slot. No "Hi, I'm recording." Hardware audio capture only — Teams, WhatsApp, Signal, Messenger, Webex, WeChat, Zoom, Discord — none of them can block it because none of them know it exists.
No subscriptions. No API keys. No cloud bills. No per-seat pricing. One-time setup, infinite meetings. The GPU is already on your desk.
Built for real-time conversation, not batch transcripts. The numbers below are the actual pipeline timings on a single NVIDIA workstation.
Every product we ship is built on the same principle: own the stack. No SaaS vendor in the middle. No per-hour billing. No third-party model reading your private conversations. If we wouldn't run it on our own workstation, we don't ship it.
Agent In The Armchair is part of The Pack — the same architecture that runs Executive Mind's own operations. If it can't survive our own meetings, it doesn't ship. If it can't run on a single NVIDIA workstation, it doesn't ship. The bar is what we trust for ourselves.
The marketing copy says "Invisible. Local. Free." That's not a tagline. It's a constraint. Every component of the pipeline is local. Every model runs on your hardware. The MIT license means zero per-hour cost. The architecture means Teams can't see it. Those aren't features — they're the design brief.
Open source (MIT). No subscriptions. No per-hour billing. Four commands on Windows, one reboot (VB-Cable):
git clone → .\install.bat → .\setup_audio.ps1 → .\start_armchair.bat
Then the agent is in the room — listening.
View on GitHub →Built by Kris Racette • Released under MIT • Source available on GitHub
Kris Racette (also known as Kristopher Racette and Kristopher Marc Racette) is the Founder of Executive Mind and the creator of Agent In The Armchair. With 18+ years of accounting experience and a Master of Commerce from the University of New South Wales, Kris brings the same engineering rigour to local AI systems that he brings to financial architecture.
Agent In The Armchair is part of The Pack — the open-source ecosystem built and run by Executive Mind. Gym Tracker handles the data athlete. ColdStack handles the crypto vault. Armchair handles the meeting room. Same philosophy across all three: local, open, no lock-in, your hardware, your data.
The source code is available for audit and contribution on GitHub. Executive Mind believes the same thing about meeting audio as it does about training data and vault data: inspectable, portable, and yours.