backchannel

Your meetings, transcribed live.
Your next move, surfaced mid-call — on your own hardware.

A self-hosted, open-source AI meeting assistant: live speaker-attributed transcripts and real-time insight agents. No bot in the call, and with the PII Shield on, no name in any prompt.

Free desktop app for Windows, macOS & Linux · self-host anytime with Docker Compose · works with any meeting — Zoom, Meet, Teams, or in the room · MIT licensed

Backchannel live during a fictional recovery-readiness review: a quiet listening bar, the live strategic-signal strip, 125 insights, an answered mid-call question, a speaker-attributed transcript, and the ask bar along the bottom.
FIG. 1A high-stakes recovery review, handled remotely: signals across the top, a question asked and answered mid-call, the attributed transcript running beside it. 125 insights from one 46-minute call, distilled into a single briefing.
New in v0.6.0

Every model gets [PERSON_1]. Only your screen gets Owen.

Turn on the PII Shield and names, companies, email addresses, phone numbers, card and ID numbers, IP and street addresses become tokens the moment a transcript line, directive, document excerpt, session name or speaker name is written. Every agent, the briefing, chat and Ask work from those tokens, whether the model runs on this machine or in someone's cloud. The database holds only tokens. The real values sit in a vault encrypted under the same master key that protects your provider keys, and come back only on the screen in front of you.

The Privacy tab's try-a-sentence box. Typed in: Owen Delacroix from Alderwake Health Network owns the identity approval, reach him at owen.delacroix@alderwake.example or 212-555-0142 before the board review. Shown as a model would receive it: [PERSON_1] from [ORG_1] owns the identity approval, reach him at [EMAIL_1] or [PHONE_1] before the board review. A legend gives each token's value and whether the on-device name model or a pattern found it.
FIG. 2The Privacy tab's scratch box: one sentence in, the sentence a model receives out, and how each token was found. Two by the on-device name model, two by pattern. Nothing typed here is stored.
  1. Written Owen Delacroix A transcript line lands, a directive is typed, a document is read.
  2. Tokenized [PERSON_1] On this machine, before the row exists. The value goes to the vault.
  3. Stored [PERSON_1] The database holds tokens only. The vault is encrypted under your master key.
  4. Sent to a model [PERSON_1] Local or cloud, every agent, the briefing, chat and Ask receive tokens and nothing else.
  5. On your screen Owen Delacroix Decoded at the edge, for the person in front of it. Every reveal is counted.
FIG. 3One value's route through a session. The same token stands for the same value for the whole call, so a model can follow who said what; nothing links a person across sessions.

Audio never leaves while the shield is on

Audio cannot be tokenized, so the shield holds transcription to a local model and switches off cloud live captions. Cloud text models stay available, because they only ever receive tokens. Uploaded documents are read on this machine and never sent as files. Privacy First decides where processing happens; the shield decides what any model, local or cloud, gets to read. Local transcription is rougher than cloud, so the optional Transcript Refiner sends the tokenized transcript to any text model to fix punctuation, casing and mishearings, and keeps a rewrite only if it carries exactly the original tokens.

Admin Privacy tab with the PII Shield on, badged Personal data tokenized. Four coverage rows: transcripts, insights, briefings, chat and documents protected; transcription audio protected, held to a local model; live captions protected; transcript refinement not covered because the refiner is off. Below them: Vault, 9 protected values across all sessions.

Detected here, never in the cloud

Pattern checks for structured identifiers, the session's own speaker roster, a protected-terms list you keep, and a small on-device name model (about 110 MB, downloaded once). No detection ever involves a cloud call.

Checkable, not claimed

Record outbound prompts logs every prompt exactly as it left for a model, badged “tokens only” or “blocked”. While the shield is on, a prompt still carrying a vault value is refused before it is sent.

Exports and the audit trail

Exports carry tokens unless you tick “Include personal data”. Every reveal, on screen or in a file, is counted. Sessions recorded before the shield was on can be protected after the fact.

The full account of the PII Shield — what is tokenized and when, where detection happens, how a value is decoded, and the limits stated plainly.

Before the call

Start from the top of one screen

Start Call and Process Transcript sit at the top of the setup screen and stay there while you scroll. A readiness line under the session name says what is configured; the steps below are collapsed cards whose headers say what they hold. The sidebar's find box understands dates, so October, oct 8, 10/8 or 2026-10-08 all bring up the sessions from that day.

The pre-call setup screen for a fictional Alderwake pilot scope review: Start Call pinned beside the session name, the readiness line Client / prospect, 2 directives, 4 participants, the audio-capture switch, an open Context card, and collapsed cards for Documents, Import a transcript or recording, and Coaching directives.
FIG. 4Client / prospect, 2 directives, 4 participants: the readiness line says what the call has, and Start Call stays in reach however far you scroll.
During the call

Everything a second set of ears should do

Built for high-stakes sales, discovery, and service-delivery calls — especially when the account team is remote. Provider-flexible, centered on a transcript you can trust.

Live diarized transcription

Silero VAD and WeSpeaker ResNet152 speaker embeddings attribute every line to a speaker, with interim text streaming seconds ahead of the final transcript. Microphone and tab/system audio are captured as separate tracks, so remote participants get their own identities.

Fictional recovery-readiness transcript with every line attributed to Me, Leah, Owen, or Maya and shown with timestamps.

Agents that work the call, not just the recap

Five live agents -- an analyst, a fast objection handler, a synthesizer, an opportunity specialist, and a strategic-signals scanner -- each run on their own trigger and push questions, objections, opportunities, and action items mid-call. Signals raised early are kept rather than overwritten, so the point someone made twenty minutes ago is still there. At call end, two independent briefing lenses draft in parallel and an arbiter reconciles them into one summary.

Speaker-attributed action items for the board-ready evidence outline and Tuesday technical working session, each badged with the agent that produced it.

Provider-routed models, on your keys

Mix Google Gemini and OpenAI per agent, transcribe fully offline with local ONNX Whisper and Parakeet, or point the agents at any number of your own OpenAI-compatible servers -- Ollama, LM Studio, vLLM, LiteLLM -- where every model an endpoint serves appears by name in each agent's picker, and run the whole pipeline with no key from anyone. Privacy First judges the destination, so an endpoint on your own machine or LAN keeps the agents running with the switch on; only cloud providers stay blocked.

Admin Connections tab: the Privacy First switch and Google and OpenAI credential cards above the Self-Hosted Models card for LM Studio, Ollama, vLLM, and LiteLLM endpoints.

Mid-call directives

Drop a note while the call runs and the analysis agents factor it into what they surface next.

Import and re-transcribe

Bring in existing transcripts or audio files, and replay any recorded call through a different model later.

Exports and cross-session chat

Transcript, insight, and summary exports, and chat across every past meeting. With the shield on, exports carry tokens unless you tick “Include personal data”.

Ask the call a question while it is still running

Twenty minutes ago someone named a number and you did not write it down. Type the question into the bar under the transcript and the answer comes back from the conversation as it stands — without stopping the recording, opening a second tool, or waiting for the post-call briefing. The answer is saved with the call, starred, and exported alongside every other insight.

The call's command bar: a Chat and Directive mode toggle, the question “What did Owen commit to sending tomorrow?” typed in, and a Gemini 3.6 Flash model chip marked Recommended.
The whole interface: two modes, one input, and the model that will answer.
A live call with two answered mid-call questions pinned at the top of the insight feed and a third being typed into the ask bar at the bottom of the screen.
FIG. 5Asked mid-call, answered from the transcript so far — and kept, so the answer is still there afterwards.

Two modes, one bar

Chat asks the call a question. Directive steers what the agents look for next. An answer you want the crew to act on converts to a directive in one click.

You choose who answers

The model that answers is picked in the bar itself — Google, OpenAI, or a model on your own hardware. Nothing is selected on your behalf.

Grounded in the room

Answers draw on the live transcript, the insights raised so far, strategic signals, your directives, and the summaries of documents you attached.

Questions answer themselves

The synthesizer keeps reading the room. When a question gets answered mid-call, it marks the card Answered, summarizes the answer, and spins off the follow-up you still owe. No bookkeeping from you.

The live feed filtered to 34 questions: the leading card is marked Answered with the synthesizer's one-line answer summary, and carries the follow-up question it spun off.
FIG. 6Mid-call — the synthesizer hears the answer land, marks the question Answered, and hands you the follow-up you still owe.
After the call

Stop the call. Open on what mattered.

Stop the call and Backchannel drains the last segments, runs the final analysis and saves the session. A finished call opens on an Overview: the briefing's top outcome first, then commitments, open loops, opportunities, risks and estimated spend, each count linking to the tab that holds the rows behind it. Below that, who spoke how much, and when in the call each insight arrived, drawn as call time with a seam at every resume.

The Overview tab of the Alderwake recovery-readiness session: a 44m 12s call plus a 2m resume, a 9 shielded badge beside Resume Call, the headline outcome Sponsor aligned on a non-disruptive pilot, and a row of counts: Commitments 4, Open loops 3, Opportunities 3, Risks 3, and Est. spend $1.14 from 1,287,379 tokens, above the start of the commitments and open-loops lists.
FIG. 7Sponsor aligned on a non-disruptive pilot: the top outcome, then five counts. The session header carries the shield's tally for this call, 9 shielded, next to the resume button.
The Overview digest scrolled into view: four commitments with due dates and owners, three open loops phrased as questions, three opportunities including a Recovery Implementation Pilot valued at $72K to $84K, and three risks with owners, each list ending in a link to the full Insights tab.
FIG. 8The digest: commitments with owners and due dates, open loops as the questions still unanswered, opportunities with a value range, risks with the person who owns them. Each list links through to the rows behind it.

The briefing does the reading for you

A 46-minute call produced 125 grounded signals. You never wade through them. Two independent briefing lenses draft in parallel and an arbiter reconciles them into one briefing that opens with an at-a-glance strip and leads with the top outcomes: risks, actions, and open questions each color-coded, owners named as people rather than identifiers.

125 1 One briefing of outcomes, objectives, and follow-ups — with every underlying insight kept, attributed, and exportable.
Alderwake recovery-readiness briefing: an at-a-glance strip of three outcomes, four actions, three risks and three open questions, the kept strategic-signal history, and the top three outcomes with named owners.
FIG. 9The settled briefing — two independent lens drafts reconciled by an arbiter, opening with the whole meeting in five seconds.

Nothing raised during the call is quietly dropped. Signals the strategic scan surfaced are kept with how often each came back, so a theme that ran through the conversation reads as a theme rather than one line you happened to catch.

Strategic Signal History expanded in the briefing: six kept signals labelled Action, Discovery Question, Opportunity, Signal and Risk, each showing how many times it was seen and when it was first and last raised.
FIG. 10Durable signals, with first and last sighting — “seen 6 times” is the difference between a passing remark and the thing the meeting was actually about.

The detail is all still there: every one of the 125, each attributed to who said it, ready to take with you as one enriched Excel workbook, HTML, or text.

Insights tab: 125 total, with 2 asked, 24 action items, 16 objections, 18 opportunities, 31 observations, and 34 questions above the answered mid-call questions.
FIG. 11All 125, kept and attributed — 24 action items, 16 objections, 18 opportunities, 31 observations, 34 questions, and the 2 you asked yourself.

What a call costs, to the token

Every model call the session made is recorded. The Tokens tab breaks usage down by source and by model, and prices cached and audio tokens at their own published rates rather than the text rate, so the estimate is honest about where the money went.

1,287,379 $1.14 One 46-minute call. 1,192,838 input and 94,541 output tokens; 176,590 of the input were audio, and 463,839 came back from the provider's prompt cache.
The Tokens tab: estimated cost $1.14 for 1,287,379 tokens, with notes that 176,590 input tokens were audio priced at the audio rate and 463,839 were served from the prompt cache at the cached rate, above a by-source table listing audio_gateway, consolidated_analyst and strategic_signals with their models, input, output, cached and audio columns and per-row cost.
FIG. 12By source, then by model. The live audio gateway is priced at the audio rate; the analyst's 168,203 cached input tokens at the cached rate.

Diarization gets close. You get it right.

Automatic diarization is fast, but it mislabels: splitting one person across several voices, or blending two into one. Rename, merge, and tag speakers by hand, then re-run the analysis so every insight reflects who actually said it.

Rename, merge, and tag

Silero VAD and WeSpeaker embeddings assign a voice to every line before anyone is named, and they don't always match reality. Map each auto ID to a real person, tag your side and theirs, and merge the duplicates a long call inevitably throws off.

Speakers tab: auto-detected speakers with name mapping, team and external tagging, merge controls, and an Enhance Insights button.

Re-run, and every insight updates

Hit Enhance and Backchannel replays the analysis over the corrected speakers — “Me to draft the SOW” instead of “Speaker 4” — so every action item and observation is re-attributed to the right person.

Action items after speaker mapping: each card names the person it belongs to, quotes the line it came from, and flags the one still needing a follow-up.

Ask across every meeting

Every transcript stays local, and stays queryable. Ask a question against one call or all of them at once, and get a grounded answer with the meetings it came from.

Cross-session chat answering what was committed and what is blocking the timeline, with a grounded answer and session scope pickers.
FIG. 13One question asked across every saved meeting, answered with the sessions it came from.

Install the way you want to run it

Use the desktop executable when you want the shortest path, or Docker Compose when you want the full self-hosted stack, development hooks, and GPU options.

Docker Compose

Self-host the full stack

Compose keeps every service isolated and is still the most flexible path for local development, server installs, and optional NVIDIA GPU diarization.

git clone https://github.com/talberthoule/backchannel.git
cd backchannel
cp .env.example .env   # optionally set GEMINI_API_KEY here
docker compose up --build

# app at http://localhost:3000 -- add API keys any time in Admin -> Connections

The first start builds images and downloads models, so give it a few minutes. Have an NVIDIA GPU? Add the override for GPU diarization: docker compose -f docker-compose.yml -f docker-compose.gpu.yml up -d --build

Built-in local batch transcription and diarization need no API key. For insight agents, connect Google, OpenAI, or a self-hosted OpenAI-compatible service, then choose explicit models under Admin. Local recommendations appear after the connected model passes Local Fit.

From spoken word to actionable insight

The live call path, end to end.

  1. Capture. The browser records mic (and optionally tab) audio and streams PCM16 16 kHz chunks over a WebSocket.
  2. Diarize. Voice activity detection cuts speech into segments; speaker embeddings assign each segment to a voice.
  3. Transcribe. Each segment is transcribed by the selected Google, OpenAI, or local ONNX model, filtered, and saved in speaking order.
  4. Shield. Optional. Names, companies and contact details become tokens before the line is stored; the real values go to an encrypted vault on this machine, and the shield holds transcription to a local model.
  5. Analyze. Text agents read the growing transcript on their own schedules and propose questions, objections, opportunities, and action items.
  6. Deliver. Deduplicated insights stream to the call view instantly; a synthesizer refines them as the conversation evolves. With the shield on, tokens turn back into names only on your screen.

Three services, two AI paths, one transcript

A React SPA, a FastAPI backend, and PostgreSQL: a low-latency interim path for feedback and a durable batch path for the record.

Backchannel architecture diagram: client layer, backend layer, AI layer, and persistence

A crew, not a monolith

Each agent has one job, its own trigger, and a configurable model and prompt. Agents start as Not selected; provider-aware Recommended tags offer starting points without forcing a choice. Gemini 3.8 Flash is the recommended Google model for every text agent and Live Ask.

AgentTriggerPurpose
Audio Bridge
audio_gateway
Continuous audio stream Listens silently via Gemini Live, OpenAI Realtime, or the on-device Parakeet captioner and streams interim transcription
Consolidated Analyst
consolidated_analyst
Every 40s + final pass Runs configurable lenses for questions, observations, opportunities, and action items in one pass
Objection Handler
objection_handler
Every 10s over the last 90s Flags objections with an immediate response and the underlying strategic concern
Principal Agent
synthesizer
New or updated insights; 75s cooldown, 120s fallback Quality-checks insights and connects them into broader objectives, initiatives, and patterns
Opportunity Specialist
opportunity_specialist
New opportunities; 55s cooldown + final match Matches opportunity insights against the configured knowledge sources without creating new insights
Strategic Signals
strategic_signals
Every 45s during the call Refreshes the live signal, risk, next question, opportunity, and action cue while linking supported cards to saved insights
Transcript Refiner
transcript_refiner
Every 45s during the call + final pass; off by default Sends the tokenized transcript to a text model, local or cloud, to fix punctuation, casing and mishearings; a rewrite is kept only if it carries exactly the original tokens, so no name ever reaches the model
Briefing Meeting Lens
brief_meeting_lens
At call end or on demand Drafts the meeting record: outcomes, decisions, blockers, commitments, and follow-ups
Briefing Discovery Lens
brief_discovery_lens
At call end or on demand Drafts the broader signal: objectives, pains, gaps, opportunities, risks, and open paths
Briefing Arbiter
brief_arbiter
At call end or on demand Reconciles the two independent lens drafts into the settled briefing

Privacy, data flow, and what it takes to run

The questions that decide whether you self-host, answered plainly.

Is Backchannel free?

Yes. Backchannel is open source under the MIT License, with no hosted tier, no seat pricing, and no feature paywall. Your only costs are your own hardware and any Gemini or OpenAI API usage you configure.

Does my meeting audio leave my machine?

Only where you route it. Voice activity detection and speaker diarization always run locally, and a fresh install uses built-in local batch transcription. Agent models stay unselected until you choose Google, OpenAI, or a self-hosted OpenAI-compatible server such as Ollama or LM Studio. With the PII Shield on, audio is held to a local model outright, because audio cannot be tokenized. Recordings and transcripts stay on your server.

What does the PII Shield keep out of the model?

People's names, company names, email addresses, phone numbers, card and national-ID numbers, IP addresses and street addresses. Each becomes a token such as [PERSON_1] the moment a transcript line, directive, document excerpt, session name or speaker name is written, so every agent, the briefing, chat and Ask work from tokens whether the model is local or in the cloud, and the database holds only tokens. Detection runs entirely on your machine. The real values live in a vault encrypted under the same master key as your provider keys and are put back only on your screen. The shield is off by default; turn it on in Admin -> Privacy.

Does it work with Zoom, Google Meet, Teams, and in-person meetings?

Yes -- any meeting, digital or in person, because it never joins the call. Backchannel captures your microphone and, optionally, tab or system audio directly in the browser, so there is no bot participant and no per-platform integration to install. For a conference room or an in-person conversation, the microphone alone carries the meeting, with diarization separating the voices. Cross-platform coverage is no longer unique: Google's take-notes-for-me now reaches meetings hosted on other providers, and Zoom's assistant can join Meet and Teams. Both do it as a visible bot in a vendor cloud, and neither can sit in a room.

How is it different from Otter.ai and other cloud note-takers?

Real-time help is no longer rare: Otter shipped Live Assist in July 2026 and Zoom's Sales Assist reached general availability a day later. Both are Enterprise-priced and tied to their own platform, and like Fireflies.ai and Granola they run in a vendor cloud. Backchannel is self-hosted, MIT licensed, and joins no call as a bot, and its objection handler generates a response to the objection actually raised, every 10 seconds over the freshest 90 seconds of speech.

Can I ask a question during the meeting?

Yes. The call's command bar answers from the conversation as it stands — the live transcript, the insights raised so far, strategic signals, your directives, and the summaries of attached documents. Recording continues while it answers, the answer is saved with the call and starred, and it exports alongside every other insight. You pick which model answers, including a model running on your own hardware.

Do I need a GPU?

No. CPU-only Docker Compose is the default. An NVIDIA GPU can accelerate diarization via a compose override, and AMD GPUs on Windows are supported with a native backend setup script.

What do I need to run it?

Use the desktop app for the easiest start: download the build for Windows, macOS, or Linux, unpack it, and run the app. Built-in local batch transcription needs no API key. Agents start as Not selected: connect Google, OpenAI, or a self-hosted service if you want its models, then choose explicit models in Admin. Recommended marks a good starting point. Use Docker Compose when you want the full self-hosted stack, local development, or GPU diarization.

Stay in the loop

Follow the project on GitHub

The source, issues, and release tags are all public. Star the repository, watch it for new releases, and read the release notes as they land — every desktop build is a free download, no account needed.

Prefer email? No weekly drip campaign — just release notes and product updates worth opening.