Voice Mode
Let end users talk to your chatbot Flow or Agent - and have it talk back.
Voice Mode
Voice Mode lets end users hold a real-time, hands-free conversation with your chatbot Flow or Agent. They talk; it transcribes, thinks, and speaks back - with no buttons to press while talking.
It's a great fit for:
- Phone-like concierges - bookings, walk-throughs, support.
- Accessibility - users who would rather listen than read.
- Hands-busy use cases - training in the field, kitchen instructions, in-car assistants.
Where Voice Mode works
Voice Mode is built for chatbot Flows and Agents. Form Flows (which take structured input fields) don't use voice.
Plan feature
Conversational voice is included with the Pro and Agency plans (and AppSumo tier 2 and above). On other plans the voice settings show an upgrade prompt, and end users don't see the voice button.
When voice is enabled and the model produces an artifact - like a generated document or chart - end users still see it on screen. Voice is for the conversational layer, not for replacing visual output.
Enabling Voice Mode
For an Agent
- Open your Agent from Agents in the sidebar.
- Go to the Identity section and find the Vibe controls.
- Below the vibe sliders, toggle Conversational voice on. It is on by default for new Agents.
- Pick a voice from the voice picker, or leave it unset to use the platform default voice.
- Publish your Agent to make the change available to end users.
For a chatbot Flow
Voice Mode is on by default for chatbot Flows on plans that include it, and uses the platform default voice. There is no voice setting in the Flow editor; to choose a voice or tune its delivery, build an Agent.
[Screenshot: Voice toggle and voice picker inside the Agent builder]
Choosing a voice
The voice picker is split into two groups:
- Your voices - Voices you've cloned or generated yourself. These let you build a branded sound.
- Library - A curated set of pre-made voices covering a range of styles and accents.
In both groups you can:
- Search by name.
- Filter by gender, accent, or category.
- Preview a voice with the play button before picking it.
Heads up - Voices come from ElevenLabs. FormWise provides a platform connection, and you can add your own ElevenLabs API key under Settings > Integrations to use the voices in your own ElevenLabs account, including ones you cloned. If voice is greyed out with "Requires an ElevenLabs API key", add a key there.
How Vibe shapes the voice
The Vibe sliders on your Agent (Formal to Casual, Terse to Chatty, Serious to Playful, Predictable to Creative) also nudge the selected voice's delivery. Roughly:
- Playful or casual - more expressive style.
- Serious or formal - steadier, more consistent delivery.
- Creative - a little more variation between turns.
You don't have to tune anything yourself - pick a vibe and a voice, and FormWise handles the blending.
What end users experience
When an end user opens a voice-enabled chat:
- They tap the voice button to start a call. There's no push-to-talk - it listens continuously.
- They see their words appear as a live transcript while they talk.
- The Agent speaks the answer back. As it speaks, a teleprompter lights up the words in time with the audio (karaoke-style). They can hide the teleprompter with an eye toggle if they want a pure-audio experience.
- Tap the orb to interrupt. End users don't have to wait for the Agent to finish - a tap stops playback and lets them speak again. This barge-in feels natural in real conversations.
- A mute button silences their mic; a red end-call button closes the voice session.
- The mic-volume "halo" around the end button gives a visual cue that it can hear them.
[Screenshot: Voice mode active with live transcript and teleprompter]
Tips
- Match voice to vibe. A "warm" vibe with a clipped, businesslike voice feels off. Preview before publishing.
- Keep answers short. Voice answers should sound like answers, not essays. Add a guardrail like "Keep voice responses under three sentences unless the user asks for detail."
- Test on mobile. Most end users will use voice on a phone. Try it on your own device before sharing.
- Plan for noisy rooms. Set expectations in your persona ("If I mishear you, I'll ask again.") rather than fighting it.
Limitations
- Voice needs a browser microphone permission. End users will see a permission prompt the first time they start a call.
- Conversational voice requires the Pro or Agency plan.
- Voice does not run inside Form Flows - only chatbot Flows and Agents.
- Cloned voices follow ElevenLabs' usage rules. Make sure you have rights to any voice you clone.