Skip to main content
Receiver experience

Read the question, speak the answer.

Keep one written question on screen while the recipient answers in their own voice. Offer writing too, and use a live call when the topic needs real-time follow-up.

Try it on a real question
5 free responses, no credit card

Why the read-and-speak pattern works

Reading is the fastest way to load a question into your head. The words stay on screen the whole time, so nothing has to be remembered. Speaking can be a natural way to explain a story or example. The recipient stays in control of when they answer, while the sender gets a recording they can review alongside a transcript and summary.

Why reading and speaking together beat typing alone

Form-based feedback asks the recipient to do two slow things at once. Read the question, then translate the answer into typed characters on a glass screen. Typing on a phone is the slowest input a literate adult uses in a normal week. It is also the most effortful, which is why so many forms get abandoned halfway.

A phone call removes the typing problem but requires both people to be free at the same time. Use it when the topic needs immediate follow-up, not by default for every first answer.

Read-and-speak keeps the best half of each. The question stays on screen as a visual anchor, no memory load. The answer goes out through the mouth, the channel humans were built for. The recipient stays in control of when they answer, and the sender gets a recording they can scan in seconds.

The cognitive math behind it

Speed depends on the person, language, device, and type of answer. Voice often suits narrative context. Writing often suits short, exact answers that benefit from editing.

The useful distinction is not a universal words-per-minute ratio. It is whether the respondent can explain this particular answer more naturally by speaking than by structuring a written response.

Speaking can preserve spontaneous phrasing and side details. Writing gives the respondent more control over exact wording. Neither format is inherently more honest.

Where the magic compounds: asynchronous

The read-and-speak pattern would still be useful in a live setting, but it gets dramatically better when nothing has to happen at the same time. The sender writes the question once. The receiver answers when their morning is calm enough. The sender listens to the answer when their own morning is calm enough. Three separate windows of attention, no calendar collision.

Submitted voice responses are transcribed and summarized. Use the summary to navigate, the transcript for detail, and the recording as the source when wording or tone matters.

When voice is the wrong choice

Read-and-speak is not for every situation. If your audience is hearing-impaired, voice cuts them out of the conversation. A text-based form is the right call there, full stop. HeySpeak should not be the only option you offer if accessibility is part of the brief.

Voice also struggles in some physical contexts. People on a quiet commuter train, in an open-plan office, or in a library may not record. Keep written responses available. A calendar can appear when the sender configured one and a live conversation makes sense.

And some people just do not like recording themselves. Plan for that. The point is not to force voice on anyone, but to make it available when it helps them answer.

Common questions

Is voice feedback actually faster than typing?
It can be faster for people who already think out loud, especially when the answer is a story or explanation. It is not faster for everyone. Writing is often better for precise names, numbers, sensitive details, or situations where speaking is awkward.
What if the recipient is in a quiet place and cannot speak?
Written replies are available by default, and a calendar appears only when the sender configured one. Recipients can also return later if the link remains active. Do not make voice the only route when privacy or accessibility may be a concern.
Why not just ask people to record a video instead?
Video is useful when facial expression, a physical demonstration, or screen context matters. Voice asks less of the recipient when you only need their words and tone. Choose the format based on the evidence you need.
Does the transcription work for accents and non-native speakers?
HeySpeak currently uses Mistral's Voxtral model. Accuracy varies by language, accent, microphone, and background noise. Keep the original recording as the source and check important names, numbers, and quotes by ear.
What about people who hate voice notes?
Some people will not record, and that is fine. Keep written responses available and add a calendar only when a live conversation is a sensible fallback. The goal is a useful answer, not forcing a preferred medium.
Can the receiver re-listen to their own answer before submitting?
Yes. The receiver can play the recording back and re-record before submitting it.

Send one question. Hear the answer in your inbox.

Five free responses to start. No credit card.

Create your first Magic Link