Docs/Tutorials/Talk instead of typing
Tutorials

Talk to your agents instead of typing

6 min read

The instruction you would actually give a person - all the context, the thing to avoid, the file you already looked at - is the one you never type. You type six words instead and the agent guesses the rest. Said out loud it takes fifteen seconds.

This walks you through dictating one message, about five minutes. You need DevThrottle open with a Gateway connected, a session selected, a microphone, and an account with Pro or the 14-day Pro trial active - dictation is part of Pro, and every new account starts with 14 days of it. Nothing to install and no key to supply. If your trial has ended and you are not on Pro, the dictation box tells you so instead of recording. It is the deeper half of The prompt bar is not the terminal, which introduces Speak as one more way of filling the same box.

Note
One thing before you start, because it looks like a fault: while you are talking, no text appears. That is deliberate. Your words are kept as audio and turned into text in one piece when you ask for them, which is more accurate than transcribing as you go.
  1. Press Speak

    The green Speak button sits beside the message box under the terminal, stacked between Send and Queue. Ctrl+H does the same thing from anywhere in the window, as long as the prompt bar is showing - so a session has to be selected, because otherwise there is nothing to dictate into.

    The prompt bar below the terminal: the message box with the Expand button in its corner, and Send, Speak and Queue stacked down its right-hand side, the Queue button red and reading Queue (3)
    The prompt bar below the terminal: the message box with the Expand button in its corner, and Send, Speak and Queue stacked down its right-hand side, the Queue button red and reading Queue (3)

    A box opens and says GETTING READY until your microphone actually delivers audio, then turns red and plays a short sound. If the wrong microphone is selected, the dropdown at the top of the box changes it.

  2. Say the whole thing

    Talk normally, in full sentences, the way you would explain it to someone at your desk. Bars move and a timer counts, and the text area stays empty - that is the part to trust. Pausing to think is fine; the timer is only telling you how long you have been talking.

  3. Press Pause to read back what you said

    Pause stops the microphone and turns everything you have said so far into text, in the box, where you can edit it by hand. It is a checkpoint rather than an ending: Resume starts a fresh stretch of talking that gets added onto what is already there, edits included. Resume stays greyed out until the text has landed.

    This is what makes a long instruction workable. Say a paragraph, pause, check it got the file name right, carry on.

  4. Choose Insert or Send

    Insert puts the text into the prompt box at your cursor and sends nothing - your typed words and your spoken words end up in one message, and you read it before the agent does. Send submits it into the session for you.

    Press Send while you are still talking and the box closes immediately: the transcribing happens afterwards, in the background, and the session shows orange with Transcribing... while it does. You do not sit and watch a progress bar. Enter is Send and Escape is Cancel. Cancel closes the box without putting anything in the prompt or sending anything to the agent - though it does not reach back and erase pieces you already checkpointed with Pause, which have been transcribed by then. The Voice page says what is kept after a transcription and for how long.

    Tip
    Use Insert for your first week. Reading what came back is how you learn which of your words it fumbles - and those are exactly the words to put in your dictionary. Switch to Send once you have stopped being surprised.
  5. Do the same thing in a browser and on your phone

    It is one dictation box, so there is nothing new to learn on the other two. In the Cockpit, open a session and Speak is beside Send under the message box. On your phone, open a session's terminal and tap Keys to bring up the controls - Speak sits next to Send there too.

    The same four buttons, doing the same four things. The one difference is that neither browser surface has a microphone dropdown; they use whichever microphone the browser is already set to.

Warning
On the desktop, while a dictated message is still transcribing, that session's prompt box, Send, Speak and Queue go grey - one dictation at a time, so two cannot land on top of each other. It is not stuck; it clears itself when the words go in. The brakes are never taken away: Stop and Interrupt - whichever of the two your agent supports; not every coding agent has both - stay live throughout.

What happens if it cannot get through

The two answers are different, on purpose, and it is worth knowing which you are standing in.

On your phone and in the browser, the recording is saved on the device before any network work happens, so a dropped signal does not lose it. While the page is open, delivery retries on its own; a closed tab or a phone that went to sleep does not deliver in the background, but nothing is lost either - the recording waits on the device, and delivery picks up the moment you are back in the app with a connection. That is what makes dictating on a walk sensible: say it, and it is safe; it goes while the app is up and there is signal, or the next time you open it.

A status strip under the session tells you where it has got to, and it tells you the awkward cases too rather than going quiet. A send that cannot get through yet shows as waiting - saved, still trying - with an Upload now button to push it the moment you have bars. A failure that will not fix itself by retrying shows Retry - the recording is still safe on the device. The one case that can actually lose a recording is when it could not even be saved on the device in the first place; that shows as a red alert that stays until you dismiss it, so it cannot slip past you.

The desktop app deliberately does not do that. A dictation sent there either goes into the session immediately or you are told immediately, with a message you cannot miss and no quiet queue retrying behind you. If the message had already been turned into text, the whole of it is put back in the prompt box so you can send it again yourself. If the failure came before the words could become text, what you had typed comes back, and the app tries to keep the last stretch of recording - when that save succeeded, the message says where to find it - but text you had already checkpointed with Pause on that dictation is not put back on this path. Dictating again is the way back.

When it keeps getting the same word wrong

It will get your product name wrong, or your repository, or the library nobody outside your team has heard of. Fixing that by hand every time is the thing that makes people give up on dictation, and it takes about five minutes to fix permanently: Teach it the words you use.

Next

For what dictation is, which surfaces it runs on, what it costs, and what leaves your machine when you use it, see Voice. If you want your agents to talk back as well as listen, that is voice mode on the phone.