Skip to content

Turn text into speech

View as markdown

Bearly can read an answer aloud or record text as audio you can download. Use Read aloud when you simply want to listen. Use Voice, under the latest message in a chat, when you want to choose the voice, direct the delivery, or compare takes.

After Bearly returns a text answer, select Read aloud beneath it. Select it again to stop playback.

You can also select Hear this reply on the Voice bar under the latest message. Bearly records the answer in your current voice and keeps the take in the chat so you can play or download it later.

If playback does not begin, check that the device is not muted and that the browser or operating system allows audio from Bearly.

  1. Open Voice

    Under the latest message, select the Voice bar to expand it.

  2. Pick what you’re making

    Under What are you making?, choose Video voiceover, Podcast intro, Audiobook, Proofread by ear, Announcement, or Wind-down. Bearly picks a model, a voice, and a delivery direction to start from. You can change any of them.

  3. Add the script

    Type or paste the words to speak. To start from the conversation, select Use reply for Bearly’s latest answer or Use my message for what you sent.

  4. Choose a voice

    Scroll the voices under Cast. Select ▶ on a card to hear that voice read the first sentence of your script, then select the card to use it. Voices suggested for your choice in step 2 come first and are marked with a dot.

  5. Direct the delivery

    Describe how it should sound, such as “warm and unhurried,” or select a suggestion.

  6. Generate

    Select Generate for one take, or Compare 3 to record the same script in three voices at once. With a keyboard, you can also press ⌘ Return on a Mac or Ctrl Enter elsewhere.

The take appears in a player with its script. As it plays, the words highlight in turn so you can follow along; the timing is estimated, so it can drift slightly from the voice. Earlier takes are listed below the player. Select one to play it again. Use the download control to save the take: Gemini voices save as WAV files and OpenAI voices as MP3.

The switch at the top of Voice sets the model:

  • Fast uses Gemini Flash-Lite. It is the quickest and suits drafts, proofreading, and everyday listening.
  • Studio uses Gemini Flash for the most expressive delivery, such as voiceovers and narration.
  • OpenAI uses GPT-4o Mini and its own set of voices.

Fast and Studio share the same 30 voices. Each card shows the voice’s character, such as Firm or Warm.

Turn on Use for Read Aloud at the bottom of Voice to have Read aloud use the model, voice, and delivery you chose. With it off, Read aloud keeps its standard voice.

Type /tts in a chat, then send the text with a short note on how it should sound. Bearly adds an audio player to the conversation with a download control. Use Voice when you want to pick the exact voice or compare takes.

  • If a name or unusual term is mispronounced, spell it phonetically in the script and generate another take.
  • If the delivery feels exaggerated, simplify the direction to one or two qualities.
  • If generation fails, shorten a very long passage and try again after checking the connection.
  • Voice samples and takes count toward your AI usage like other generations. Compare 3 records three takes.
  • If an older recording says it is no longer available, generate it again from the original text. Download important finished recordings rather than relying on an old chat as permanent audio storage.

Team administrators can turn text to speech off for members with a team policy.