VibeVoice ASR

VibeVoice ASR for 60-Minute Speech-to-Text

Up to 60 minutes · Speaker labels · Timestamps · 50+ languages

Upload Audio

Add a meeting, podcast, interview, lecture, or other long-form recording to transcribe.

No file selected

Custom Hotwords

Enter names, brands, acronyms, technical terms, or other context that may appear in the recording.

Structured Output

Transcript Preview

Who · When · What
00:00:03Speaker 1Welcome, everyone. Today we’re reviewing the product roadmap and the next release.
00:00:14Speaker 2I added VibeVoice ASR and the key product terms to the hotword list before transcription.
00:00:25Speaker 1Great. Let’s capture the decisions and assign the follow-up items.
Explore VibeVoice ASR ↓

What Is VibeVoice ASR?

VibeVoice ASR is a unified speech recognition model for long-form audio. It combines automatic speech recognition, speaker diarization, and timestamp generation instead of treating them as separate tasks.

The VibeVoice ASR output answers three practical questions:

Whowhich speaker is talking
Whenwhen each segment occurs
Whatwhat the speaker says

Because these elements are processed together, users receive a structured transcript instead of one long block of text. VibeVoice ASR also supports customized hotwords, multilingual speech recognition, and code-switching speech recognition.

VibeVoice Speech to Text for Long-Form Audio

VibeVoice ASR provides a unified workflow for long-form speech recognition. Instead of requiring users to divide an hour-long recording into many short clips, the model can process up to 60 minutes of continuous audio within a 64K-token context.

This longer context helps VibeVoice ASR follow ideas across a full conversation, maintain more consistent speaker tracking, and preserve connections between earlier and later sections. It is useful for meetings, podcasts, interviews, lectures, webinars, and customer calls.

A VibeVoice transcription includes recognized speech, timestamps, and speaker labels in one structured output. For users searching for VibeVoice speech to text, VibeVoice audio transcription, or long audio to text, the benefit is a simpler long-audio workflow.

Key VibeVoice ASR Features

60-Minute Audio Transcription in One Pass

VibeVoice ASR supports 60-minute audio transcription without requiring users to manually split a recording into short sections. It accepts up to one hour of continuous audio in a single pass, helping preserve wider conversational context. VibeVoice ASR can use more of the recording when interpreting references, shortened names, and repeated topics. Users can upload a longer file and review one connected result.

Speech to Text with Speaker Labels

VibeVoice ASR creates speech to text with speaker labels so readers can distinguish participants in the same recording. A transcript may separate voices as Speaker 1, Speaker 2, Speaker 3, and so on. VibeVoice ASR is especially useful for meeting transcription with speaker labels and multi-speaker audio transcription.

Speaker Diarization Transcription

Speaker diarization transcription determines when different voices are active and assigns separate labels to their speech. VibeVoice ASR combines diarization with recognition and timestamping, reducing the need to run multiple tools. The model does not verify a person’s real-world identity. It separates speakers and applies labels that users can review or rename later. It makes long conversations easier to scan.

Speech to Text with Timestamps

VibeVoice ASR generates speech to text with timestamps, connecting each transcript segment to its approximate position in the recording. Users can return to an important comment without replaying the entire file. They can also create podcast chapters, review quotations, prepare captions, locate clips, and organize meeting summaries.

Customized Hotwords ASR

VibeVoice ASR supports custom vocabulary and background context. Customized hotwords ASR allows users to provide names, brands, acronyms, technical terms, product names, locations, or other words that may be difficult to recognize. Before starting VibeVoice ASR, users can enter terms likely to appear in the recording. This can improve recognition of uncommon vocabulary in medical, legal, technical, academic, or business content. Important spelling and terminology should still be reviewed before publication.

Multilingual Speech Recognition

VibeVoice ASR supports multilingual speech recognition across more than 50 languages. Users do not need to manually choose one language for every recording. This makes the model useful for international meetings, multilingual interviews, global webinars, and regional content.

Code-Switching Speech Recognition

Code-switching speech recognition matters when speakers move between languages during the same sentence or conversation. VibeVoice ASR is designed to recognize these changes within and across spoken segments. For bilingual podcasts, international interviews, and multilingual meetings, the model can reduce the need to separate a recording by language before transcription.

Automatic Speaker Identification

VibeVoice ASR provides automatic speaker identification by separating voices and assigning labels throughout the transcript. Here, automatic speaker identification means distinguishing participants, not confirming their real identities. Speaker labels can be renamed after transcription when the user knows who each participant is. This makes VibeVoice ASR easier to use for meeting records, interview notes, podcast scripts, and research documentation.

Why Long-Form Speech Recognition Matters

A speaker may introduce a name near the beginning, shorten it later, and return to the same subject near the end.

VibeVoice ASR keeps more of the conversation available during processing. By handling up to 60 minutes in one pass, it is designed to maintain stronger semantic continuity and more consistent speaker tracking.

This approach is valuable for users who need context across a complete meeting, podcast, interview, lecture, or webinar. The model also simplifies file preparation because users do not need to manage many separate clips.

VibeVoice ASR Use Cases

Meeting Transcription with Speaker Labels

Use VibeVoice ASR for planning sessions, project reviews, customer calls, and team discussions. It can help teams identify who made a suggestion, when a decision occurred, and what follow-up work was discussed. Custom hotwords can improve recognition of internal names and project terminology.

Podcast Transcription

VibeVoice ASR supports podcast transcription for solo shows, interviews, panel episodes, and multi-host programs. The transcript can be reused for show notes, chapters, articles, captions, social posts, quotations, and searchable archives. Speaker labels separate hosts and guests, while timestamps help editors locate important moments.

Interview Transcription

VibeVoice ASR makes interview transcription easier for journalists, researchers, recruiters, and content teams. It can separate questions from answers, add timestamps, and preserve the structure of longer conversations. Users can add names, organizations, locations, and specialist terms as hotwords.

Lecture Transcription

Use VibeVoice ASR for lecture transcription, seminars, classroom discussions, and training recordings. A transcript can support study notes, searchable learning materials, accessibility workflows, captions, and lesson summaries. Educators can add course names and technical vocabulary as custom hotwords.

Webinar Transcription

VibeVoice ASR supports webinar transcription for presentations, product demonstrations, online events, and audience Q&A sessions. It can separate presenters, moderators, and participants while adding timestamps for key topics. Teams can reuse the output for documentation, articles, FAQs, and summaries.

Multi-Speaker Audio Transcription

VibeVoice ASR is suitable for multi-speaker audio transcription across panel discussions, meetings, podcasts, interviews, focus groups, and customer research. Because the model combines recognition, diarization, and timestamps, users receive one structured result showing who spoke, when they spoke, and what they said.

How to Use VibeVoice ASR

  1. 1.

    Upload Your Audio

    Choose the meeting, podcast, interview, lecture, webinar, or other recording you want to transcribe with VibeVoice ASR. Clear audio with consistent volume and limited background noise usually produces better results.

  2. 2.

    Add Custom Hotwords

    Enter names, brands, technical terms, acronyms, locations, and background information before starting VibeVoice ASR. This optional step can improve recognition of uncommon vocabulary.

  3. 3.

    Start the Transcription

    Run VibeVoice ASR to process the audio. The model recognizes speech, separates speakers, and generates timestamps within the same workflow.

  4. 4.

    Review the Transcript

    Check the VibeVoice ASR output for speaker labels, timestamps, names, numbers, technical terms, language changes, and overlapping speech. Use the reviewed result for notes, summaries, captions, chapters, research, or documentation.

VibeVoice ASR vs. Traditional Speech-to-Text

CapabilityVibeVoice ASRTraditional Workflow
Continuous audio inputUp to 60 minutesOften split into short chunks
Long-context processing64K-token contextVaries by system
Speech recognitionIncludedIncluded
Speaker diarizationIncludedOften separate
TimestampsIncludedMay require alignment
Custom vocabularyCustomized hotwordsSupport varies
LanguagesMore than 50Depends on the model
Code-switchingSupportedSupport varies
OutputWho, When, WhatOften plain text

VibeVoice ASR is best suited to long or multi-speaker audio where context, speaker labels, timestamps, language changes, and specialist vocabulary matter.

VibeVoice ASR FAQ

What is VibeVoice ASR?

VibeVoice ASR is a long-form speech recognition model that converts audio into structured text with speaker labels, timestamps, and recognized content.

How much audio can it process?

VibeVoice ASR is designed to process up to 60 minutes of continuous audio in one pass.

Does it identify different speakers?

Yes. VibeVoice ASR performs speaker diarization and assigns separate labels to different voices.

Does it generate timestamps?

Yes. VibeVoice ASR includes timestamps as part of its structured transcription output.

Can I add custom vocabulary?

Yes. VibeVoice ASR supports customized hotwords such as names, brands, technical terms, acronyms, and background information.

Does it support multiple languages?

Yes. VibeVoice ASR supports more than 50 languages and can handle code-switching within and across spoken segments.

Is it the same as VibeVoice TTS?

No. VibeVoice ASR converts spoken audio into text, while VibeVoice TTS converts written text into generated speech.

What audio can I transcribe?

VibeVoice ASR can be used for meetings, podcasts, interviews, lectures, webinars, panel discussions, customer calls, and other long-form recordings.

Transcribe Long Audio with VibeVoice ASR

Use VibeVoice ASR to turn up to 60 minutes of audio into structured text with speaker labels, timestamps, multilingual recognition, code-switching support, and customized hotwords.

Start Transcribing