Tutorials

How to transcribe audio or video without coding

A practical guide to tools for working with interviews and other archive material (both audio and video files) whether they're hosted online (e.g., YouTube) or sitting on your own computer. The focus here is on getting a usable transcript of what was said, not on visual analysis of the footage.

YouTube videos

In general, all the methods below rely on the same underlying source: YouTube's own captions either uploaded by the creator or auto-generated by YouTube. Nothing here independently "listens" to the audio; the caption text is what everything is built on. This matters because caption quality (or its absence) sets a ceiling on how good any of these results can be.

Gemini 

Recommended - The only major consumer LLM that extracts and transcribes the video itself

  • Paste the YouTube link directly into a chat, no manual steps needed.

  • Can produce several different output formats depending on what you ask for: a verbatim transcript, a verbatim transcript with speaker labels, a literary-edited (cleaned-up) transcript, or a brief summary of general topics.

  • It still relies on the underlying captions, but it can also correlate them with the video itself to add speaker labels — something the captions alone don't provide. With several speakers, this correlation can still get it wrong.

Manual extraction + other LLMs (Claude, ChatGPT, DeepSeek, Grok)

  1. Open the video on YouTube.

  2. Below the video, click the "...more" (or "...show more") text under the description.

  3. Click "Show transcript".

  4. The transcript appears in a panel next to the video, with timestamps for each line.

  5. Select all the text and copy it, or use the three-dot menu on the transcript panel to toggle timestamps on/off before copying.

  6. Paste the copied text into any LLM for summarizing or analysis.

Notes and limitations:

  • Only works if the video has captions available at all.

  • Auto-generated captions can have mistakes: no punctuation, misheard words, no indication of who's speaking. Quality varies a lot depending on audio clarity and accent.

  • YouTube's captions do not distinguish between speakers — an interview with several people comes out as one continuous stream of text with no speaker labels. If you need to know who said what, you'll need a separate diarization step (see the Whisper note below) or to mark speakers manually.

  • If a video has no captions at all, this method won't work — you'll need Whisper (see below) after downloading the audio.

Once you have the transcript, paste it into Claude, ChatGPT, DeepSeek, or Grok (only for English texts) to get summary or main topics of the audio.

Local files - Whisper

Whisper (OpenAI's open-source speech-to-text model) turns any local video or audio file into a text transcript, entirely offline and free — useful when there are no existing captions to fall back on.

Buzz Desktop (Free, Open Source) - Recommended

Website: Buzz website

GitHub: Buzz GitHub repository

How to use:

  1. Download and install Buzz.

  2. Open the app.

  3. Drag and drop your audio file.

  4. Choose a Whisper model (e.g. medium or large-v3 for the best accuracy).

  5. Click Transcribe.

  6. Export the transcript as TXT, SRT, or VTT.

Whisper Desktop

GitHub: Whisper Desktop GitHub repository

How to use:

  1. Download the latest release.

  2. Install and launch the application.

  3. Open your audio file.

  4. Select the language (or Auto Detect).

  5. Click Transcribe.

  6. Save the transcript.

Notes:

  • For long interviews or meetings, I'd recommend Buzz. It's actively maintained, works offline, supports speaker identification, and provides a polished interface powered by Whisper.

  • Model size is a trade-off: smaller models (tiny/base) are faster but less accurate; larger ones (medium/large) are slower but much more accurate.

  • Whisper doesn't distinguish speakers. If you need to know who said what, that requires a separate diarization step (e.g., pyannote.audio).

OpenAI ChatGPT

On paid accounts, an option to upload a video directly is gradually starting to appear, but it's not yet widely or reliably available — access varies by plan, app, and rollout stage.



CONTACT US
 INSTITUTE FOR EUROPEAN, RUSSIAN AND EURASIAN STUDIES 1957 E St NW Washington, DC 20052

1957 E St., NW, Suite 412,
Washington, DC 20052

russiaprogram@gwu.edu
+1 (202) 9946340

CONTACT US
 INSTITUTE FOR EUROPEAN, RUSSIAN AND EURASIAN STUDIES 1957 E St NW Washington, DC 20052

1957 E St., NW, Suite 412,
Washington, DC 20052

russiaprogram@gwu.edu
+1 (202) 9946340