Skip to content Accessibility statement

Transcribing audio and video recordings

Convert your audio and video recordings into written text. 

These tools support multiple languages and work best with structured dialogue, such as interviews, involving fewer speakers. Use this page to find the right tool for your data's security classification, file format, usage limits and access requirements.

See our separate guide if you want to transcribe meetings in progress (Google Doc).

Before you begin

Before processing any files, you must:

Privacy and information handling

In order to keep University data secure and comply with data protection laws, staff and students must only use GenAI tools that are provided and licensed by the University. This ensures appropriate contractual provisions are in place to protect University data. Do not use personal accounts with the same or alternative providers as these do not provide the same level of protections.

If you are processing personal data that is likely to result in a high risk to individuals’ interests you must undertake the Data Protection Impact Assessment (DPIA) screening process. This will help to ensure compliance with Data Protection legislation. For support with this, please contact the Data Protection team.

Approved tools

Remember, AI can make mistakes. After processing your recording, it's important that you review and verify the transcription for accuracy.

Google Gemini (Standard and Pro licences)

  • Approved data classification: Public, Internal, Confidential.
  • Data processing: The data is processed in Google's cloud – they do not guarantee UK or EU geolocation, but it is protected by the terms of the Google Workspace agreement. Your data is not used for training or human review.
  • Access and licencing: Standard use is free to all staff and students using their University account. The Pro version is funded by departments – to discuss your needs contact IT Services.
    • Standard limits: Up to ten minutes of audio total per prompt. You can upload up to ten audio files per prompt (under 100 MB each), but their combined duration cannot exceed ten minutes.
    • Pro limits: Up to three hours of audio per prompt.
  • Restrictions: There is no daily or monthly cap, but running multiple file chats back-to-back can temporarily trigger a usage limit error.

Explore Gemini

Otter.ai (Pro and Business licences)

  • Approved data classification: Public, Internal.
  • Data processing: AWS West (United States) using Standard Contractual Clauses. Cyber Risk Assessment (CRA) approved.
  • Access and licencing: You must purchase a paid Pro or Business plan using your University of York Google account.
    • Pro licence: Includes 1,200 in-app recording minutes per month, ten monthly file imports and up to 90 minutes per meeting.
    • Business licence: Includes unlimited meetings, unlimited in-app recordings, unlimited file imports and up to four hours per meeting.
  • The free version is not permitted for University use.
  • Approved only for transcribing pre-recorded audio or video files. Do not use for live transcription or captioning.

Explore Otter.ai

Caption.ed

  • Approved data classification: Public.
  • Data processing: Google data centres in the UK.
  • Access and licencing:
    • Funded through Access to Work or Disabled Student Allowance (DSA) based on individual needs.
    • You must enable Multi-Factor Authentication (MFA) to access the service.
  • Capabilities:
    • Provides live captioning, synchronised transcripts, timestamped notes and automated media saving.
    • Uses AI to generate structured summaries and chapter divisions.

Explore Caption.ed

OpenAI Whisper (Local)

  • Approved data classification: Public, Internal, Confidential (University-managed devices only).
  • Data processing: Processed entirely on your local device.
Access and setup instructions

If you're already using Homebrew on your Mac, installation is as simple as:

% brew install openai-whisper

You can then transcribe audio files:

% whisper audio-sample-1.mp3

For further help, such as setting it up on Windows, please contact IT Services.

OpenAI Whisper (Viking)

  • Approved data classification: Public, Internal, Confidential.
  • Data processing location: Sweden (on University of York infrastructure supported by Alces Flight). Data is not used for model training or review.
  • Access: Requires an active account on the Viking cluster linked to a project. (For help, see Viking documentation.) Accessible via command line or web browser using OpenOnDemand.
  • Capabilities: Supports WhisperX and CrisperWhisper modules with unlimited speech processing and no daily usage caps.
  • Data handling: Files are stored on Viking storage in Sweden during processing before you transfer them back to local University storage or Google Drive.