Wikier

Speech to text

Speech to text is an artificial intelligence-based automatic transcription service, that can be used for public and internal data. On this page you will find information about what you can use the service for and how to use it.

Topic page on research data | Topic page on software

Norwegian version: Tale til tekst

Illustration with laptop, speech bubble and transcribed document

What is Speech to text?

Speech to text is a service for automatic transcription using AI, developed at NTNU. Two different language models are available: Whisper from OpenAI and NB-Whisper Large from the National Library. Whisper recognizes 98 languages, including Norwegian and English. NB-Whisper only recognizes Norwegian and English, and is recommended for Norwegian. You can either transcribe text in the same language or translate into English.

Speech to text can be used to transcribe most types of audio and video files containing public (green) or internal (yellow) data, such as audio recordings, recordings from Zoom or Panopto. The file type .opus unfortunately does not work. The service can also be used to streamline the subtitling of video. Read more about tools for captioning video.

The quality of the automatic transcription with Whisper will vary, and the text should be reviewed and corrected manually. Both the sound quality of the recording and the language or dialect to be transcribed, will affect the accuracy of the transcript.

Get started

 Log in to Speech to text

  1. Log in to Speech to text with your NTNU user (Feide login). You must be connected to NTNU's network, either directly from campus or via VPN.
  2. Choose language and model. NB-Whisper is recommended for Norwegian. OpenAIs Whisper allows you to choose automatic recognition. Check the box if you want the audio files translated into English.
  3. Select the audio or video files you want to transcribe by dragging over files or clicking "browse files". Remember to read the terms of use and check that you have classified your data. The transcription process may take some time, depending on how long the queue is. You will receive an e-mail when your transcription is complete.
  4. Download the files in the desired format: txt, vtt or srt. Choose srt-format if you are going to add captions to the video.

Uploaded audio/video files are deleted when the transcription is complete. Transcribed files are automatically deleted after 14 days, unless you delete them yourself.

Guidelines for use

Classification of personal data

Personal data - including research data containing personal data - is usually classified as internal or confidential. Speech to text may NOT be used on material containing special categories of personal data, as this type of data is classified as confidential. Special categories include topics such as health, religion, race or political opinion. In addition, trade secrets or research subject to export controls will usually be classified as confidential.

Please note that all audio in uploaded files will be transcribed. Be aware if confidential or sensitive topics are recorded, even if they are unforeseen or small in scope. See NTNU's guidelines for information classification and data storage for more information.

Data management and privacy

Speech to text is set up on the servers and infrastructure of NTNU's IT department. Your audio or video content does not leave NTNU, and is not shared with others. The service is provided by the Research Data Project at NTNU, which has also conducted a risk assessment and privacy impact assessment (DPIA).

See the privacy statement for Speech to text.

Contact us

Contact NTNU Hjelp if you have questions about using Speech to text.

If you have questions about research data, visit Forskningsdatahjelpen (Research Data @NTNU).