← Manual

Appendix A

Text Style Guide

OpenSpeaks

This guide is for writing down spoken words in language documentation: as captions on screen, as subtitles in another language, or as a full written transcript. Captioning writes what is said, and important non-speech sounds, as on-screen text in the language actually spoken, mainly for deaf and hard-of-hearing viewers. Subtitling translates those captions into a different language. Transcribing writes the spoken words down as a plain document, without time codes.

A.1 Captioning

Captioning helps people who are deaf or hard of hearing read what they cannot hear. It only works for languages that have a script, and for viewers who can read that script. If a language has no script of its own, you can still caption it using the script of a neighbouring language taught locally in school.

Accuracy. Write what you hear: captions should match the spoken words exactly, including slang and informal language, if that is how the speaker actually talks. If a speaker uses words from a different language, set those words in italics.

Timing and format. Keep each line to around 30–35 characters so it stays easy to read; break a long sentence across two lines, or into two or three meaningful parts, rather than cramming it onto one line. "We are keeping the crops in a closed bin in winter," for example, reads better split as "We are keeping the crops" and "in a closed bin in winter." Never show more than two lines of text on screen at once. Time each caption to appear when the speaker begins and disappear when they stop, and keep it on screen long enough to read comfortably, between three and seven seconds. Where a sentence is long, break it where the speaker naturally pauses, never in the middle of a phrase.

Giving background. Use brackets for sounds that matter to the story, for example, (DRUMMING) or (LAUGHING). When more than one person is speaking, name the speaker at the start of their line; if a full name is long, shorten it to a first name plus a surname initial, "Surendra S.P." rather than "Surendra Singh Pangtey," and note the full name separately in the recording's description text.

A.2 Subtitling

This subtitle convention is based on the BBC Subtitle Guidelines (Version 1.2.5, March 2026), itself based on the EBU-TT-D standard. It covers multilingual interviews, where an interviewer and interviewee speak different languages or switch between languages within the same recording.

File format. Use the SRT (.srt) or VTT format for subtitles. Wikimedia Commons uses TimedText, which is close enough to SRT that you can open your SRT file in a plain text editor and copy its contents straight into TimedText. Each subtitle block needs a sequential index number, a timecode in the format HH:MM:SS,mmm --> HH:MM:SS,mmm, one or two lines of subtitle text, and a blank line separating it from the next block:

1
00:00:01,292 --> 00:00:02,001
Your name?

2
00:00:02,210 --> 00:00:03,086
Sukra Dhangdamajhi

Language layers. Subtitle each recording in at least two layers: the source language, the primary language spoken, transcribed as closely as possible, including code-switches (see "Code-switching" below); and a translation language, a more widely spoken language the source subtitles are translated into. Keep each layer as its own SRT file, named LanguageCode-RecordingID-LayerLanguageCode.srt — for example, Bfw-Munaremo-SukraDhangdamajhi.or.srt for the Odia (or) layer of a Bonda (Bfw) recording. Use the ISO 639 code for languages that have one; for languages without an assigned code, use the Glottolog code or the language's full name. Dialects are not usually captured separately in subtitles.

Identifying speakers. The first time each speaker appears in a subtitle block, name them in full, on a separate line above their speech, followed by a colon and a line break. After that, only repeat their name when the speaker changes, or after a gap of 30 seconds or more:

1
00:00:01,292 --> 00:00:02,001
GOBARDHAN PANDA:
Your name?

2
00:00:02,210 --> 00:00:03,086
SUKRA DHANGDAMAJHI:
Sukra Dhangdamajhi

Off-screen speakers. When you can hear a speaker but they are not visible on screen, put a single quote before and after their speech, for example, 'What all do you do?'. If the off-screen speech runs across several lines, keep one quote at the very start and one at the very end.

Inaudible speech. When speech cannot be heard, explain why, in capital letters if you are writing in Latin script: POURS WATER INTO GLASS. INAUDIBLE SPEECH.

Whispering. Label whispered speech the first time it occurs, WHISPERS: I knew it. and drop the label afterwards if the whispering continues. For a long whispered subtitle, put brackets around the whole line instead: (We often pickled raw mangoes in summer and ate throughout the year.)

Code-switching. Code-switching is when a speaker shifts from one language to another within the same speech, whether mid-sentence or across part of a recording; it is common wherever interviewers and interviewees share a language other than the one being documented. Name the language in capital letters where the switch happens:

IN NEPALI: Our worship rituals are known to the elders.

For a switch mid-line, mark it inline in angle brackets at the point it happens:

We speak Raji, {NEPALI}, during our worship rituals

Use a line break once the speaker returns to the primary language. Name the speaker, the first time they appear, alongside the language tag if an off-screen speaker code-switches:

UDAY AALEY: And in which places {NEPALI} do you use the Raji language?

Only mark a genuine language switch this way, not a single loanword, a place or a person's name.

Hesitations, fillers, and overlapping speech. Keep meaningful hesitations, such as "Hmm..." or "Okay...", but do not transcribe every filler if one recurs often. Too many on-screen makes a subtitle unreadable.

Indigenous and community-specific terms. Where a term has no real translation into another language, keep the original word, and set it in capital letters to mark it apart from the rest of the subtitle: I am the DISARI of this village.

Line length and reading speed. Recommended:

  • Maximum two lines per subtitle.
  • Maximum 42 characters per line, including spaces.
  • Minimum display duration: one second.
  • Target reading speed: 160–180 words per minute for a general audience; allow more time for recordings with technical, cultural, or Indigenous vocabulary.

Break lines at a natural pause, especially where a sentence must split mid-way — for example, "We built this small house" and "to keep our domestic animals."

Metadata for each subtitle file. Recording this alongside each SRT file makes the collection usable later. For Wikimedia Commons, add it to the file's own information page. Keep:

  • Primary language: the language spoken, with its ISO code.
  • Subtitle language: the language of this particular SRT file.
  • Speaker(s) and interviewer(s): full names, where possible.
  • Subtitler(s): full names and method, for example, "transcribed from Bonda by X; translated to Odia by Y; English translation and Odia editing by Z."
  • Code-switch language(s), if any, with their codes.
  • Notes: anything else worth recording, including any departure from this convention.

Available in: English

Cite this appendix