Blog••Journalism

Focus Group Transcription: Formats, Speaker Labels and Turnaround

Sarah Johnson

Sarah Johnson

Writes about field sales, meeting notes and voice-first workflows at ParrotNotes. Every article is reviewed by the ParrotNotes product team before it goes live.

Focus Group Transcription: Formats, Speaker Labels and Turnaround

Dana ran six focus groups for a regional grocery chain: eight shoppers per session, 90 minutes each. The recordings sounded fine in the room. By session three of transcribing, she couldn't tell the two retired teachers apart, and one heated minute about delivery fees had four people talking at once. (Dana is a composite, not a real researcher.)

That's what makes focus group transcription harder than transcribing an interview. One interview has two voices. A focus group has eight, they interrupt each other, and half the value sits in who agreed with whom.

The good news: most of the fix happens before anyone types a word. Below: which format to pick, how to set up the room for clean speaker labels, a template you can copy, a turnaround planner and how to anonymize a group transcript.

Running groups this month? Get ParrotNotes free, record on the phone you already own, and start every transcript from a draft instead of a blank page.

How focus group transcription differs from an interview transcript

The basic job is the same: accurate, searchable text. Our guide to how to transcribe an interview covers the correction pass that applies to any recording.

Groups add four problems:

  • Many speakers. The University of Wisconsin Extension's focus group tip sheet says there are usually 8 to 12 participants, and that sessions typically run one to two hours. That's a lot of similar voices to keep apart.
  • Overlap. People agree, interrupt and laugh over each other. Crosstalk is where words go missing.
  • Interaction is data. In an interview you mostly need what was said. In a group you also need who said it, who pushed back and whether the room agreed.
  • Shared confidentiality. Everyone in the room hears everyone else, which changes consent and anonymizing.

So focus group transcription is only as good as its speaker labels. Plan for them first.

Choose a format for focus group transcription

Pick the level of detail from what you'll do with the transcript. For a full breakdown of the styles, see verbatim vs clean verbatim transcription.

Study purposeFormatKeepDrop
Academic thematic analysisVerbatim, light notationEvery word, laughter, (crosstalk), inaudible marksMost "um" and "uh"
Discourse or interaction analysisClose verbatimFillers, false starts, overlaps, who interrupts whomLittle
Market research or UX reportClean verbatimEvery point and quote, speaker labelsFillers, false starts, repeated words
Internal debrief or program reviewSummary notes plus key quotesThemes, agreements, disagreements, quotable linesWord-for-word record

Keep speaker labels in every format. Without them, you know a point was made, but not whether one loud participant made it five times. Running one-to-one interviews too? Our qualitative interview transcription guide has a full notation protocol.

Before the session: set up for clean speaker labels

Labels go wrong in the room, not at the keyboard. Four habits fix most of it:

  1. Draw a seating chart. Number seats clockwise from the moderator (P1, P2...) and keep the chart with the recording.
  2. Run an introductions round on tape. Ask each person to say their first name or ID and one sentence about themselves. Those first 60 seconds become the reference sample for every voice in the session.
  3. Place the recorder in the middle. Lay the phone flat on the table, equal distance from everyone, away from cups and air vents. For a long table, use two devices. Our guide to how to record an interview covers test recordings and room noise.
  4. Use a note-taker. The Wisconsin tip sheet recommends an assistant moderator who records "an identifier of who said what." That note is gold when the audio is unclear.

Consent from every participant

Everyone in the room is recorded, so everyone needs to agree, in writing and again on tape. The Teachers College, Columbia University IRB advises that the consent form should disclose which activities will be recorded and how the recordings will be stored and used.

The same guidance makes a point unique to groups: "researchers cannot guarantee confidentiality to their participants," because other participants hear everything. Say so on the form, and ask the group to keep what's said in the room.

For phone or video sessions with people in different US states, recording laws vary. See our overview of one-party and two-party consent states, and check with your ethics board or legal team.

Speaker labels with 6 to 10 voices

Agree on these rules before the first session, and use them in every transcript in the study:

  • MOD for the moderator, AM for the assistant moderator, P1 to P10 for participants by seat.
  • One speaker per paragraph. A new paragraph every time the voice changes.
  • Never rename mid-study. P4 in session 2 is the seat, not the person. Keep a roster that maps each session's IDs to your participant codes.
  • Unsure? Mark it. Use P? with a timestamp rather than guessing.
  • Crosstalk gets a tag. Write (crosstalk) where several people talk at once, then transcribe the words you can attribute.
  • Group reactions count. Note (several agree), (laughter) or (P2 and P5 shake heads, per notes) when the reaction matters.

The note-taker's speaker log

While the session runs, the assistant moderator writes one line every time a new person starts a point:

TimeSeatFirst wordsNote
00:14:05P3"Honestly the app is..."Strongly against delivery fee
00:14:40P6"I'd pay it if..."Agrees, with a condition
00:15:02P1 + P4(crosstalk)Both laughing, P1 louder

When the draft comes back, match each time and first words to the transcript. A wrong label takes seconds to fix, instead of replaying audio to guess between two similar voices.

Before your next session: download ParrotNotes free and make a two-minute test recording from every seat, so you know each voice will be audible.

Timestamps and a focus group transcript template

Choose one timestamp rule and keep it: at every speaker change, or every two to three minutes for long sessions. Use [hh:mm:ss] so any team member can jump straight to the audio.

Copy this template to the top of every transcript:

FOCUS GROUP TRANSCRIPT
Study:            [study name or code]
Session:          [3 of 6]
Date and mode:    [2026-10-11, in person / video]
Recording file:   [FG03_2026-10-11.m4a, 01:31:12]
Format:           [clean verbatim, speaker labels, timestamps per turn]
Moderator:        MOD = [initials]   Note-taker: AM = [initials]
Transcribed by:   [name or "AI draft", date]
Checked against audio and speaker log by: [name, date]
Anonymized:       [yes, log ref ANON-FG03 / no]
Consent:          [form version; all 8 confirmed on tape at 00:00:40]

SPEAKER ROSTER (by seat, clockwise from MOD)
P1 = [code F-01]  P2 = [F-02]  P3 = [F-03]  P4 = [F-04]
P5 = [F-05]  P6 = [F-06]  P7 = [F-07]  P8 = [F-08]

---

MOD: [00:12:30] Let's talk about delivery. What would make you pay for it?

P3: [00:14:05] Honestly the app is fine. It's the fee I won't pay.

P6: [00:14:40] I'd pay it if it came the same day.

(crosstalk, P1 and P4) [00:15:02]

P1: [00:15:06] Same day, sure, but not five dollars. (laughter)

Keep one turn per paragraph, so analysis tools and spreadsheets can treat each turn as a row.

Turnaround: how long focus group transcription takes

Turnaround depends on two things: the draft, and the checking. Drafts are fast now. Checking is where the hours go.

In a 2024 study in the European Journal of Cardiovascular Nursing, Helen Eftekhari transcribed one-to-one interviews of 40 to 131 minutes with speech recognition in Microsoft Teams. Checking and cleaning took "between 1.5 and 3.5 h depending on accuracy and interview length." That was with two voices. A 90-minute group with eight voices won't take less.

Here's a planner for Dana's six-session study. The per-session hours are our planning assumption for groups, set above the study's interview range, not a measured figure:

StepPer 90-minute sessionSix sessions
AI draftMinutes after uploadSame day
Fix speaker labels with the log0.5 to 1 hour3 to 6 hours
Accuracy check against audio2 to 3.5 hours12 to 21 hours
Format and anonymize0.5 to 1 hour3 to 6 hours
Total3 to 5.5 hours18 to 33 hours

To shorten it, check each session the day it runs, while voices are fresh, and give one person the speaker-label pass for every session. Hiring a focus group transcription service instead? Ask whether they charge per extra speaker, label by seat from your roster and how fast they turn around multi-speaker audio, then spot-check one session yourself.

Anonymizing a group transcript

Groups name each other. "Like Maria said" can identify Maria even after you've replaced her name in her own turns.

The UK Data Service's guidance on anonymisation for text data recommends replacing identifiers with bracketed pseudonyms or descriptions, keeping a log of every change separately, and not over-anonymizing so the data loses its meaning. For groups, add three checks:

  • Search for every first name from the roster, including nicknames, across all turns.
  • Watch employers and places that only one participant would mention.
  • Check combinations. A job, a town and a story can point to one person even with names gone.

Work on a copy, and keep the log with the original audio.

Focus group transcription with ParrotNotes

ParrotNotes records on your phone or Mac with no bot joining, so an in-person group in a community hall works as well as a call. Recording works offline; transcription runs once you're connected.

  • Long sessions. Pro records up to 3 hours per recording, with 3,000 minutes a month, enough for a full study.
  • Speaker identification on Pro. Turns are separated before your check. The introductions round and the speaker log make the labels easy to confirm, especially during crosstalk.
  • Search across sessions. Keyword search on every plan finds the quote you half remember. Pro adds AI semantic search.
  • 99+ languages. Pro transcribes and translates, useful for groups held in Spanish, German or a mix.
  • Export. Text or markdown on every plan, DOCX or PDF on Pro.

The Free plan includes 100 minutes a month with an AI summary on every recording. Pro is $19.99 a month, or $14.99 a month billed annually.

Kofi, a UX researcher, ran four groups in English and Portuguese. He recorded each on his phone, confirmed speakers against the note-taker's log the same afternoon, and sent clean verbatim transcripts two days after the last session. (Kofi is a composite.)

Get the labels right, and the rest follows

Good focus group transcription comes down to a few early decisions:

  • Pick the format from what you'll do with the transcript
  • Seat people by number, record an introductions round and keep a speaker log
  • Get consent from every participant, and say confidentiality can't be guaranteed in a group
  • Budget hours for checking, not typing
  • Anonymize names participants use for each other, not only their own

Planning your next session? Download ParrotNotes free and start each transcript from a searchable draft.

Frequently Asked Questions

How do you transcribe a focus group?

Record the session with consent from every participant, get a first draft from software or a transcriber, then fix speaker labels using a seating chart and a note-taker's log. Check every line against the audio, apply your format, and anonymize on a copy.

How do you label speakers in a focus group transcript?

Use MOD for the moderator and P1, P2 and so on by seat, with a roster that maps seats to participant codes. Start a new paragraph for every speaker change, mark uncertain turns as P? with a timestamp, and tag crosstalk.

How long does it take to transcribe a focus group?

The draft can be ready in minutes; checking takes the time. One 2024 study found checking one-to-one interviews took 1.5 to 3.5 hours each. Plan more for a group with many voices: roughly 3 to 5.5 hours per 90-minute session.

Should a focus group transcript be verbatim?

It depends on the analysis. Academic thematic or discourse work usually needs verbatim. Market research reports often use clean verbatim. Every format should keep speaker labels.

Do all focus group participants need to consent to recording?

Yes. Get written consent from each participant, confirm it on tape at the start, and tell them what is recorded, how it will be stored and that confidentiality can't be fully guaranteed in a group.