“The audio is bad” is almost never the real problem statement.
What people mean is: “I cannot understand the words.”
In training rooms and hybrid classrooms, speech intelligibility is the make-or-break metric. Video quality matters, but when learners miss a sentence, the entire session gets harder to follow. Remote attendees are impacted even more because they only hear what the room microphones deliver.
The good news is that intelligibility is not mysterious. It is measurable, and it improves dramatically when you fix the right things in the right order.
Loud is not the same as clear
A room can be loud and still be unintelligible.
Speech clarity drops when any of these three conditions get worse:
-
Background noise is too high
-
Reverberation is too long
-
The capture chain is wrong (microphones, loudspeakers, DSP, or configuration)
Classroom acoustics standards make this very explicit. ANSI/ASA S12.60 sets limits for both background noise and reverberation time in unoccupied, furnished learning spaces because those two factors are core to speech communication.
The four numbers that explain most “they can’t hear us” complaints
You do not need an acoustics lab to get practical value. But you do need a few measurable targets.
1) Background noise level (dBA)
For core learning spaces, ANSI/ASA S12.60 limits the greatest one-hour average A-weighted background noise level to 35 dBA for both interior and exterior sources (evaluated independently).
2) Reverberation time (T60)
ANSI/ASA S12.60 sets maximum reverberation times in core learning spaces as:
-
0.6 seconds for volumes up to 10,000 ft³
-
0.7 seconds for volumes from 10,000 to 20,000 ft³
3) Signal-to-noise ratio (SNR)
For meeting and conferencing spaces, Shure notes a common industry guideline: an absolute minimum SNR is 20 dB in a conference room.
4) Speech Transmission Index (STI) or STIPA
If you want one score that tracks intelligibility, STI is the standardized index.
IEC 60268-16 describes the Speech Transmission Index (STI) as a fast, objective method that expresses speech transmission quality as a value between 0 and 1.
NTi Audio explains that STIPA uses a known test signal and reports speech intelligibility as a single number from 0 (unintelligible) to 1 (excellent intelligibility) under the IEC 60268-16 method.
Quick reference targets
(Use these as starting points when designing, upgrading, or troubleshooting.)
| Metric | Typical target |
|---|---|
| Classroom background noise | 35 dBA |
| Classroom T60 | 0.6 to 0.7 s |
| Conference room minimum SNR | 20 dB |
| STI / STIPA scale | 0 to 1 |
Why hybrid makes intelligibility harder
Hybrid and distance learning introduce a few challenges that in-person-only rooms can sometimes “get away with”:
-
Microphones are farther from the talker than a headset or lapel mic would be.
-
The room’s loudspeaker audio can leak back into the microphones.
-
Echo and noise are far more noticeable once audio is digitized, compressed, and delayed.
This is exactly where conferencing DSP matters.
Acoustic Echo Cancellation is not optional in real hybrid rooms
Bose explains AEC in plain terms: Acoustic Echo Cancellation prevents far-end participants from hearing their own voices echoed back to them. The issue becomes obvious in full-duplex calls when the far-end audio plays through near-end loudspeakers and gets picked up by the near-end microphones, then sent back to the far end with enough delay to be disruptive.
If AEC is missing, misconfigured, or double-applied, the far end often says “I hear myself” or “it sounds like you’re in a tunnel,” even when the in-room audience thinks it is fine.
The most common root causes in training rooms and hybrid classrooms
1) Too much reverberation
Hard surfaces like glass, concrete, whiteboards, and untreated drywall reflect speech. The reflections overlap with direct speech and smear consonants. The room sounds “live,” but words become harder to distinguish.
A fast clue: the clap test. If you hear a long ring or flutter echo, intelligibility is at risk.
2) HVAC noise and mechanical rumble
If the room is already loud before anyone speaks, every microphone system is forced to amplify noise too.
Remember: ANSI/ASA S12.60 sets 35 dBA as the target background noise in core learning spaces, and that is an unoccupied baseline. If your empty room is already above that, you are starting behind.
3) The wrong microphone strategy
This usually shows up in one of two ways:
-
The instructor is too far from the mic (ceiling mic only, no wearable, no boundary mic, no proper gain plan).
-
The room relies on “the laptop mic” or a single tabletop device in a big space.
In training rooms, the “primary talker” is often one person. You usually get the biggest clarity boost by capturing that voice close to the source.
4) Untuned DSP and incorrect gain structure
Even great hardware sounds mediocre if:
-
the gain is too low (noise dominates)
-
the gain is too high (clipping or distortion)
-
EQ is not shaped for speech
-
AEC is not properly referenced
5) Loudspeaker placement that fights the microphones
If your loudspeakers blast into the mic pickup area, the AEC has to work harder, and your intelligibility drops under real conditions.
A practical troubleshooting sequence that works
If you try to fix audio by swapping equipment first, you can spend a lot and still fail.
This order works because each step reduces variables for the next one.
Step 1: Measure the room before touching the system
Do these in a quiet room:
-
Background noise (dBA)
-
A basic reverberation estimate (or RT60 measurement if you have tools)
If you want classroom-aligned targets, use ANSI/ASA S12.60 as your reference point for core learning spaces: 35 dBA and 0.6 to 0.7 seconds depending on volume.
Step 2: Fix acoustics and noise first
The best microphone cannot fully overcome a bad room.
Common fixes that move the needle:
-
ceiling absorption (acoustic tile or cloud panels)
-
wall absorption at first reflection points
-
soft treatments on large reflective surfaces where feasible
-
HVAC adjustments to reduce diffuser noise and low-frequency rumble
Step 3: Choose the right capture model for the use case
Pick based on how people actually talk in the room.
Typical patterns
-
Instructor-led training: wearable mic for instructor plus room mics for Q&A
-
Discussion-based classroom: distributed mics (ceiling array, table mics, or a hybrid)
-
Hybrid collaboration room: balanced pickup for multiple talkers plus robust AEC
Step 4: Confirm AEC is present and correctly referenced
AEC needs the correct far-end reference feed and a stable loudspeaker-to-mic relationship.
Bose’s explanation highlights why this matters: full-duplex conferencing requires AEC to prevent the far end from hearing their own audio echoed back due to round-trip latency.
Step 5: Commission with intelligibility metrics, not “it sounds fine”
If you want objective verification, use STI or STIPA.
IEC 60268-16 defines STI and positions it as an objective rating of speech intelligibility, with results expressed from 0 to 1.
NTi Audio describes STIPA measurement as reproducing a known test signal and measuring intelligibility at listening positions, again producing a 0 to 1 score.
NTi Audio also notes a practical detail that matters when simulating a talker: IEC 60268-16 references a typical talker level of 60 dBA at 1 meter for a human voice level when reproducing the test signal in rooms without installed audio systems.
What “good” looks like operationally in 2026
A training room or hybrid classroom is truly ready when:
-
Remote participants can understand words without asking for repeats.
-
In-room listeners can understand speech from anywhere you expect instruction.
-
The system remains intelligible when the room is occupied and noisy.
-
AEC works reliably during full-duplex conversation.
-
Support can troubleshoot using known metrics and a consistent room standard.
Where VIcom makes it simple.
We understand that speech intelligibility is not a single product. It is the result of design, integration, and commissioning across the room and the UC stack.
VIcom helps organizations get from “it’s installed” to “it’s intelligible” by:
-
assessing room acoustics and noise issues against measurable targets
-
designing microphone and loudspeaker strategies that match real use cases
-
integrating conferencing DSP with correctly configured AEC
-
commissioning and validating performance with objective intelligibility methods (STI/STIPA where appropriate)
-
creating repeatable standards so every room behaves predictably
If your organization is planning 2026 upgrades, this is one of the highest-return improvements you can make because it directly affects comprehension, meeting outcomes, and training effectiveness. We’d love to speak with you, schedule a free consultation today by filling out the form below!
