two audio waveforms with a mic and headphones and an audio interface

Audiobook Audio (V3) Part 6: Leave Headroom

Why headroom is a performance tool

This part is about gain staging, which sounds like the dullest part of audiobook recording until the exact moment it matters: a line that suddenly gets louder than the rest.

In a chapter I recorded recently, there was a scene where a policeman shouts at a crowd. Not a full-on theatrical yell, but definitely a shout; the kind where you want the intensity to jump without the sound turning into “new room, new mic, new narrator”.

This time I did something I used to avoid. I kept my distance from the mic exactly the same – the distance where I had already decided my tone lives – and I just raised my voice.

My old instinct would have been to lean back slightly, just to be safe. But I did not want the shout to sound as if the policeman had suddenly stepped into a different room. And, miracle of miracles, it did not clip.

Later, in editing, I brought the shout down to a sensible loudness with clip gain, and to my delight and faint surprise, it sounded right. It still sounded like a shout, just not offensively loud.

The important thing was that the tone matched the surrounding narration, because I had not changed the acoustic setup mid-line. That is the core idea here: headroom lets you keep tone decisions stable while loudness moves around.

So, why does headroom matter beyond just not clipping?

Everyone knows not to clip. That part is not exactly a revelation. The more interesting problem is what happens before clipping, when you are close enough to the ceiling that you start narrating defensively.

When the meter becomes the thing you are managing, performance changes:

  • emphasis gets smaller;
  • consonants get careful;
  • breath gets less natural;
  • intensity gets contained to stay safe.

Over a full audiobook, those tiny compromises become a pattern. Later, in editing, that pattern is obvious, not because the audio is broken, but because the performance feels managed.

Headroom is what stops that. It buys freedom, but in the least dramatic possible way: by giving the loud bits somewhere to go.

The nerdy bit is crest factor: why speech peaks jump around

Speech has a high crest factor. That means the peaks can be much higher than the average level. A paragraph can sit at a steady average level, then a single consonant, plosive, or emphatic word spikes far higher than expected. That is normal. It is how speech works.

If the recording is already hot, those peaks force behaviour: backing off, pulling emphasis, or turning the shout into a polite version of itself. That is how a line that should feel big ends up feeling strangely small.

Modern digital capture makes recording hot unnecessary. With 24-bit recording, there is enough dynamic range that the practical noise floor is the room long before it is the format. So the trade-off has shifted. There is very little reward for pushing the level, and a lot of risk.

So this brings me to practical targets: a repeatable window

The numbers matter less than repeatability, but I do like having a target window. Otherwise I will absolutely start negotiating with myself. For narration, my useful boring window is:

  • normal narration peaks around -12 dBFS;
  • bigger moments rising naturally up to around -6 dBFS;
  • no clipping.

That gives me enough room for real performance, including the shout, without making ordinary narration pointlessly tiny.

I set that with the input gain on the interface or microphone. The DAW fader in Logic is not usually what prevents clipping; it changes what I hear or how things sit later. If the signal is already clipped on the way in, the DAW cannot really rescue it.

The important distinction is that monitoring level is separate from record level

A lot of recording too hot comes from trying to hear yourself better. Record level and monitoring level are different systems: record level determines what gets written into the file, and monitor level determines what I hear in my headphones.

So the rule is simple:

  • set record gain once to hit the target window;
  • then set headphone or monitor level for comfort.

If it needs to be louder in my ears, I change monitoring, not input gain. That one distinction prevents a surprising amount of nonsense.

Making a demo: recreating the shout problem on purpose

This is an easy thing to prove, and it is worth doing once because the lesson sticks much better when you hear yourself flinch.

First is take A: with headroom. Set gain so normal peaks sit around -12 dBFS, and just shout naturally, without moving away. I do this in the podcast version, and we get a nice, above-speaking loudness that just fits nicely.

Then Take B: with it too hot. Set gain so normal narration peaks around -3 dBFS, and just shout naturally, without moving away. This one was actually not terribly bad in the podcast, but the voice was a bit harsher.

After recording, I then needed to level-match both takes in the edit. This matters because headroom is not about leaving the recording quiet forever. It is about capturing safely and shaping loudness later, and in the too-hot take, one of two things usually happens:

  • it clips;
  • or you self-censor, and the shout becomes careful.

In mine, I let it clip.

iZotope RX11 screenshot of the Take A shout recorded with headroom
Take A in iZotope RX11, after being reduced to a comfortable listening loudness. This is the shout recorded with headroom, so the level can be shaped afterwards without having baked clipping into the file.
iZotope RX11 screenshot of the Take B shout recorded with the gain too high, with clipped peaks highlighted
Take B in iZotope RX11, also reduced to a comfortable listening loudness, with the clipped peaks highlighted. The loudness can come down afterwards, but the clipping from recording too hot is already part of the audio.

When the takes are matched, the difference is usually clear. The headroom take stays stable in tone, even through the intensity jump. The hot take tends to sound constrained, and it tolerates processing less gracefully because the peaks are already crowded.

But why not just back off?

It’s easy to give it a try, and you might be surprised. In my recording, I left it set at that -3 dBFS position, but backed away from the microphone to prevent clipping. I was still shouting naturally, but further away.

What we found was that backing away from the mic does not just change loudness. It changes the sound.

Usually you get some mix of:

  • reduced proximity effect, so the voice can get thinner;
  • more room reflection;
  • a different consonant balance;
  • and sometimes a different noise relationship.

So the shout can end up sounding as if it happened in a different place, even if the level is technically controlled.

For now, the simple constraint is: keep distance constant to keep tone constant. Let the performance create intensity. Let headroom absorb it. Then set loudness later with clip gain.

Distance can be used as an effect, but only once it is a deliberate choice rather than a safety reflex.

So, here is the calibration habit

The habit that makes this boring, in the good way, is a calibration read. Before recording properly, I use a consistent paragraph and include my loudest plausible delivery. Not my loudest imaginable delivery – I am not testing for opera or disaster – just the loudest thing this book is likely to ask of me. Then I set gain so normal narration peaks around -12 dBFS, and the loud line still stays safely below clipping.

The point is not to worship those numbers. The point is to stop thinking about them once the chapter starts. Then I leave input gain alone. The card is:

  • normal peaks: about -12 dBFS;
  • big peaks: up to about -6 dBFS;
  • input gain fixed;
  • monitoring adjusted separately.

So, where does that leave us?

Here is the rule: if my performance is being driven by the meter, I have already lost. Set gain once, leave headroom, and set monitoring separately so I am not performing into the red.

Next time, I will tackle the thing that can ruin perfect gain staging anyway: the noise you notice. Trains, birds, air-con, and the myth of silence.

Leave a Reply

Your email address will not be published. Required fields are marked *

Scroll to top