"Make the quiet parts louder" runs a compressor — so the loud parts get louder too
Crisp's Level audio does not find the quiet passages and lift only those. It runs a dynamic-range compressor with a 6× makeup gain and a lookahead limiter, and the makeup gain applies to the whole track. Measured on a test clip: quiet passages +15.98 dB, loud passages +10.85 dB, and the loud↔quiet gap closed only from 21.93 dB to 16.80 dB. Everything gets louder; the quiet parts just gain more.
The measurement
How the numbers were produced, so you can disagree with them: a 12-second test file — pink
noise band-limited to 180–3800 Hz, roughly the speech band, in alternating 1.5-second quiet and
loud blocks — through Crisp's audio lane in levelize mode, measured block by block
with ffmpeg's volumedetect. A synthetic signal rather than a recording of a room,
on purpose: it isolates the level behaviour. Figures are means across the four quiet and four
loud blocks.
| Quiet blocks | Loud blocks | Gap | |
|---|---|---|---|
| Before | −43.08 dB | −21.15 dB | 21.93 dB |
| After Level audio | −27.10 dB | −10.30 dB | 16.80 dB |
| Change | +15.98 dB | +10.85 dB | −5.13 dB |
The quiet talking did come up, by about 16 dB, which is why anyone presses this button. But the loud parts finished 10.85 dB louder than they started, and the distance between loud and quiet — the thing actually bothering you — closed by 5.13 dB out of 21.93. If the problem is a door slam blowing your ears off, this does not fix it. It makes the slam louder too, just more slowly.
Why the loud parts rise: the filter, term by term
The whole operation is one ffmpeg filter chain, and it is short enough to read
(backend/crisp_engine/audio.py):
acompressor=threshold=-22dB:ratio=4:attack=20:release=300:makeup=6,alimiter=limit=0.95
makeup=6is a linear multiplier, not decibels. 6× is +15.56 dB, and it is applied to every sample in the track. Checked by running the same compressor withratio=1so no compression happens at all: −43.0 dB became −27.4 dB. That +15.6 dB is the entire "lift".threshold=-22dB:ratio=4only touches audio above −22 dBFS. The quiet blocks sit at −43 dB, far below it, so they receive the full makeup gain and nothing else. The loud blocks cross the threshold and take 4:1 gain reduction, which claws back about 5 dB — which is precisely the gap that closed.alimiter=limit=0.95is a lookahead limiter on the end, so the makeup gain has somewhere to land instead of running off the top.
That is why "boost only the quiet bits" is not something a compressor can do. It narrows the distance by holding the loud down; the sensation of the quiet parts being lifted is the makeup gain afterwards, and makeup gain is not selective. Nothing in the chain knows which passage is dialogue and which is a slammed door.
The word you type picks a different filter
Crisp's plain-English box maps a request onto one of eleven audio modes. Two of them sound like the same English sentence and are not the same operation at all. Run against the shipped parser, not read off a list:
| What you type | What runs |
|---|---|
| "level the audio" | Normalize |
| "level the volume" | Normalize |
| "level the sound" | Normalize |
| "even out the volume" | Normalize |
| "even out the audio" | Normalize |
| "level the dialogue" | Level audio |
| "even out the dialogue" | Level audio |
| "even out the speech" | Level audio |
| "level the voices" | Level audio |
| "level the dynamics" | Level audio |
| "make the quiet parts louder" | Level audio |
| "even out the loud and quiet" | Level audio |
| "the quiet parts are too quiet" | Level audio |
The verb does not decide — the noun does. "Level" or "even out" plus audio / sound / volume / levels goes to Normalize. The same verb plus dialogue / speech / voices / dynamics goes to Level audio.
And Normalize is a
genuinely different operation, not a milder one: loudnorm=I=-16:TP=-1.5:LRA=11,
EBU R128 loudness normalisation to −16 LUFS. On the same test file it moved the quiet blocks
+5.65 dB and the loud blocks +5.32 dB and left the gap at 21.60 dB — where it started. It
decides where the clip sits; it does not change the distance between its loud and
quiet parts.
So: whole clip too quiet, or several clips that need to match each other → Normalize. Quiet talking and loud bangs inside one clip → Level audio. And if you phrase the second request with the word "audio" in it, you will get the first one.
What it costs
- The noise floor comes up with everything else. In the test, the quiet blocks were the noise, and they rose 15.98 dB. Room hum, air conditioning and tape hiss live in the quiet passages by definition, so they take the same +16 dB the dialogue does. Clean first, level second: Denoise for steady hiss and hum, Cut wind for buffeting on the mic.
- The output lands just under the ceiling. Measured on the levelled file, the loudest sample reads about −0.4 dBFS. That is a hair below full scale, so there is very little headroom left for anything you stack on top of it afterwards.
- You cannot dial it back. Those filter values are fixed; there is no strength slider for this mode. Heavy compression flattens a performance, so if a levelled clip sounds lifeless, the remedy is not to level more gently — it is to lift the whole clip instead.
What it will not do
- It cannot lift the voice and leave the music. Ask, and Crisp refuses in its own words rather than doing something else quietly: "Crisp works on the whole soundtrack — it can't split the music from the voices. It can mute it, even out the loudness, reduce background noise, or save the audio on its own." There is no stem separation in the build, so every audio mode acts on the mixed soundtrack.
- It works a clip at a time. Crisp changes a whole clip at a time, so to level one stretch of a recording you split the clip where the change should start (press S, or right-click → Split here) and apply it to that piece.
- It refuses silence instead of re-encoding it. A track that peaks below −80 dBFS gets "This clip's audio track is silent — there's nothing to even out." A file with no audio track at all gets "This clip has no audio track — there's nothing to adjust." Both checks run before any work starts.
Running it on a Mac
-
Drop the clip into Crisp
Get Crisp for Mac and drag your video onto the window. The whole job runs on your Mac — nothing is uploaded.
-
Select the clip, then Audio in the Tools panel → Level audio
Or type "make the quiet parts louder" in the prompt bar, which reaches the same mode. "Level the audio" does not — see the table above.
-
Press it, then check the loud parts, not the quiet ones
The video stream is copied with
-c:v copy, so the picture is byte-for-byte untouched; only the audio is re-encoded, to AAC at 256 kbps. Every audio track in the file is processed, not only the first — a screen recording with separate mic and system-audio tracks keeps both, and both get levelled. If the loud moments are now too loud, that is expected: they rose too. Lower the whole clip afterwards, or use Normalize instead if what you wanted was one consistent overall loudness.
Level a clip and measure it yourself
Crisp runs entirely on your Mac — no account, no upload, no queue. Free to try; free exports carry a "Made with Crisp" watermark.
Download Crisp for MacApple Silicon · macOS 26+ · Notarized
Questions people actually ask
Does making the quiet parts louder leave the loud parts alone?
No. The makeup gain applies to the whole track. Measured: quiet +15.98 dB, loud +10.85 dB, and the gap between them closed by 5.13 dB. That gap closing is the real effect, and it is smaller than most people expect.
What is the difference between Level audio and Normalize?
Two different filters. Level audio is a compressor plus limiter and narrows the loud↔quiet distance inside a clip. Normalize is EBU R128 loudness normalisation to −16 LUFS and sets where the clip sits, leaving that distance roughly as it was (measured: gap 21.93 → 21.60 dB).
Why does "level the audio" not do the same thing as "level the dialogue"?
Because the noun routes the request, not the verb. "Audio", "sound", "volume" and "levels" go to Normalize; "dialogue", "speech", "voices" and "dynamics" go to Level audio.
Does levelling the audio re-encode the video?
No — the video stream is copied, so the picture is untouched and no quality is lost on it. Only the audio is re-encoded, to AAC at 256 kbps.
Can it raise the dialogue and leave the music where it is?
No. Crisp works on the whole soundtrack and says so when asked; there is no stem separation in the build.
Related
- Normalize the loudness of a video on a Mac — the other half of this page
- Remove background noise — do this before you level, not after
- Extract the audio as an MP3
- Mute a video
- Levelling a lecture recording
- Levelling a service recording