Crisp › How-to › How Auto-Montage chooses

Auto-Montage scores every second — and the scoring changes with the footage

Quick answer

Crisp scores your footage one second at a time. Each second gets ten measurements — motion, loudness, voice-band ratio, audio onsets, edge sharpness, scene cuts, luminance, and face, main-subject and on-screen-text readings from macOS Vision — combined by one of five scoring formulas. Crisp picks the formula from the footage itself, then takes the highest-scoring seconds, pushing the picks apart so one hot stretch can't fill the reel.

Crisp · built in Grand Rapids · published 2026-07-07 · last updated 2026-09-15

The unit of decision is one second

Crisp walks the whole clip first and builds a record for every one-second window. Four decode passes fill in most of it: scene-cut detection, a tiny-frame motion/sharpness/luminance pass, a loudness envelope of the audio, and a second envelope band-limited to the voice range. The arithmetic that turns those into a decision is ffmpeg plus plain Python — the montage engine ships without numpy, scipy or torch. The only learned components are Apple's own Vision detectors, and they supply three of the ten signals rather than making the call.

SignalWhat it actually measures
motionFrame-to-frame difference on downscaled frames.
audioRMS loudness envelope of the mix.
transAudio onset — the positive jump in loudness. A bat crack, a golf strike.
speechShare of a second's audio energy in the ~200–3500 Hz voice band against the full mix. A likelihood someone is talking — not recognition of any word.
scene1 if an ffmpeg scene cut falls inside the window, 0 otherwise.
sharpEdge energy. An out-of-focus second scores low.
lumaMean luminance, absolute.
faceOn-screen face area, from the bundled macOS Vision helper.
subjectAttention saliency — how prominent the main subject is.
textFraction of frame covered by recognised text (Vision OCR). This is what separates a title card from footage.

Six of them are normalised inside their own clip, so the best second of a dull clip reads 1.0 exactly like the best second of an exciting one. Text and luminance stay absolute — a title card is a title card and black is black regardless of what else is in the file. The Vision helper is bundled and the build refuses to produce a Crisp without it; if it is unavailable anyway those three signals return empty and selection falls back to motion and audio. Longer clips are sampled more coarsely so analysis stays bounded: past about 15 minutes the Vision pass drops below one frame per second, and the motion pass falls from 2 fps toward 0.4.

Five formulas, and Crisp picks one from your footage

Crisp does not apply one notion of "interesting". It classifies the footage from global statistics of the signals it just measured, then chooses the scorer to match. The tests run in order and the first match wins:

Note how small the speech weight is outside the talking profile — 0.04 balanced, 0.02 in the sports profiles. That is deliberate: an earlier build weighted talking heavily enough that a reel filled up with whoever was speaking instead of whoever was doing something. The subject-action term carrying most of the balanced score is a product, not a sum — a salient subject multiplied by how much it is moving — so a static talking head and empty scenery both score low. Without Vision it degrades to plain motion, never worse than before.

Two disqualifiers every profile obeys

Every scorer routes through one shared shot-quality step, so no profile can forget the rule: it nudges the score by sharpness, and demotes recognised text hard — a penalty of 3.0, which takes a full title card down by 27–84% while a jersey number or scoreboard loses almost nothing. That exists because of a real failure. On two public-domain films, a dissolve between two title cards spiked motion sevenfold while the lettering read as crisp and the narration read as speech, and the window scored third highest in the entire film. Nine sampled frames of real footage measured exactly 0.0 for text; six title cards measured 0.09–0.28. The threshold sits at 0.02.

Demoting is not enough alone, because a chunk is scored by its strongest window — a title card at the edge of an otherwise good five seconds is invisible to the score. So a second pass trims caption and near-black windows off the head and tail of each picked chunk, whole windows only, never below the minimum clip length. The near-black threshold is a mean luminance of 0.10, measured against those same films rather than guessed.

Highest score does not simply win

If Crisp just took the top-scoring seconds, a reel of a football match would be three minutes of one goal. Selection tiles the source into candidate chunks, scores each by its best window, then picks greedily with a diversity penalty: a candidate near an already-picked chunk is discounted, with the spacing derived from how many picks your requested length implies. The genuinely best stretch still wins; one merely near something already chosen has to be clearly better. The tiling spans the whole source rather than a fixed budget of tiles, so on a long recording Crisp strides further between candidates instead of only considering the opening. The picks are then sorted back into chronological order — with several clips, in the order you dropped them in. Crisp never reorders your footage to build a better arc.

The other two modes score differently on purpose

Smart Condense skips the profile detector entirely — it ranks every one-second window with the balanced scorer, keeps the best until it reaches your target, then merges adjacent kept seconds back into continuous runs. It preserves chronology and trims dead air. Be precise about what that means: there is no silence detector anywhere in this. "Trim the dead air", "cut the boring parts" and "condense this" are recognised phrases that route to the montage lane — and what they get is interest scoring, not a search for quiet.

Music throws most of the scoring away and scores on motion alone, minus the same two disqualifiers. A music reel wants movement rather than the subject and speech weighting; a beat-length chunk of title card is still a beat-length chunk of nothing. Darkness is deliberately not a term here — demoting dark windows did not change how many segments were picked, it changed which, concentrating them at the bright end of the source.

Where the cuts land in a music montage

Beats are onset peaks in a loudness envelope built at a 10 ms hop, with a minimum 0.28 s between accepted beats (about 210 BPM). The envelope is band-limited to 30–150 Hz, and that band is the most consequential number on this page. A kick lives there; a melody at 220–330 Hz and a sustained pad do not. Measured at 128 BPM, 128 true beats over 60 seconds, same peak-picking either way:

Test signalBroadband envelope
(the earlier detector)
30–150 Hz
(what ships)
bare click127 found, 5.0 ms128 found, 5.0 ms
+ sustained bass127 found, 3.8 ms128 found, 3.7 ms
+ offbeat hats127 found, 3.8 ms128 found, 3.7 ms
+ pad and melody174 found, 130.0 ms128 found, 3.7 ms

Offsets are the median distance from the true beat. A melody note starting mid-beat raises a broadband envelope exactly like a kick does — which is why a percussion-only test bed can make a broadband detector look perfect.

Landing the cut is the other half. Each segment is trimmed and resampled to a constant frame rate independently, so its length is rounded to a whole frame on its own — and at 128 BPM a beat is 14.0625 frames at 30 fps, so the same fraction is dropped at every cut, always in the same direction. Crisp rounds the cumulative time to frames instead: each cut lands on the nearest frame to where the beat actually falls, the error is bounded by half a frame (16.7 ms at 30 fps, the floor for any frame-based edit), and it never accumulates. A music montage also runs the length of the track rather than a length you set, and gets 768 segments where Highlights and Smart Condense get 64 — a 3-minute track at 128 BPM asks for 384 cuts, one at 174 BPM asks for 522.

What Auto-Montage does not read

It has no speech recognition — no transcript, no subtitles, no searching a recording by what was said. The voice-band ratio is a loudness measurement, not a word. There is no face restoration or beautification anywhere in Crisp. And the montage chooses whole seconds of a clip; a request to apply an edit to only part of a clip is refused, with instructions to split it first.

Running it

  1. Drop your clips in

    Get the free Crisp app for Mac and drag in one long clip or several. Analysis and edit both run on your Mac; nothing is uploaded.

  2. Press Make a montage from these clips

    It is in the Project panel — click Back to project first if a clip is selected. Or type it: "make a 30 second highlight reel", "condense this", "cut it to the beat of this song". Crisp maps the phrasing to the same tool and sets the mode and length.

  3. Read the result line

    It names the mode, the in and out durations, and — for Highlights — which content strategy fired. That is the fastest way to find out why a reel picked what it picked.

Crisp is a one-time $129 · or 4 × $32.25, not a subscription. On an unpaid machine a montage export carries a Crisp watermark.

See what it keeps from your own footage

Free to try on your Mac. Drop in a long clip, pick a length, and check the result line to see which strategy your footage triggered.

Download Crisp for Mac

Apple Silicon · macOS 26+ · Notarized

Questions people actually ask

Does Crisp listen to what people are saying to pick the highlights?

No. Crisp has no speech recognition of any kind. What the montage measures is a voice-band ratio — how much of a second's audio energy sits in the roughly 200–3500 Hz band against the full mix. That tells it someone is probably talking; it never tells it what was said, and Crisp cannot transcribe, subtitle or search a recording by its words.

Why did Crisp pick completely different moments from two similar clips?

Because it may have classified them as different kinds of footage. The global statistics decide which of the five scorers runs, and the tests are ordered, so a clip sitting near a threshold can land either side of it. Game capture is scored by audio onsets and in-frame change; field sport by camera motion plus sustained loudness. Those pick genuinely different seconds out of the same minute.

Can I see which scoring profile it used?

Yes, for Highlights. The result line names the strategy the content triggered — sports, golf/tennis-style, game capture or talking. The balanced default prints no style name, so a blank there means none of the four matched.

Why does a music montage cut far more often than a highlight reel?

Different segment ceilings: 64 for Highlights and Smart Condense, 768 for Music. Highlights is bounded by how much material is worth keeping; a beat-matched reel is bounded by the song.

How close to the beat do the cuts actually land?

Within half a frame — 16.7 ms at 30 fps, the floor for any frame-based edit — because the running total of the reel is quantised to frames rather than each segment separately.

Can I see which scoring profile Crisp used?

Yes, for a Highlights montage. The result line names the strategy the content triggered — sports, golf/tennis-style, game capture or talking — beside the mode and the in and out durations. The balanced default prints no style name.

Related