How long should a short-form clip be? Why we settled on 20–45 seconds

Our first clips came out at 6 and 12 seconds. What the platforms reward, what the data on our own output showed, and why we settled on 20–45 seconds.

4 min read

The first real complaint we got about clip quality wasn't about the moments the tool chose. It was about how short they were: a 7-minute comedy upload produced clips of 6 seconds and 12 seconds. The model had found the funny line, cut tightly around it, and thrown away the setup that made it funny.

Those clips were technically accurate and completely unwatchable. This post is about what we changed and why the numbers are what they are.

What the platforms actually reward

There's a lot of folklore here, so it's worth separating what's known from what's repeated.

Completion rate matters more than length. Every short-form algorithm weights the fraction of the clip people watched, and whether they watched it more than once. A 12-second clip that people finish is not automatically better than a 40-second clip that people finish, because the second one also delivered 40 seconds of watch time. Extremely short clips have a ceiling on both.

The first two seconds are a separate contest. The scroll decision happens before the clip has really started. That's about the hook, not the length, and it's its own subject.

60 seconds is a real boundary. YouTube Shorts allows up to three minutes now, and TikTok goes much longer, but the Shorts shelf and the "short" classification on every platform have their roots at 60 seconds, and the audience's expectation of a clip (as opposed to a video) still sits there. Past a minute, the thing is no longer a moment; it's an excerpt, and it gets judged as one.

Under 15 seconds, there's no room for an arc. A hook, a build and a payoff take time to say out loud. Fifteen seconds is about 40 words. Most complete thoughts are longer than that.

So the useful range is roughly 15 to 60 seconds, and within that the question is what the content needs, not what the platform prefers.

What the content needs

We classify the video before picking anything, because the unit of a clip is different per type:

Content The unit of a clip Typical length
Comedy, pranks, reaction setup → escalation → punchline → the reaction 20–35s
Podcast, interview question or claim → development → landing 30–50s
Educational, talks hook question or claim → explanation → takeaway 30–55s

The comedy row is where the 6-second clips came from. A model asked to "find the funniest moment" finds the punchline, because that's the line with the laugh under it. But the punchline is the end of the unit. Cut there and you've clipped a joke told to silence.

The podcast row runs longer because the good moment is usually a story or an opinion with a reason attached, and the reason is the part worth posting.

What the data on our own output showed

We logged where the model's picks landed against the actual sentence boundaries in the transcript. One example from a real job, on a comedy upload:

  • the joke's setup started at 49.0s
  • the model clipped 56–62s

Six seconds, starting seven seconds after the bit began. The model had anchored on the punchy line and amputated everything before it. That pattern, punchline-first and setup-lost, showed up in most of the short picks, not just this one.

The instinct is to fix this by asking the model for longer clips. That half works. The thing that actually fixed it was changing what happens to a short pick after the model returns it.

What we changed

1. Never trim, always widen. Our earlier hard cap was 35 seconds, with automatic trimming to fit. That was the tightest cap in the market and it reproduced the market's most common complaint, mid-sentence cuts, by design. The cap is now 60 seconds, and nothing is ever trimmed to fit it. A pick that runs long is a pick that runs long.

2. Short picks get widened, setup first. Anything under 18 seconds is expanded towards about 26 by pulling in whole sentences from the transcript, and the expansion goes backwards first. The model reliably finds the payoff and reliably loses the setup, so the setup is what's missing. Widening forwards would add the next joke's opening instead.

3. The target range is stated, not implied. The selection is asked for 20–45 seconds, with 15 as a floor and 60 as the ceiling, and told explicitly that the setup of a good bit is the hook and must not be sacrificed to start on the punchiest line.

4. The reaction is part of the clip. For comedy and reaction content the instruction now says the laugh, scream or silence after the line is the payoff, and the clip may not end before it. We also feed the model audio-energy markers from the waveform, so it can see where the loud reaction is even though the transcript can't.

What to do with this if you clip by hand

Whatever tool you use, or if you use none:

  • Start one sentence earlier than feels necessary. The line you think is the start is usually the line that made you notice. The real start is before it.
  • End after the reaction, not on the line. Hold for the laugh, the pause, the "wait, what?". It's rarely more than two seconds and it's the difference between a moment and a quote.
  • Distrust anything under 15 seconds. If it genuinely works that short, fine. Most of the time it's a fragment.
  • Don't trim to a number. If the moment is 52 seconds, post 52 seconds. Cutting it to 45 to hit a target loses the part that made it worth posting.

The tool's job is to get this right by default. Ours didn't at first, and we measured how badly before we fixed it. The 20–45 range isn't magic; it's just where complete thoughts tend to fit.

Try it on your own video

3 videos a month free, no card. Paste a link and see what comes back.