Cutting an 11-second clip from a two-hour YouTube video with 5.9 MB instead of 370

Sep 3, 2026

VideoToShorts takes a YouTube link, shows you the transcript, and turns the lines you pick into a vertical captioned clip. To render the clip we need the source video for those seconds, and only those seconds. For a while we were paying for the whole video anyway.

Two days of real traffic, measured on 1 September: 35 source fetches, 3.1 GB moved. The average clip cost 90 MB of source. The worst one was an eleven-second cut from a two-hour film that cost 369.7 MB, and the next clip from the same film cost it again. The bytes are metered, so this was a real bill for a product that charges about a quarter per clip.

This post is about how that eleven-second clip now costs 5.9 MB, using nothing more exotic than HTTP Range requests and the MP4 container format.

Why we were taking whole files

The normal answer is yt-dlp --download-sections, which asks ffmpeg to seek into the DASH streams and pull only the window. That works when ffmpeg can reach YouTube directly. It does not work on the route our worker has to use, where ffmpeg cannot fetch at all and we pull the streams ourselves with curl.

Pulling DASH streams ourselves ran into a wall we spent four rounds of probing on. A Range request into any adaptive stream is served for roughly the first minute of the video and refused with a 403 past it. The limit is in seconds, not bytes. Here is the measurement, one session, four audio formats:

formatbitrateservedrefused
itag 13949 kbps100 KB (16 s)1 MB (164 s)
itag 140129 kbps1 MB (62 s)3 MB (185 s)
itag 251129 kbps1 MB (62 s)3 MB (186 s)
itag 258388 kbps3 MB (62 s)8 MB (165 s)

Chunking the request does not lift it. Two 30-second chunks land and the third is refused, so the ceiling is on how far into the stream we may reach, not on how much we ask for at once. A PO Token on the URL changes nothing. The range= query parameter a real player uses behaves the same way.

One format is exempt: itag 18, the progressive muxed file, h264 plus aac at 640x360. The android client hands it out without a token, and it downloads through the same route from start to finish. So for any clip that starts past the first minute, itag 18 was the only source we could get, and we took it whole. The note in the code called that acceptable, measured against a ten-minute source at 28.5 MB. At two hours it was neither acceptable nor necessary.

Two facts that make the whole file unnecessary

itag 18 answers Range requests at any depth. We checked at 72 MB into a file and at 90 percent of a 377 MB file. Both came back 206. The reach wall applies to the adaptive streams and never applied to this one.

A progressive MP4 knows where every second lives before you have any of the payload. The container keeps its index in the moov box: per-sample sizes in stsz, chunk offsets in stco or co64, and the sample-to-chunk map in stsc. YouTube writes itag 18 with moov at the front of the file, the layout usually called faststart. So the first few megabytes of the file are enough to answer "which bytes hold seconds 2355 to 2366", exactly, with no guessing.

Put those together and the plan is three requests. Fetch the header. Read the byte span for the window out of the index. Fetch the span. Everything in between stays a hole.

Fetching the header

We resolve the media URL with yt-dlp, keeping the file size it reports, then ask for the first 64 KB. That is enough to hold the box headers of anything we have seen. Walking them is a dozen lines: an MP4 box is a 4-byte big-endian length followed by a 4-byte type, and a length of 1 means a 64-bit length follows.

def _moov_end(head: bytes) -> int | None:
    at = 0
    while at + 8 <= len(head):
        size = int.from_bytes(head[at : at + 4], "big")
        kind = head[at + 4 : at + 8]
        if size == 1:
            if at + 16 > len(head):
                return None
            size = int.from_bytes(head[at + 8 : at + 16], "big")
        if size < 8:
            return None
        if kind == b"moov":
            return at + size
        if kind == b"mdat":
            return None
        at += size
    return None

Reaching mdat before moov means the index sits at the end of the file. Everything below depends on reading it from the front, so that case returns None and the clip is refused with that reason. If the header is bigger than our first bite, a second Range request fetches exactly the rest of it. A two-hour progressive file carries a few megabytes of sample tables; we give up past 32 MB, because a file shaped like that is not one we understand.

The bytes land in a sparse file. We truncate a temp file to the full reported size, then write each fetched range at its true offset. Nothing else gets written. The operating system stores the holes as nothing, and every offset in the index still points at the right place.

Reading the span out of the index

We do not parse stco and stsz ourselves. ffprobe already does, and it will tell you where each packet in a time interval sits:

ffprobe -v error \
  -read_intervals "${START}%+${LENGTH}" \
  -show_entries packet=pos,size \
  -of json ranged.mp4

Run against the sparse file, ffprobe walks the index and lists every packet in the interval with its byte offset and length. The payload behind those offsets is still a hole, which reads as zeros for free, and we throw the packets away. All we keep is the smallest pos and the largest pos + size.

Two details in that command matter. The interval is padded by three seconds on each side, so the span includes the keyframe before the cut. And it lists both streams, because audio and video are interleaved in a progressive file, and a span that covers only the video packets renders as a clip with no sound.

We then add 1.5 MB of margin at each end. The index is exact, but a cut lands in the middle of a group of pictures more often than not, and the margin is cheap next to a decode that comes up short.

Why not just estimate

The first attempt guessed. Eleven seconds of a two-hour file is 0.15 percent of it, so take the file size, multiply, and fetch a window around that offset.

On the film that started all this, the eleven seconds the visitor asked for occupied 0.69 MB, starting 82.1 MB into the file. The proportional guess pointed at 72.4 MB. ffmpeg decoded nothing there. The bitrate of a real video is nowhere near even, and a guess that is off by ten megabytes is worse than no guess, because you have paid for bytes you cannot use. The span is always read from the index and never estimated.

Fetching the span and rendering

The span goes out as Range requests in 8 MB chunks, each written at its true offset in the same sparse file. Chunking costs nothing and means a refusal stops at a chunk boundary instead of after one enormous request.

curl -sL --max-time 300 \
  -H "Range: bytes=80600000-84300000" \
  -o span.part -w "%{http_code}" "$MEDIA_URL"

Then ffmpeg gets the sparse file with -ss at the clip start. It reads the index from the header, seeks to offsets we fetched, and never touches the holes. From its point of view this is an ordinary file with an ordinary seek.

Measured end to end on that same eleven-second clip: 5.9 MB pulled against 369.7 MB, a 98.4 percent saving, and the render succeeded. The worker now logs bytes pulled, file size and the fraction on every fetch, and the cost record on each clip names the method, because "the whole window was fetched the cheap way" and "the whole file was fetched" had been the same line item.

The bug that cost 181 MB anyway

The first shipped version kept the whole-file download as a fallback behind the ranged path, on the theory that cheaper is never worth a clip that does not exist. The day after it shipped, a 35-second cut from a 49-minute video cost 181 MB.

Walking the steps by hand found the file was faststart and the window was 6 MB. What had happened: the ranged path tried one outbound address, failed on it, returned None without saying why, and the whole-file path behind it succeeded on its own retries. YouTube binds a media URL to the address that resolved it, so a refused range cannot simply be retried elsewhere. It needs a fresh resolve on the new address, and the ranged path did not do that.

Two changes. The ranged fetch now retries per address, resolving again each time, and it distinguishes a property of the file from a property of the route. An index at the end of the file, or a window the index places no packets in, is the same on every address, so it stops and reports. A 403 on the bytes is the route, so it moves on.

And the whole-file fallback is gone. A window the ranged fetch cannot take is a refused clip with the reason attached. A silent fallback had turned a bug into a bill, and it had hidden the bug for a day, because from the outside a 181 MB success looks exactly like a success.

What this does not do

It only works on a progressive file with the index at the front. Adaptive streams are cut differently, and we can only reach the first minute of those on this route.

itag 18 is 360p. A 640x360 source under a 720x1280 canvas is soft but watchable, so clips past the first minute of a YouTube video render at 720p. We refuse 1080p from that source rather than sell an upscale as native quality.

And YouTube can change any of it. The reach wall, the exempt format, the faststart layout of itag 18: every one of these is a measurement, not a guarantee, and the worker treats a change as a refusal with a reason rather than something to route around silently.

None of it is specific to YouTube, though. Any server that honours Range requests and serves a faststart MP4 can be cut this way, and the header plus the span is all a clip ever needed.

If you want to see what the clip side looks like, the tool is here: paste a link, pick lines from the transcript, download the vertical cut. If all you need is the words, the YouTube transcript extractor gives you the transcript with a timecode on every line.

lumian

Builds VideoToShorts