Thirty seconds is roughly 75 spoken words. That is the entire budget of a short product video: 75 words, a handful of shots, and one action you want the viewer to take. Most product videos fail because they spend that budget like a brochure — a logo animation, three features, a montage — instead of like an argument. A high-converting 30-second video is an argument with four beats: hook, problem, proof, ask. Each beat has a time slot, and each slot has a job.

0–3 seconds: the hook is a promise, not a logo

Never open with a logo animation. The first three seconds decide whether the video gets watched at all, and a logo answers a question nobody asked. Open instead on the most visual moment your product has — the after state, not the setup. A stain remover opens on the stain vanishing. A scheduling tool opens on a chaotic calendar snapping into order. If your product's payoff isn't visual, open on a bold outcome statement rendered as an on-screen caption, because the sound is probably off.

There is one hard test for the hook: by second three, a stranger should know what category of thing this is. Not the brand name, not the feature list — the category. Confusion in second three becomes a scroll in second four.

3–10 seconds: name one pain, in the customer's words

This is where the viewer decides the video is for them. Name exactly one pain — not three, one. Pull the phrasing from real customer language: support tickets, reviews, sales-call notes. If your customers say the project plan died in week two, say that; do not translate it into alignment challenges across stakeholders. Specific pain makes the right viewer lean in and the wrong viewer leave, and both outcomes are good — you are buying qualified attention, not raw attention.

10–22 seconds: the demo shows hands, screens, or results

Twelve seconds of proof, and only proof. The structure inside this slot: show the mechanism for three or four seconds, slow slightly on the single aha moment, then land the result. Cut everything that explains setup, onboarding, or menus — nobody converts on a settings screen. And demo exactly one capability: the one that resolves the pain you just named. A product with ten features has ten videos, not one video with ten features.

A worked example for a project-management tool: 0–3s, a tangled timeline auto-arranges itself into a clean plan. 3–10s, caption reads: your plan died in week two — again. 10–22s, a hand drags one deadline; every dependent task reflows on screen, and a teammate's phone pings with the update. 22–30s, caption: start free, no card, and a URL. Four beats, one feature, one action.

22–30 seconds: a CTA the viewer can execute in one tap

The close introduces nothing new. One verb, one destination, risk removed: try it free, no card needed. If the product is cheap, show the price — a visible price is a filter that improves the quality of every click. Hold the end card for a full two to three seconds with the logo, the CTA, and the URL together, because plenty of viewers decide during that hold. A CTA that requires typing, searching, or remembering is a CTA that requires forgetting you by lunchtime.

Captions: design for sound-off first

Assume the sound is off — social feeds autoplay muted, and viewers in offices, queues, and beds rarely turn it on. That means captions are not an accessibility afterthought; they are the primary script channel. Write them as their own layer, not a verbatim transcript of the voiceover. Compress: the voice can say twelve words while the caption says five.

  • Maximum two lines and roughly six words on screen at once — captions are read in glances.
  • High contrast, inside safe margins, and clear of platform UI zones (usernames, buttons, and progress bars eat the edges).
  • Emphasize the one keyword per line with weight or color, not the whole sentence.
  • The video must make its full argument with captions alone; sound should be a bonus, never a requirement.
  • Keep captions inside the tightest crop you will publish, so one caption layer survives every placement.

Aspect ratios: one master, three crops

Placement decides shape. Vertical 9:16 owns Reels, TikTok, Shorts, and Stories, where the video fills the phone. Square 1:1 or portrait 4:5 wins in-feed placements, where taller frames occupy more scroll real estate than 16:9 ever will. Widescreen 16:9 remains right for YouTube in-stream and website embeds. The efficient path is to compose one master with all critical action — hands, screens, product, captions — held in a center-safe zone, then export all three crops from it. Remember that vertical platforms overlay their own UI along the bottom of the frame, so nothing important should live there.

If paid social is your main placement, build the 9:16 first and derive the rest, because vertical is the least forgiving frame. Auto-reframing tools — including a platform like AI BOSS — will happily generate every crop and re-fit captions in one pass, but review each output with human eyes: automated cropping is notorious for cutting off the exact hand or screen your demo depends on.

If a stranger can't tell what your product does with the sound off by second ten, no call to action can save the last twenty.

Measure the beats separately. The first-seconds hold rate grades your hook, mid-video drop-off grades your problem and demo, and click-through grades your CTA. Iterate the hook first — it moves the numbers most because it gates everything downstream — then the CTA, then the demo. Re-cut one beat at a time, re-run, and let the placement data pick the winner. Thirty seconds is small enough to treat like copywriting: draft, test, tighten, repeat.