AI Video Ad Generator Workflow: A Sportswear Running Commercial, Shot by Shot

Six shots for a fifteen-second sportswear spot, generated twice: once with three fixed sentences about the light pasted into every prompt and once without, with the color spread across the six measured both ways.

*No credit card required
Two panels: an invented brand spec listing name, wordmark, primary color, secondary color, line and rule; and a light block of three sentences about color temperature, key light ratio and camera height.
CapCut
CapCut
Sep 24, 2026

A fifteen-second spot is six shots and five joins. The shots are the part everyone generates. The joins are the part that decides whether you have a commercial or a folder of nice frames, and a join fails for a dull reason: the two shots either side of it were lit differently.

So this ai video ad generator workflow is built around one question. Six shot descriptions were sent twice, identical except that one set had three extra sentences about the light stapled to the end of every prompt. Then the color of all twelve was measured.

Invent the brand before you write a single shot

An ad needs a brand, and a generated ad needs an invented one. HALFMILE is the name used throughout here; it is not a company, and the point of the spec below is that you replace every line of it with your own.

Two panels: an invented brand spec listing name, wordmark, primary color, secondary color, line and rule; and a light block of three sentences about color temperature, key light ratio and camera height.

The brand spec is what you replace. The light block is what you paste.

One line of that spec does the most work and is the easiest to get wrong: the brand name stays out of the prompt. Each of the shot prompts used here ends with "no text, no logos, no printed words and no brand marks anywhere in the frame." The wordmark goes on as a text layer in the editor afterwards, where it is sharp, spelled correctly, and in your own typeface.

Six shots, and the three sentences that go on all of them

The shot list is the ordinary one for this kind of spot: the set at the start line, a side view at full stride, a shoe hitting water, a wide shot with the runner small, a macro for effort, and a product frame. Each one was written as a single sentence, with the same runner description in front of it.

A lone adult runner in plain unbranded black running kit and gray trainers. No text, no logos, no printed words and no brand marks anywhere in the frame. Close on the runner hands and front foot set on a wet running track at the start of a sprint. Cold overcast daylight at 5600 kelvin, key light from camera left about three stops brighter than the shadow side, camera at the runner hip height.

The bold part is the only thing that changes between the six. The sentence before it and the three sentences after it are constant.

Six generated stills in two rows: hands at a start line, a side view at full stride, a shoe spraying water, a wide shot against sky, a thigh macro and a pair of shoes on a block, all in cold grey light, each labelled with its a* and b* readings.

The six with the light block. Each is labelled with its median a* and b*.

The same six, with the light block deleted

The same six shot descriptions without the light block: a blue-sky wide shot, a warm golden macro of an eye, a bright overexposed side view and a grey sole shot, visibly mismatched.

The same shot sentences, the same model, the same 16:9 and 2K. Only the three sentences are gone.

These are all defensible photographs. The wide shot has a blue sky and white clouds. The macro is a warm golden close-up of an eye. The side view is bright and almost blown out. Any one of them would be fine on its own, and no two of them belong in the same fifteen seconds.

What an ai video ad generator workflow buys with three sentences

A scatter plot of median a* against b* with one dot per shot: six blue dots in a small box and six orange dots spread across a much larger box.

One dot per generated still. The box around each set is its bounding box.

Measuring the median a* and b* of each frame — the two color axes in CIELAB, where a* runs green to red and b* runs blue to yellow — the six shots with the light block sit inside a box 2 wide and 5 tall. The six without it need a box 15 by 18. On the blue-yellow axis, which is where a warm shot and a cold shot separate, the spread fell from 18 to 5.

That is the whole argument for the block. It costs nothing, it is three sentences, and it is the difference between six shots that share a color and six that do not.

Two things the block did not fix are worth being straight about. Median brightness still ranged from 132 to 208 across the six, because a wide shot full of sky is brighter than a close-up of a hand whatever the light is. And the third sentence, about camera height, was written but not measured here — color is easy to read off a frame, and camera height is not.

Two of the six came back off-brief

Generating six shots produced four usable ones. The sole shot put a photographer with a camera and a flash into the frame, which nothing in the prompt asked for. The macro was written for a bead of sweat at the temple and came back as a close-up of a thigh. The product frame left the runner in the background instead of clearing the frame for the shoes.

The pattern across the three failures is that the tightest briefs failed. A shot described by what fills the frame — hands, a stride, a runner against sky — came back as asked. A shot described by a small detail inside a larger scene needed the scene to be got right first, and the generator filled the rest of the frame with whatever it thought belonged there. Budget a re-roll for the macro and the product frame, or plan the spot around the four that work.

Three of them, animated, and the color holds

Three of the four usable stills were sent back through as five-second clips, each with one sentence of motion attached and the framing pinned down.

The runner holds a steady full stride and the road surface moves past beneath. Nothing else in the frame moves and no text or logo appears. Locked-off tripod shot, no zoom, no push in, no dolly, no pan and no crop, and the last frame has the same framing as the first frame.

Shot 2 of three: 1280 by 720, 24 frames per second, 5.04 seconds. Shown here as frames in sequence; the full video and a GIF version are supplied alongside this document.

Three frames side by side from the three finished clips, each labelled with its mean a*, b* and median lightness.

The three clips, measured the same way the stills were.

Across the three finished clips the a* spread is 1.0 and the b* spread is 4.5 — the same range the stills came back in. Whatever the light block did at the image stage survived the step into footage, which is the part that matters, because footage is what you cut.

Fifteen seconds, laid end to end

A CapCut timeline showing three video clips laid end to end, with the project duration reading 00:15:01.

The three shots on one timeline in the editor, reading 15.01 seconds. The clips were added by double-clicking each one in the Media panel; dragging them onto the timeline did not take.

Three five-second shots is fifteen seconds with no trimming at all, which is the plainest version of this spot and the one worth building first. The faster rhythm — cutting each shot down so nothing in the first eight seconds runs longer than a second and a half — is a second pass on the same three files, and it needs six shots rather than three to have anything to cut to.

The export step did not complete in this session. The assembly above is the state of the project as saved, and the three clips are delivered individually; the spot has not been rendered to a single file here.

What the shot list cost

A cost table: twelve stills at one credit each, one re-rolled still at two credits, three clips at 25 credits each, with the balance moving from 24,015 to 23,925.

The cost lines as shown under each request, read on 14 September 2026.

Request
Configuration shown
Cost line
Count
Six shots, twice over
Image mode, Seedream 4.5, 16:9, 2K
Generating 1 item will consume 1 credit.
12
One still re-sent
Image mode, Seedream 4.5, 16:9, 2K
Generating 1 item will consume 2 credits.
1
Three shots as clips
Video mode, Seedance 2.0 Mini, 16:9, 5s, 720p
Generating 1 item will consume 25 credits.
3

The balance in the header moved from 24,015 to 23,925 across the whole session, which is 90. Stills are the cheap half of this and the half that decides whether the spot works, which is the argument for settling the look on 1-credit frames before spending 25 on any of them. The stills and clips were made through the AI image generator and AI video generator flows. These figures are for these configurations on this account and this date, and other models, durations and resolutions are charged differently.

One request is worth singling out. The product shot came back with a cost line reading two credits rather than one, because the request planned two items instead of one; it then failed with "Couldn't generate. Credits returned." and was re-sent. The cost line is priced per request, not per shot, and the number in it is set before you press send. On a job with a dozen frames in it, that line is the thing to read before each send.

Cut the six-shot list down to the four that behave

The six-shot list at the top of this piece is the conventional one, and two of its six came back as something else. The four that behaved have something in common: each is described by what fills the frame. Hands on a line. A stride. A runner against sky. Shoes on a block, once the runner is told to leave.

So the shot list to write is a list of frames, not a list of details. Put the three light sentences on all of them, generate the whole list as 1-credit stills first, and read the color off the contact sheet before anything becomes a clip. The spread across those stills is the number that tells you whether you have a commercial or a folder.

HALFMILE is invented for this article and so is the runner. Keep it that way in your own version: an existing brand's name, wordmark, colors, slogan or endorser has no business in a generated ad, and a generated spot that looks like one of theirs is a problem whoever made it. The same goes for faces — the runner here was generated from a text description, and putting a real athlete, or anyone else, into a commercial they did not agree to is not a prompt question.

Say somewhere that the footage is AI-generated when you post it. The platform requires you to have rights and consent for what you upload, and the advertising rules that apply to a filmed commercial apply to this one too.

Also not covered: placement and channel strategy, beat-aligned cutting, color grading in the editor, voiceover, vertical versions and 4K export.

Written 14 September 2026. Every frame and clip was generated in CapCut for this article; no real person, product or brand appears in any of them, and HALFMILE is not a company. The twelve stills are 2560 by 1440 and were scaled to 1280 by 720 before measurement; the three clips are 1280 by 720, 24 frames per second, 121 frames, 5.04 seconds. Color readings are the median CIELAB a* and b* of the whole frame, and for clips the mean of those medians across all 121 frames. Spread means the largest reading minus the smallest across the six shots in a condition. Median lightness is quoted but not attributed to the light block, because a frame filled with sky is brighter than a close-up whatever the lighting is; the camera-height sentence was written into every prompt but not measured. Interface labels and cost lines reflect a single session on that date and the configurations named above, and may change.

Hot and trending