In the five-second clip this article starts from, the best second and the worst second are the same second. At 1.58 seconds the picture is changing more than anywhere else in the clip, the dress is out in a full circle, and her hands have stopped being hands. Everything that makes the frame worth posting is also what broke it.
So the edit is not decoration. It is the thing that decides whether the clip is usable at all: which half second comes out, what the cut costs when it comes out, and where the one speed change is allowed to sit. All of it was made at 16:9 in CapCut from a single generated photo, and every number below was measured on the exported file.
The finished piece, 1280 by 720, 30 frames per second, 4.90 seconds, with a track from the audio library under it. Shown here as frames in sequence; the full video and a GIF version are supplied alongside this document.
The second that changes most is the second that breaks
Measuring the mean change per pixel between each pair of frames gives one number per frame, and it says where a clip is fast. The fast clip averages 3.02 across its five seconds and peaks at 5.14 at 1.58 seconds. A second clip made from the same photo, asked for a small step instead, averages 1.10 and stays well below that peak.
The two generated clips measured the same way. The shaded band is the 0.65 seconds that was later removed; the peak inside it is the frame shown below.
The peak is where the damage is. At 1.58 seconds her right hand is a smooth pink shape with no fingers in it, and her left hand is a smear against the wall. The slow clip at the same moment has five fingers and a wrist.
The same moment in the two clips, and the same hand enlarged four times in each. Nothing was retouched; these are frames from the downloaded files.
This is worth stating plainly because it sets the whole plan: fast movement is what the format is for, and fast movement is what the model gets wrong. You are not choosing between a good clip and a broken one. You are choosing between a slow clip with nothing to fix and a fast clip with half a second to remove.
What the photo has to give the model
Whatever the photo does not show has to be invented once the figure starts moving, and in this run the invented parts are the parts that came back broken. So the photo is chosen to hide as little as possible.
The source photo with the four things that matter marked. Generated at Seedream 4.5, 16:9, 2K.
Four rules, in order of how much they cost you when broken. The whole body has to be in frame with space above the head and below the feet, because a body cropped at the knees gives the model nothing to stand on and the legs it invents wander. Both hands have to be clear of the body, since hands crossing a torso are read as one shape and come apart the moment they separate. The floor line has to be visible, or the figure drifts up and down against a background with no ground in it. And the background should be plain, because a wall with nothing on it cannot be redrawn wrongly.
Photoreal photograph, 16:9 landscape. A woman in her twenties stands full length in the centre of the frame on a plain concrete floor in front of a plain pale wall, both arms held out clear of her body, feet apart in a wide stance, wearing a fitted red dress and flat shoes, looking at the camera. Even soft light with no shadow on the wall, sharp focus, the whole body inside the frame with space above her head and below her feet. No text, no logos, no brand marks.
That request was sent through the AI image generator flow in image mode. A photograph of your own works in the same place, and the same four rules apply to it.
Two ways to ask for the same dance
Both clips were made from that one photo through the AI video generator flow at Seedance 2.0 Mini, 16:9, five seconds, 720p. The difference is one paragraph of wording.
The camera does not move. She dances a fast turning step on the spot: she spins a full turn, throws both arms up and out, steps wide to one side and spins back the other way, moving quickly the whole time, her dress swinging out. Her feet stay in the same place on the floor. No cut, no zoom, no pan.
The camera does not move. She dances a slow step on the spot: she shifts her weight from one foot to the other, raises one arm above her head and lowers it again, and turns her shoulders a little each way. Both feet stay on the floor the whole time and her arms stay clear of her body. No cut, no zoom, no pan.
Three phrases in both are doing work rather than describing. The camera does not move and no cut, no zoom, no pan keep the clip as one continuous shot, which is what makes a cut inside it possible later. On the spot and her feet stay in the same place keep the figure in the frame instead of walking out of it. In the slow version, her arms stay clear of her body is the same rule as the photo, carried into the motion.
The fast clip is roughly three times as busy as the slow one, and that is the whole trade. Both clips have one larger step at the very start, where they leave the photo and begin to move: 3.94 in the fast clip and 4.70 in the slow one, between the first two frames. After that the slow clip stays under 1.95. Both came back on the first request.
Taking the broken window out, and what the cut costs
The fast clip was uploaded to the CapCut web editor and cut on the timeline: the playhead at 1.30 seconds, split; the playhead at 1.95 seconds, split; the middle piece deleted. Deleting a piece on the video track closes the gap behind it, so the two remaining pieces meet on their own and the total drops by the 0.65 seconds that came out.
The two frames that now sit next to each other, and how large that step is compared with the rest of the piece.
On the export the join lands at 1.23 seconds and measures 14.8 against a typical 2.6 for a neighboring pair of frames. That is a visible jump, and it should be: she is bent forward in one frame and upright with her arms raised in the next. Removing broken frames does not hide anything, it swaps one problem for another, and the second problem is the cheaper one because a jump cut is a normal thing for a dance clip to contain and a hand with no fingers is not.
There is no way to soften this by choosing a better cut point either. The window that has to go is set by where the model failed, not by where a cut would look good, so the only free choice left in the piece is the speed change.
Where the speed change goes
Two things have to agree at a speed change: an accent in the music, and a moment when the dance is not doing anything interesting. Both can be measured rather than guessed.
The track came from the audio library in the editor, dragged from the music row onto the audio track and trimmed to the length of the picture. Measured on the export, its accents fall about every 0.53 seconds, which is a little under 111 beats a minute, at 0.45, 0.98, 1.51, 2.04, 2.58, 3.11, 3.64 and 4.17 seconds. The slow points in the dance come from the same frame-to-frame measurement as before, taken on the cut piece before the speed change went in.
The two clocks side by side. The speed change went in at 2.10 seconds, between the accent at 2.04 and the slow point at 2.13.
The first attempt put the speed change at 1.85 seconds, on a slow point, and measuring it against the track afterward put it 5.7 frames from the nearest accent, which at 30 frames a second is nearly a fifth of a second late. Moving it to 2.10 seconds puts it 1.7 frames after the accent at 2.04 and one frame before the slow point at 2.13. The end of the slowed section, at 3.10 seconds, lands 0.2 frames from the accent at 3.11.
The change itself is half a second of the clip set to 0.5x in the Speed panel, which plays for one second. Two frames of tolerance is the working rule here: at 30 frames a second that is 0.07 seconds, and the closest an accent and a slow point come anywhere in this piece is 1.6 frames.
The finished timeline. The badge on the selected piece reads 0.50x; the cut at 1.23 seconds sits inside the left box.
Two things about that second are worth knowing before you use it. At 0.5x the editor repeats frames rather than making new ones, so in the slowed second every other frame is identical to the one before it and the measured change alternates between about three and zero. And a slowed second reads as slow in the numbers as well as on screen: 1.21 average change inside it against 2.38 across the rest of the piece.
The timeline toolbar also carries a Beats detection control for a selected audio clip. On this account, on this date, the switch did not turn on for the library track used here, so the accents above were measured on the exported audio instead.
Why every cut stayed inside one clip
Two clips were generated, and only one of them is in the finished piece. That was a measured decision rather than a preference.
The two clips agree about the picture. Their first frames differ by 0.71, which is as close to identical as two files get, because both were built from the same photo. They also keep her at the same distance: measured as a share of the frame, the red of the dress averages 2.8 percent in the fast clip and 3.0 percent in the slow one. So the thing people usually worry about when joining generated clips, that the person changes size between them, did not show up in these two.
What does not match is the pose. Cutting from the end of the fast clip to the start of the slow one measures 18.4, and the other way around 11.0, against 14.8 for the jump cut inside the fast clip. The cost of a join is about what the body is doing on either side of it, not about which file the frames came from. Keeping the cuts inside one clip is what let this piece have exactly one visible jump instead of two.
What the three requests cost
The cost lines as CapCut showed them under each request, read on 10 September 2026.
Fifty-one credits for a 4.90-second piece, and no cost line appeared for the editing, the speed change, the library track or the export. Ten generated seconds became just under five, which is the shape of this kind of edit: one clip is spent on the material you keep and one on the comparison that tells you what your material is worth. These figures are for these configurations on this account and this date and come from the line the composer shows before each request; other models, durations, resolutions and modes are charged differently, and prices change.
The fix that does not work
The obvious alternative is to slow the broken part down instead of cutting it out, on the theory that a speed change hides damage. It does the opposite here. In this export, the 0.5x speed change repeated frames rather than creating new intermediate detail, so a frame with a broken hand stayed visible longer. Nothing was smoothed in this test, and the artifact became easier to notice.
The other alternative is to ask for the slow step and skip the whole problem. That clip is intact, it needs no cut, and it also does not do the thing the format is for. Across the five seconds the red of the dress covers between 2.20 and 3.18 percent of the frame in the slow clip, a swing of one point, against 1.84 to 3.78 percent in the fast one, a swing of nearly two. That is the difference between a skirt that hangs and a skirt that flies. If the slow version is what you want, generate it and stop there; the editing in this article is only worth doing on material that moves enough to break.
What is not covered here: stored dance templates and template packs, captions and subtitles, transitions, filters and effects, image enhancement and 4K export, and any editing of a real person's photograph. If you use your own picture, use one of yourself or of someone who has agreed to it.
The planning row for this page called it a template video. No stored template was used in this session: the clips were generated from a photo through the AI video generator flow and cut by hand in the web editor, so the word came out of the title. The address of the page is unchanged, and templates are a different route to a similar result through the template explorer.
Written 10 September 2026. The photo and both clips were generated in CapCut for this article and assembled in its web editor; the dancer is invented and no real person, choreographer, track or account appears. Interface labels and cost lines reflect a single session on that date and the configurations named above, and may change. Frame-to-frame change is the mean absolute difference per pixel between neighboring frames of the downloaded files; music accents were found by measuring the rise in the exported audio's spectrum and fitting one spacing across the piece; the slow points were measured on the cut piece with the speed change taken back out.