Method

How to animate a product photo with AI: from catalogue still to video in an hour

4 August 202615 min read
Key points
  • A catalogue shot on white is the best source for animation, not the worst: a clean separation of product from background gives the model geometry it can read.
  • Never feed in infographics with text on them. Take the frame from before the graphics were added, and put the text back in the edit.
  • One long clip from a text prompt falls apart. What works: four segments of 3 to 5 seconds, each generated from its own start frame.
  • Physics is still the weak spot of every model: the best score on the academic VideoPhy-2 benchmark is 32.6%, against 100% for real footage.
  • There is no public data on video lifting conversion on the Russian marketplaces Wildberries and Ozon. The only figure a marketplace has ever quoted is plus 24% to click through rate, and it dates from 2022.
  • You cannot animate a supplier's photo or someone else's listing image: reworking a picture with a neural network is still an infringement of exclusive rights.

You can animate a product photo with AI in about an hour: take the catalogue frame from your listing, generate three or four short segments of 3 to 5 seconds, each from its own start frame, and cut them together. That hour is for a simple product with a good photo. A complex cross section, clear glass and small print on the packaging do not fit into it, and below is why.

Three generation modes that people keep confusing

Articles about animating photos lump everything together. There are three modes, and the choice decides whether the product stays yours.

Text-to-video. Generation from a text description. The model invents the product from scratch. Fine for advertising an abstract idea, no good for your specific SKU.

Image-to-video. You supply a photo and the model animates that exact image. There is a formal metric for how tight the anchor is: in the VBench benchmark the I2V track separately measures how close the video stays to the supplied picture, and models score around 97% there. Text-to-video has no such anchor at all: there is nothing to compare against.

Generation from a start and an end frame. You set not only the beginning of the segment but its end. Every top model of 2026 has this mode: Kling handles a first and last frame, Veo gives you start and end frames, Seedance accepts several frames pinned to timecodes. It is the most controllable option, and the method in this article is built on it.

Why one long clip falls apart

We made a clip for a door seal: a black profile with seven chambers in cross section and a red adhesive strip. First attempt: a single 15-second generation from a text description. The result: the cross section morphed into mush, the number of chambers changed from frame to frame, the black drifted towards grey, and the red backing strip stayed hanging on the door after the seal was stuck on, even though in reality you peel it off and throw it away.

The client summed it up in one line: "that is some kind of dragon, not my profile".

This is not a bug in one particular model, it is a failure class described in the literature. VideoPhy-2, presented at ICLR 2026, tests the physical common sense of generated video: the best model scores 32.6% on the full set and 21.9% on the hard subset. On the Physics-IQ scale, where a pair of real videos scores 100, the best generator gets 24.1, and the most visually realistic model only 8.7. A manual review of a sample found physical defects in 83% of clips.

Put simply: the model prioritises a plausible picture over preserving mass, volume and shape. The strip left hanging on the door is exactly that kind of break in object permanence.

Which photo works as input

Good news for sellers: a catalogue shot on a white background is the best source, not the worst. A clean separation of product from background gives the model readable geometry and honest depth, while a busy environment makes it guess what to move.

Works:

  • a shot on white or light grey where the product fills 70 to 80% of the frame: that is exactly what Wildberries, one of Russia's two big marketplaces, requires of a main image;
  • three quarter and side angles: they let the model read volume;
  • macro details: texture, a seam, a port, a thread.

Does not work:

  • screenshots and over compressed JPEGs pulled out of messengers: generation amplifies compression artefacts;
  • frames with hard shadows that hide the shape;
  • a photo that does not contain what you are asking to see. If you only shot the front, the model will invent a 360 degree turn rather than show one.

Resolution is simple: Wildberries accepts from 700×900 and recommends 900×1200 at a 3:4 ratio, and Ozon, the other big Russian marketplace, recommends the same 900×1200. That is enough for generation too. Go vertical from the start: generating in 16:9 and cropping to vertical afterwards cuts the product off at the edges.

If you have exactly one usable photo, you will still get a clip, just a shorter one. A single frame yields two safe segments: a slow push in and a light parallax. Everything else needs angles the source does not have, and the model will start inventing them.

Do not feed in infographics

This gets its own section, because half of a seller's listing is frames covered in labels, arrows and spec callouts.

Infographics do not work as input. The model rewrites letters, changes captions and draws in elements that were never there: it redistributes pixels between frames, it does not understand what the text means. On video this shows up more than on a still, because every glyph has to match across each consecutive frame.

Take the source frame from before the graphics were laid on. If you need text, put it back over the finished video in the edit: that way you keep both the kerning and the exact wording.

The same goes for the barcode, the ingredients list and the regulatory markings on the packaging: you never hand that area over to motion.

There is a second reason not to drag finished infographics into generation. Wildberries bans marketing copy, discount labels and watermarks on photos and fines sellers $357 for them, while descriptive infographics with specifications are allowed. The line is thin, and there is no reason to run those frames through a model that likes to add elements that were never there.

Shot by shot assembly instead of one clip

The method that works looks like this: the clip is assembled from segments of 3 to 5 seconds, each with its own approved start frame.

Why not longer: past ten seconds every model accumulates artefacts, shape drift, stuttering motion, background deformation. A short segment physically does not have time to fall apart.

Why from a start frame: the model has no freedom to invent an in between shape. It builds the motion out from what is already in the frame.

With the door seal this worked immediately: as soon as we replaced the single fifteen second clip with five segments built from agreed frames, the cross section stopped floating and the colour stopped drifting.

References: making the model repeat your product, not a lookalike

One frame holds the shape inside a segment, but between segments the product can drift. The fix is a reference set: attach 2 to 3 photos from different angles to every generation. Kling accepts up to four reference images, Seedance takes up to twelve files with assigned roles, but the practice is the same everywhere: more than three usually hurts. The model starts averaging and gives you a product that resembles yours instead of yours.

If you have renders from the manufacturer, attach them: the geometry is cleaner than in a photo, and the model reads the shape more accurately.

A working set looks like this: a wide view on white, a three quarter angle, and one macro of the spot the buyer uses to recognise the product. For the door seal that was the end face with the chambers, for a bottle it would be the cap, for clothing it is the fit at the shoulder. That is the spot the client checks first, and the spot where they catch a substitution.

If the product has already drifted in generation, adding references is useless, you need to cut instead: keep the two cleanest frames, strip every extra movement out of the prompt and generate again. The model fails not from a lack of data but from contradictions in it.

Which movements hold the shape and which break it

Ranked by risk, from safe to dangerous.

Movement Risk Comment
Slow push in, product centred low the default choice for product shots
Parallax, layers moving at different speeds low works well on a clean background
Slow 20 to 30 degree orbit medium needs a source where the volume reads
A finger or a hand interacting with the product medium a strong shot, but it takes retakes
Full 360 degree fly around high the model invents the side you cannot see
Assembly or disassembly into parts high also the most striking, when it lands
Liquids, fabric, fur very high physics breaks measurably

The main rule: one movement per segment. A request like "the camera flies around, the product opens up and the background changes" is close to a guaranteed reject. Keep the speed slow: models lose detail in fast movement.

Products AI cannot handle yet

An honest list, so you do not burn credits for nothing:

  • jewellery and anything mirrored: metal turns to mush, stones rearrange between frames;
  • anything transparent: glass, bottles, packaging with a window. Segmentation algorithms fail on transparent surfaces many times more often than on matte ones;
  • watches: the numbers on the dial get rewritten;
  • cosmetics with embossing and small ingredient print: the same problem as with text;
  • electronics with ports and screens: the model invents the interface;
  • fur, knitwear, liquids: physics is the first thing to go.

The pattern is simple: the more the buyer's eye is trained on a detail, the worse it goes. A matte object with a simple shape comes alive in one take, a perfume bottle does not come alive in ten.

The take acceptance checklist

Look at every generated segment and reject without regret. There is no single published norm for the number of takes: practitioners quote anywhere from two to ten per usable shot, and it depends on the product more than on the model.

  1. The shape of the product matches the source, nothing appeared and nothing vanished.
  2. The count of details is the same: chambers, buttons, fastenings, facets.
  3. Colour and finish are unchanged, matte has not turned glossy.
  4. The logo is intact, the letters have not warped.
  5. Text is legible, or there is no text in the frame at all.
  6. Proportions have not slipped by the end of the segment.
  7. The first and last two seconds are clean, because artefacts build up at the edges.
  8. The physics does not lie: nothing passes through anything, floats or disappears.

Point eight is the one we tripped over with the adhesive strip.

It helps to count not takes but the cost of a usable shot: divide every generation spent on the scene by the number accepted. That metric shows straight away where the prompt is bad and where the product is simply hard. It also shows when to stop: if after five runs not a single shot has been accepted, it is not bad luck, it is the staging of the scene, and the scene is what you rewrite instead of generating a sixth time.

What the built-in Wildberries and Ozon tools give you

A fair question: why generate anything yourself if the marketplace can do it.

In January 2026 Wildberries launched Video Covers: video generation from a photo right inside the seller dashboard, around a minute per asset. Access comes with the Jem subscription, whose Standard plan has cost $328 for 30 days since March 2026. Separately, since 25 December 2025 the free Photo Editor 3.0 has been open to all sellers based in Russia. Ozon has free photo generation in its dashboard: the source frame mode removes the background and substitutes a new one, returning up to four variants in 20 to 60 seconds. Ozon has had automatic video cover generation since 2023.

Here is the difference. The built-in tools mostly apply effects to a still image: zoom, soft transitions, pulsing, slideshows. The camera moves, the product stands still. Real generation moves the object itself: the profile compresses, the cap opens, the fabric gives under a hand.

If you need a tidy animated frame for a cover, the built-in tool is enough. If you need to show how the product works, it is not.

The one hour route

For a simple product with a catalogue photo already in hand:

  1. Pick one frame per segment: the wide shot, the detail, the product in use. Five minutes.
  2. Gather 2 to 3 references from different angles. Five minutes.
  3. Generate four 4-second segments, budgeting 2 to 3 takes for each. One clip takes 1 to 2 minutes to generate, and with selection that comes to around forty minutes.
  4. Cut it together, lay on the text, export. Ten minutes. Put the joins on a change of shot, then the transition does not read; if two neighbouring segments come from the same camera position, separate them with a short dip to black a couple of frames long.

That is the shape in which an hour is realistic. If your product is on the list above, or your only photo is a straight on shot on white, the hour turns into several sessions: we broke down hour by hour how long a full clip takes and what the deadline is made of.

For comparison: a studio video review of 15 seconds costs around $71, and that is the cheapest option with a real set, no hands in frame and no model. The catalogue shot you have already paid for cost $7 to $11.

Price up a clip for your product →

Rights: to other people's photos and to your own result

The two questions sellers ask most often.

You cannot animate someone else's photos. Reworking an image is a separate form of use under subparagraph 9 of paragraph 2 of article 1270 of the Russian Civil Code, and it requires the rights holder's consent. A neural network changes nothing in that chain: the result is still a derivative work. Compensation is awarded under article 1301 of the Civil Code, and on Wildberries the dispute goes to Digital Arbitration: 10 days to reply, and after two confirmed infringements the category is closed for 30 days.

On rights to generated video, Russia currently holds two opposing positions. In the deepfake case involving Keanu Reeves the court awarded $7,143 and stated that a neural network is an additional processing tool, and that the authors' creative contribution to the script and the edit does not disappear. Then on 2 August 2026 a Moscow district court refused protection to images reworked by a neural network, calling the entering of commands "simple mechanical actions"; that is a first instance ruling, and there was no data on it entering into force at the time of publication.

The practical conclusion from both positions is the same: whether the result can be protected rests on a demonstrated human contribution, the script, the storyboard, the selection of takes, the edit. A clip assembled shot by shot from approved frames stands on firmer ground than a button press in a dashboard.

The truth about conversion

Here we have to disappoint you.

Numbers circulate online: video lifts conversion by 15%, by 16%, raises add to cart from 18% to 35%. We checked each one. Not a single one leads to a study, a methodology or a sample: they are agency blogs reprinting each other, citing internal analytics with no document behind it.

The only figure a marketplace has stated publicly: video covers raise the click through rate of a listing by 24%. That is an Ozon statement from September 2022. In 2026 it is republished with no date, as if it were fresh statistics.

What can be said honestly: both marketplaces are investing in generating video from photos, Ozon since 2023, Wildberries launching its own tools at the end of 2025 and the start of 2026. Companies do not spend development budget on things that do not move sales. But nobody has promised you a specific lift on your own listing, and if someone promises you a number, ask for the link to the study.

What the generations themselves cost and how you pay for them is broken down in what it costs to make a video with AI. What goes into the price of a finished turnkey clip is covered in how much an AI video ad costs.

FAQ

Can I animate a product photo for free?

Partly. On Wildberries, Photo Editor 3.0 has been free since 25 December 2025, and Ozon has free photo generation inside the seller dashboard. But the built-in marketplace tools mostly move the camera across a still image rather than the product itself. Real motion generation costs credits, and we break those prices down in our article on what it costs to make a video with AI.

Is one photo enough, or do I need every angle?

For one short clip a single frame is enough. But to stop the product drifting between clips, it is better to gather 2 to 3 references: straight on, three quarters, and a macro of the detail. More than five usually hurts, because the model starts averaging them out.

Why does the product warp and change shape, and can a prompt fix it?

Only partly. A generative model builds a plausible scene, it does not preserve catalogue accuracy for your SKU. What helps is not the wording but the method: short segments, anchoring to a start frame, minimal movement, and rejecting bad takes.

What if the packaging has text, an ingredients list or a barcode?

Keep that area out of the motion. Models rewrite small text almost every time: they redistribute pixels between frames, they do not understand what the writing means. Frame the movement so the text sits outside the active area, and lay the wording you need back over the top in the edit.

Can I animate a supplier's photo or a screenshot of someone else's listing?

No. Reworking an image is a separate form of use under subparagraph 9 of paragraph 2 of article 1270 of the Russian Civil Code, and it requires the rights holder's consent. The fact that a neural network did the reworking changes nothing. On Wildberries such disputes go to Digital Arbitration: you get 10 days to reply, and after two confirmed infringements the creation of listings under that brand in the category is closed for 30 days.

Do I have to label the clip as made with AI?

As of August 2026, no. Federal law 243-FZ on the development of AI technologies comes into force on 1 September 2026, while the labelling provisions apply from 1 March 2027, and for the author of the content labelling is described as a right, not an obligation. The duty to provide the technical means for labelling sits with large platforms. Russian language blogs constantly mix up those two dates.

Will the marketplace reject a clip made with AI?

Not for being made with AI. It will be rejected if the clip does not match the product. Ozon moderation checks more than 12 million listings a day and hides around 50 thousand of them. If the generation flattered the product by removing a scratch, shifting a shade or adding bulk, that is grounds for rejection as not matching the description. Format requirements are a separate matter, and each marketplace sets its own.

Is it worth paying for a subscription just to get the marketplace's built-in AI video?

Do the maths on volume. Video Covers on Wildberries come with the Jem subscription, whose Standard plan has cost $328 for 30 days since March 2026. That is reasonable if you use the rest of the subscription, and expensive if all you want is clips.

About the author
Сергей — specialist at Blind Chameleon studio
Сергей

Runs orders at Blind Chameleon studio

Quotes clips and works with Veo, Kling and Seedance every day: the prices in this article come from our own invoices, not from someone else’s round-up. He will help you size a budget for your job and say plainly if the job is cheaper to solve without us.

Contacts

Get in touch directly

For a price, use the calculator above. If you have a question, need an invoice for a company or have an urgent brief, write to us and we will answer within the hour during working hours.

Индивидуальный предприниматель Мироненко Анна Сергеевна · Tax ID 237310336801. We work under a contract and issue an invoice for every payment. Full company details and postal address on request by email.