Scale or control: the industry just placed its bets

Clip from a video generation made with LTX 2.5 Source: https://ltx.io/model/ltx-2-5

Two bets are being placed on AI video right now, and both are coming from the labs that build the actual models, not from tool makers layering something on top.

One bet is scale. This month Alibaba’s Wan 3.0 opened public beta with 30-second single-pass generation and document-to-video input, meaning you can go straight from a slide deck to a shot. Longer clips, bigger inputs, more of everything a demo reel loves to show off. It is an impressive engineering achievement and it does not touch the actual problem professionals run into, which is getting the model to do the specific thing you meant on the first or second try.

The other bet is control. Lightricks shipped LTX-2.5 as open weights this month too, and for the first time made hand-drawn trajectory control a first-class, documented part of the model’s own toolchain. Not a workaround somebody bolted on in a community graph. A supported input mode, built by the people who make the engine. That is a different kind of signal than a feature request. It is an admission that a text box was never going to be enough on its own.

Scale gets you a longer clip. Control gets you the clip you meant, on the take you actually wanted, without rerolling the whole thing and hoping the dice land differently. Those are not the same product, and they are not for the same buyer.

We are not neutral here, and we have not been for a while. Production work runs on the second bet. Nobody making a real deliverable, on a schedule, for a client who is going to look at the frame closely, can afford to keep pulling the lever and hoping. That is the whole reason Fossa Tether exists as keyframes and motion paths inside a timeline an animator already knows, instead of another chat box promising to understand you eventually.

We spent the past two weeks on our own version of this same argument: pulling finished shots apart instead of only building new ones. Multi-object tracking now finds every subject in a frame on its own, a clean-plate fill covers what was hidden behind them, and each piece exports as its own layer with proper full-resolution alpha, not a blocky matte you fight in post. It is still on the workbench, not in customers’ hands yet. Early tests are clean enough that we wanted to say so before it ships, not only after.

The tools getting good this year are the ones that hand you a path, not the ones that hand you a longer prompt box. That is the bet we are making, and this month two of the biggest labs in the field started making it too.

Next
Next

The integration problem: why AI video demos never survive contact with a pipeline