AI Video Made Crystal Clear · chapter 2: How Video Models Work
The generation pipeline, the anchor schema for this book
2026-09-09
Drawn from the Sora technical report (Brooks, Peebles, et al., 2024), the Veo 3 technical report and the Wan paper (2025); the six stages are common to all three.
Below: the paragraph from the book that builds this idea, then the diagram itself (Figure 2.1), and a recap. About a minute of reading.
Every video generator on sale in September 2026, from Google's Veo 3.1 to Alibaba's open-weight Wan, is a variation on one design. OpenAI's 2024 Sora report described it, Google's Veo 3 technical report describes it in almost the same words, and the Wan paper describes it again for an open model. The vendors differ in scale, training data and polish, not in shape. Figure 2.1 is that shape. Chapter 3 zooms into its input side (what you can feed it), chapter 6 into the prompt stage, chapter 7 into why identity does not survive between runs, and chapter 8 into the audio branch.

Recap
- The idea: Drawn from the Sora technical report (Brooks, Peebles, et al., 2024), the Veo 3 technical report and the Wan paper (2025); the six stages are common to all three.
- The picture: Figure 2.1, from chapter 2 ("How Video Models Work") of AI Video Made Crystal Clear.
- Go deeper: the chapter builds this step by step, with recipes and sources at the end.
This diagram is one of many in AI Video Made Crystal Clear.
Every chapter opens with the gist, draws the hard ideas, and ends with recipes and sources.
Get the book
