AI Video Made Crystal Clear · chapter 2: How Video Models Work

Two ways to build the denoiser

2026-10-02

Left, the U-Net lineage described in Weng (2024); right, the diffusion transformer over spacetime patches introduced by Peebles and Xie (2022) and used by Sora (2024), Veo 3 and Wan (2025). Both take a noisy latent plus the encoded prompt and return a slightly cleaner latent.

Below: the paragraph from the book that builds this idea, then the diagram itself (Figure 2.3), and a recap. About a minute of reading.

Sora, whose 2024 report was co-authored by Peebles, took the DiT to video with one more idea. Language models split text into word pieces called tokens; the report's line is "Whereas LLMs have text tokens, Sora has visual patches." A spacetime patch is a small cube cut out of the latent video: a few latent pixels wide, a few tall, a few frames deep. Cut the whole latent into these cubes, lay them out in a row, and you have a token sequence a transformer can read. The report calls the cutting step patchify. Figure 2.3 puts the two designs side by side.

Figure 2.3: Two ways to build the denoiser. Left, the U-Net lineage described in Weng (2024); right, the diffusion transformer over spacetime patches introduced by Peebles and Xie (2022) and used by Sora (2024), Veo 3 and Wan (2025). Both take a noisy latent plus the encoded prompt and return a slightly cleaner latent.
Figure 2.3: Two ways to build the denoiser. Left, the U-Net lineage described in Weng (2024); right, the diffusion transformer over spacetime patches introduced by Peebles and Xie (2022) and used by Sora (2024), Veo 3 and Wan (2025). Both take a noisy latent plus the encoded prompt and return a slightly cleaner latent.

Recap

  • The idea: Left, the U-Net lineage described in Weng (2024); right, the diffusion transformer over spacetime patches introduced by Peebles and Xie (2022) and used by Sora (2024), Veo 3 and Wan (2025).
  • The picture: Figure 2.3, from chapter 2 ("How Video Models Work") of AI Video Made Crystal Clear.
  • Go deeper: the chapter builds this step by step, with recipes and sources at the end.

This diagram is one of many in AI Video Made Crystal Clear.

Every chapter opens with the gist, draws the hard ideas, and ends with recipes and sources.

Get the book

All diagrams