Wan: Video From Text and From a Photo
Open-source video right in ComfyUI - bring an image to life or build a clip from scratch, plus a realistic look at hardware
What Wan from Alibaba is, how Wan 2.1 and 2.2 differ, how to generate video from text (T2V) and bring a photo to life (image-to-video), why Wan 2.5 with audio is a closed API, and how to run video generation on a modest GPU. Realistic expectations on clip length and render time.
Video Stopped Being Magic for the Select Few
Not that long ago, generating video locally was nearly impossible - either a paid service or a lot of painful workarounds. In 2026 open-source video arrived in ComfyUI, and the ruler here is Wan Wan is a family of open video models from Alibaba. Wan 2.1 (early 2025) and Wan 2.2 (mid-2025) were published with open weights and work in ComfyUI. They can generate video from text and bring a still photo to life. from Alibaba. Download the model and you’re making short clips on your own GPU.
Two Modes: From Text and From a Photo
- Text-to-video (T2V). You describe a scene in words and the model builds a clip from scratch. Useful when you have no source image: “coffee slowly pours into a cup, steam rising.”
- Photo-to-video (I2V, image-to-video). You give it an existing image and the model brings it to life - adds movement, wind, light glints, subtle animation. This is where the magic is for creative work: generate a frame in FLUX or Qwen, then use Wan to turn it into a living clip.
The combo worth remembering: FLUX/Qwen draws the frame - Wan brings it to life. First the perfect image, then the movement. That gives you far more control than asking for video straight from text.
Which Wan to Pick
Inside Wan 2.2 there are options for different hardware - the same “light vs heavy” principle as with images.
- A single compact checkpoint for both modes (text- and photo-to-video).
- Lighter on video memory - realistically runnable on a home GPU.
- Faster - great for trying things out and iterating.
- Higher quality - cleaner motion and detail.
- Separate checkpoints for text- and photo-to-video.
- Demands a powerful GPU and more time to render.
Where’s Wan With Audio?
You’ll come across Wan 2.5 - it has synchronized audio, 1080p, and longer clips. Sounds great, but there’s a catch: Wan 2.5 weights are closed Wan 2.5 is the version from September 2025 with synchronized audio, 1080p resolution, and longer clips. But its weights are NOT published - access is through a paid API only; it can't be run locally. - it’s a paid API, not a downloadable model. In ComfyUI you work with the open 2.1 and 2.2. Audio is usually added as a separate step or with a separate tool after the video is generated.
Running It and Setting Realistic Expectations
- Update ComfyUI to a recent version - Wan is supported natively and ready-made templates are in the menu.
- Download Wan 2.2: weak hardware - TI2V-5B, powerful hardware - A14B. Very weak - look for a quantized (GGUF) build.
- Decide on your mode: bring an existing photo to life (I2V) or build from text (T2V). Newcomers - go I2V.
- Open the matching Wan template from the ComfyUI examples.
- For I2V, provide your image and a short description of the movement. Run it and give it time to render.
Common Mistakes with Video
- Running the heavy A14B on a weak card. It will crash with an out-of-memory error. For home hardware - the light TI2V-5B or GGUF.
- Expecting a film. Local open-source video means short clips. Realistic expectations save your nerves.
- Going straight to text-to-video. Less control over the frame. Start with animating an existing photo.
- Looking for Wan 2.5 to download. Its weights are closed. Locally - only 2.1 / 2.2.
- Describing overly complex motion. Simple, clear movement is what Wan handles cleanly - ten simultaneous actions in one frame is a recipe for chaos.
TL;DR - если коротко
- Wan (Alibaba) - the main open video models of 2026. Wan 2.1 and 2.2 have open weights and work natively in ComfyUI.
- Two modes: text-to-video (a clip from scratch based on a description) and photo-to-video (image-to-video) - bring an existing image to life. The second one is the exciting part.
- Wan 2.2 comes in a light variant (a single compact TI2V-5B checkpoint) and a heavy one (A14B - cleaner, but hungrier for video memory).
- Wan 2.5 added audio and 1080p, but it's already a closed API - you can't run the weights locally. In ComfyUI you work with the open 2.1 / 2.2.
- Realistic expectations: video eats memory and time. Clips are short (seconds), render is not instant. Weak hardware - go with the light TI2V-5B or a quantized build.