Continue character identity with reference video and create short narratives with sound
wan2.6-r2v is a video reference generation model in Alibaba's Wan 2.6 series, designed to bring character appearances from existing videos into new scenes and stories. It establishes character identity from reference material, then combines text descriptions to arrange actions, environments, and narratives. With multi-character and audio synchronization capabilities, it is suited for continuous character shorts, branded character content, and creative shot production.
Input parameters and result formats vary by service. Use the public API for this model and follow its guide for generation, task retrieval and editing operations.
Specifications and API Features
Creation method
Joint generation from reference video and text prompts
Native resolution
720P, 1080P
Native duration
2–10 seconds
Native video output
30 fps, MP4
Characters and audio
Multi-character, narrative, audio synchronization
Platform media input
POST /wan/videos; reference_video_urls submits reference video URLs
Platform task results
Supports asynchronous tasks; returns video URLs, dimensions, thumbnails, and other information
Resolution, duration, frame rate, and format are native specifications; media submission and result retrieval use this platform's Wan video API.
Core Capabilities
Bring the same character into new scenes
The focus of video reference is character identity, rather than guessing a character's appearance again with text. wan2.6-r2v can extract character appearances from existing clips and use them to create new scenes. Prepare material in which the character is clearly identifiable, then describe the target environment, intended clothing, and actions. This is better suited for continuously creating content around a fixed character.
Organize short narratives around multiple characters
The model supports multiple characters and narratives, making it suitable for shorts depicting characters meeting, interacting, or completing actions together. When creating, specify each character's identity, position, and sequence of actions so that events within a short duration have a clear storyline. There is no need to fit an entire complex story into one clip; it can be split into segments by plot.
Include sound in shot design
wan2.6-r2v's native capabilities include audio synchronization, making it suitable for planning sound requirements together with character actions and scene atmosphere. Prompts can clearly specify which events need sound accompaniment, rather than only listing visual tags. After production, check the correspondence between visuals and sound, then decide whether additional editing or voice-over is needed.
Use Cases
Series of Short Films with Fixed Characters
Using an existing character video as a reference, enter the location, actions, and emotional changes for this episode to generate new clips that preserve the character's identity. Suitable for creations such as having a character move from indoors to the street or enter different story environments. The deliverable is a short video that can continue to be edited, allowing the production team to review character traits and plot continuity clip by clip.
Situational Concepts for Brand Characters
Design different usage scenarios, interaction methods, and opening actions around existing brand characters or on-camera talent, and use reference videos plus creative scripts to generate candidate shots. Suitable for first validating how characters perform in new scenes, then selecting appropriate clips for advertising production; product logos and text details should still be checked during the final production stage.
Storyboard Trials for Multi-Character Interactions
Organize a short plot into shot descriptions covering who appears, who acts, and what happens, then generate interaction clips with character reference materials. The deliverables can be used to evaluate composition, character relationships, and narrative pacing, helping the team discuss filming plans. Test key interactions first, then combine selected shots into a more complete story.
How to Choose This Model
For Existing Video Characters, Consider r2v First
If the core task is to place a character from an existing video into a new scene, wan2.6-r2v better fits the workflow than describing the character from scratch. When you only have a text concept, you can choose wan2.6-t2v; when starting with a still image, consider wan2.6-i2v. Both have publicly available durations of 2–15 seconds, while this model supports 2–10 seconds, so consider both the material type and shot length when choosing.
Choose Between the Standard Version and Flash Based on Your Goal
wan2.6-r2v explicitly supports audio synchronization, multiple characters, and narrative capabilities, making it suitable for short films that need characters, events, and sound designed together. wan2.6-r2v-flash is a reference-video variant that emphasizes fast generation and can be considered when iteration speed matters. Do not assume that their sound controls are fully identical based solely on their names; test each one according to your actual creative goals.
Get Started
Prepare the Inputs for This Task First
Put the reference video in reference_video_urls, and describe the character and new scene in prompt; this model does not directly trim or edit the original video.
Select the Actual Model and Output
Specify model=wan2.6-r2v for /wan/videos, submit according to the reference video guide, and start with 5 seconds at 720P; Wan 2.6 does not use Wan 3's 30-second duration, automatic duration, or all-purpose media parameter.
Retrieve the Video and Check Audio and Visuals
Use async=true to save the task_id, query /wan/tasks or receive a callback; after completion, retrieve video_url, check the subject, actions, and sound, then proceed to editing.
Trial suggestion: character enters a new scene
Input and goal
Use the person in the reference video as the character, have them walk into a bookstore, pick up a book, and look at the camera, preserving their appearance, clothing, and vocal characteristics, with warm scene lighting.
Acceptance and next steps
Use reference_video_urls to submit identity references; check the character, scene transitions, and audiovisual performance, and do not treat the task as frame-by-frame editing of the original video.
Usage boundaries
The native output duration of this model is 2–10 seconds, making it suitable for short shots rather than generating a long narrative in one go. Stories beyond this range should be split into several segments, with character actions and plot highlights arranged separately, then connected through editing; longer durations generated from text or images in the same series cannot be applied directly.
Transferring a character's appearance does not mean copying the original video frame by frame, nor does it mean performing precise replacements or motion reenactments on existing video. If the goal is to preserve the original camera movement, modify backgrounds frame by frame, or replace actors, choose the corresponding editing or animation task instead of treating video-reference generation as deterministic editing.
Reference generation is suitable for continuing a character's identity, but character details, complex occlusions, and multi-person interactions should not be regarded as absolutely locked. It is recommended to use material with clear characters and easily distinguishable identities, check clothing, faces, and motion continuity segment by segment, and retain a post-production correction step for brand elements that must be rendered precisely.
Frequently Asked Questions
What is the difference between wan2.6-r2v and regular image-to-video?
It uses the character identity in a reference video as the basis for creation, applying the person's appearance to a new scene; image-to-video uses a static image as the generation starting point. If you already have a video character and want to continue telling their story, r2v is a better fit; if you mainly want to animate a single image, consider wan2.6-i2v.
How long can the generated videos be, and at what resolution?
The native output specifications of wan2.6-r2v are 2–10 seconds, with support for 720P and 1080P, and videos are 30 fps MP4 files. When designing scripts, arrange actions around short shots, and do not treat the 15-second limit of wan2.6-t2v and wan2.6-i2v as the output range for this model.
How do I submit a reference video and obtain the generated result?
Submit model=wan2.6-r2v to POST /wan/videos, provide the reference video URL with reference_video_urls, and use prompt to describe the new scene. Set async=true to first obtain a task_id, then query the final status through /wan/tasks; after success, obtain the video URL for playback or download.
How should prompts be written for character transfer?
First explain the identity of the reference character in the new clip, then describe the environment, action sequence, and shot intent. For multi-character tasks, distinguish who does what, and avoid using only abstract terms such as “cinematic feel.” Keep the plot focused on one clear event, and check whether the generated character appearance and actions meet the creative goal.
Does it support audio, and is it equivalent to voice-over editing?
wan2.6-r2v has native audio synchronization capabilities and can be used for short narratives with sound, but this does not make it an independent voice-over editing, precise lip-sync correction, or voice cloning tool. When creating, describe audio requirements together with scene events, and check the synchronization effect after completion; when precise dialogue or audio processing is needed, arrange post-production afterward.