Inputs
Note: The
length parameter is processed in steps of 4 frames, and the node automatically applies temporal scaling when building the latent space. When ref_image is provided, only its first frame is encoded (resized to width x height) and attached to the conditioning as reference latents. When control_video is provided, it is trimmed to length frames, resized, encoded, and placed into the concat latent used by the conditioning. The concatenated latent is duplicated along the channel dimension and its channel layout depends on the VAE’s latent channel count (48 channels uses the Wan 2.2 format, otherwise the Wan 2.1 format). The start_image parameter is referenced in the execution logic but is not exposed in the node’s input schema, so it cannot be set from the node interface.
Outputs
This documentation was AI-generated. If you find any errors or have suggestions for improvement, please feel free to contribute! Edit on GitHub
Source fingerprint (SHA-256):
731b848f15c13ddc662f19230acb55d195f934bad7d9ae516a288e0ed8f8d899