Skip to main content
Creates video latents for the Cosmos Predict2 image-to-video workflow. The node can produce an empty video latent of a given size and length, or insert encoded start and/or end images into the sequence so those frames are kept during generation. Any supplied image is resized to the requested width and height and encoded with the provided VAE before being placed at the beginning and/or end of the latent sequence.

Inputs

Note: When neither start_image nor end_image is provided, the node simply returns an empty latent of the requested size and length. When one or both images are provided, they are resized to width and height, encoded with the vae, and placed at the beginning and/or end of the latent sequence. The corresponding regions are marked in the noise mask so they are preserved during generation. The encoded latents are converted with the Wan 2.1 latent format, and the resulting latent and mask are repeated batch_size times.

Outputs

This documentation was AI-generated. If you find any errors or have suggestions for improvement, please feel free to contribute! Edit on GitHub

Source fingerprint (SHA-256): 842bd2b8cda438e7b938439d4eba280478939e3302dc1846d52595d40082ff05