Skip to main content
TomePatchModel applies Token Merging (ToMe) to a diffusion model to reduce computational cost during inference. It works by merging similar tokens inside the model’s attention mechanism, so the model processes fewer tokens while keeping output quality largely intact.

Inputs

Note: If the number of tokens in an attention block is small enough that no downsampling is needed, the merging functions are replaced with no-ops, and the model runs unchanged for that block.

Outputs

This documentation was AI-generated. If you find any errors or have suggestions for improvement, please feel free to contribute! Edit on GitHub

Source fingerprint (SHA-256): 1202c0df17f357440cd156fa0920f70c18a318e32c41dc04cecff11613f0072f