🎨 FWKV-Vision — Text to Image

A ~40M-parameter rectified-flow DiT where self-attention is replaced by FWKV (bidirectional decayed-accumulator time-mixing, computed with an exact O(log T) parallel scan), with cross-attention to a frozen CLIP text encoder and denoising over a frozen sd-vae-ft-mse latent space. Model: FWKV/FWKV-Vision · Running on CPU

1 100
0 15
Examples