🎨 FWKV-Vision — Text to Image
A ~40M-parameter rectified-flow DiT where self-attention is replaced by
FWKV (bidirectional decayed-accumulator time-mixing, computed with
an exact O(log T) parallel scan), with cross-attention to a frozen CLIP
text encoder and denoising over a frozen sd-vae-ft-mse latent space.
Model: FWKV/FWKV-Image · Running on CPU
1 100
0 15
Examples