Hi Maarten, incredible guide! The decisions made throughout this paper are genuinely inspiring. I wanted to ask more details on switching between causal attention to bidirectional attention for the Gemma 4 model as you had mentioned "more on this later" earlier. Would love to know how you differently you expect it to perform if it was trained bidirectionally from scratch as well.
When I saw that DiffusionGemma was released the first thing I googled was "Visual Guide to DiffusionGemma" and you didn't disappoint 🤭
Hi Maarten, incredible guide! The decisions made throughout this paper are genuinely inspiring. I wanted to ask more details on switching between causal attention to bidirectional attention for the Gemma 4 model as you had mentioned "more on this later" earlier. Would love to know how you differently you expect it to perform if it was trained bidirectionally from scratch as well.
The last two paragraphs in 'Multi-Canvas Sampling' seem repetitive, especially with the identical opening sentence. Was this intentional?
Could you please hint why a small FFNN is added for self-conditioning?
You can find an answer in this paper ! https://arxiv.org/pdf/2510.19304