all science is math · all models are wrong
Jin Ziqi
金子琪
2nd-year Ph.D. student at Nanyang Technological University, advised by Prof. Aixin Sun. I build language models that reason, and probe why they work.
- NOW
- ByteDance
- PREVIOUSLY
- SEA AI Lab · MiroMind
- RESEARCH
- Efficient LLM · Optimisation · Diffusion LLM
01 / FEATURED
Featured
- research
MARS: Several Tokens per Forward Pass, Without Giving Up Autoregression
One SFT checkpoint, two modes: full-quality AR at τ=1.0, or 1.8 tokens per forward pass for 1.56–2.16× speedup, with an exact, free fallback to plain AR.
02 / WRITING
Blogs
- research
How Much Do Diffusion LLMs Memorize? (vs Autoregressive)
Same size, same memory ceiling, but the diffusion model needs ~10x the training to reach it, and eventually just can't.
- dog
Living With a Dog: Treat It Like an RL Agent
A dog doesn't get reasons, only consequences. Treat it like a reinforcement-learning agent and almost every training problem collapses into one: what exactly are you rewarding, and what are you punishing?