Latent-GRPO: Reinforcement Learning in Continuous Thought Space
When you sit down to solve a complex puzzle or plan three moves ahead in chess, do you narrate every synaptic firing to yourself in full, grammatically correct English sentences? Of course not. Human
Sep 23, 202610 min read4
