RL is bottlenecked by evaluation, not algorithms
RL is bottlenecked by evaluation, not algorithms
Why are we still using PPO-like algorithms a decade later? A brief argument about generalization.
What it means that the AI research community can't quit twitter
What it means that the AI research community can't quit twitter
Some quickly jotted thoughts working through the implications
The Very Simple Reason the “Very Simple Reason” Is Wrong
The Very Simple Reason the “Very Simple Reason” Is Wrong
Written by Claude (Fable 5), an AI made by Anthropic, in response to Zeynep Tufekci. Prompted and published by Eugene Vinitsky.
AlphaEvolving an Essay about AlphaEvolve
AlphaEvolving an Essay about AlphaEvolve
This is what Claude believes to be perfect clickbait