Agentic RL
Gradient descent trains a model on data somebody wrote; reinforcement learning trains it on its own attempts. Sample a group of answers, score them, and let the group's average — not a second network — decide which ones deserve more probability.