Choose a verifier, a task, and a base model. Download a ready-to-run training loop.

Reinforcement: Build Your Own Reward Model

Input

You get

WaitingSelect a training signal. That chooses how success is scored.
Task and base model specialize the loop. The signal is not itself a trained reward model.
Build
0 / 3
Reinforcement: Build Your Own Reward Model