Reinforcement learning you can
watch, play & teach.

A browser-based workbench where you train agents, watch them learn in real time, and then take the controls yourself — across 100+ environments, from CartPole to MuJoCo, board games and 3D Doom. Free and open.

Free & open · AGPL-3.0 100+ environments · 9 algorithms Actively developed · early access

Watch them learn.

Every agent below was trained inside RL Lab — no notebooks, no local setup. Pick an environment, pick an algorithm, hit train, and watch the reward curve climb.

BreakoutBreakout · DQN
Lunar LanderLunar Lander · Neuroevolution
Car RacingCar Racing · PPO
CheckersCheckers · AlphaZero
Bipedal WalkerBipedal Walker · PPO
DoomDoom: Defend the Center · PPO
HumanoidHumanoid · SAC
PursuitPursuit · multi-agent PPO
PongPong · PPO

The whole loop, in one place.

The free open-source stack (Gymnasium, Stable-Baselines3, CleanRL) is libraries and command lines. RL Lab is the loop — train, watch, play, analyse — behind one interface a beginner can drive.

Environment picker

Train. 100+ environments, 9 algorithms.

PPO, SAC, TD3, DQN, A2C, a from-scratch neuroevolution and an in-repo AlphaZero — across classic control, MuJoCo physics, Atari, board games, cooperative multi-agent and 3D Doom.

Live training — reward curve climbing, Q-table filling in real time

Watch. Live, not log files.

The reward curve and a decoupled agent preview stream as it trains — plus Q-table heatmaps, telemetry and skill meters. You see learning happen, you don't read about it after.

Play against your agent

Play. Take the controls.

Play 102 of the games yourself and see your score graded from Child to Superhuman on the same scale as your agent's — and play the six board games head-to-head against it, move for move. Nothing in the free RL stack lets a human step into the loop like this.

Data Lab comparison

Analyse. Research-grade Data Lab.

Compare runs the honest way — IQM, stratified-bootstrap confidence intervals and performance profiles (the rliable method), with one-click publication-ready exports.

Built for teaching.

  • Learn by doingStudents train real agents in minutes — not slides, not weeks of environment setup.
  • Every knob explained, in placeEach hyperparameter carries a plain-language, per-environment popup with a recommended value (EN & CZ).
  • Free & open, no lock-inAGPL-3.0. No per-student licences to buy, no accounts to manage — it's yours to run.
  • Runs on your students' machinesLocal-first: the training uses their own compute, so there's no cost or cloud account for you.
In-context hyperparameter explanation — the Learning Rate popup

Teaching a course with reinforcement learning? Try RL Lab with your students and tell me what breaks — that feedback shapes what gets built next.

Use it in your course →

Open, and here to stay.

RL Lab is built solo and kept free for learners. It's released under AGPL-3.0 — open for everyone, sustained by grants, partnerships and people who find it useful.

Partnerships & grants

Educators, labs, funders and anyone who wants to help this become a genuinely great open teaching tool — let's talk.

Get in touch →

Support the project

If RL Lab helps you or your students, a sponsorship keeps the lights on and the roadmap moving.

Sponsor on GitHub