t3rl: An Open Workspace for RL Research
I’m building t3rl as a workspace for programming and reinforcement learning research. The goal is to connect a hypothesis to the code, training run, evaluation, and evidence needed to judge it, covering both control tasks and language-model post-training.
The project is an independent open-source fork of T3 Code, retaining its core coding workspace and adding a project-scoped RL Lab. You can build applications, fix bugs, and refactor code with Codex, Claude Code, Cursor, Grok, or OpenCode using your own provider accounts. Projects and conversation threads organize work alongside an integrated terminal, file browser, app previews, and external-editor integration.
The coding foundation includes Git branches, isolated worktrees, diffs, commits, and workspace checkpoint restore. GitHub integration supports cloning and publishing repositories, pushing changes, and creating and reviewing pull requests; GitLab, Bitbucket, and Azure DevOps integrations are also available. Permission modes, customizable shortcuts, usage tracking, and remote access through web, desktop, and mobile carry over from T3 Code.
For experiments, a Node server supervises replaceable Python workers through a versioned NDJSON protocol. Web and desktop clients display authoritative run state; disconnecting a client leaves training running on the server. High-volume metrics stay outside the agent conversation.
Stable-Baselines3 handles control experiments, while TRL and Axolotl adapters provide LLM post-training paths. Each run records its resolved configuration, seeds, source revision, environment, metrics, and artifacts. The lab supports comparisons across seeds, checkpoint lineage, evaluation inspection, and evidence exports with a standalone verifier. Agents can inspect this evidence and help prepare the next experiment; the current Autoresearch workflow proposes one iteration for review.
t3rl is still alpha. Recorded checks exercise real training, checkpoint resume, and adapter loading, but use bounded fixtures on one host. They establish parts of the workflow; useful model improvement and independent reproduction remain to be demonstrated.
The next steps are concrete:
- Complete a preregistered investigation with a baseline, candidate, ablation, independent holdout, and at least three training seeds, then reproduce it from a clean environment.
- Validate GPU study scheduling, web and desktop workflows, remote reconnect, and measured performance budgets.
- Expand LLM training methods with RLOO and a version-scoped PPO integration as research requires, followed by distributed execution, recovery, and mobile monitoring.
The ambition is an open research workspace where an agent can propose a change and a researcher can trace the resulting conclusion back to its evidence. The repository and development roadmap track that work.