My name is Jake and I have a snake. Currently building Recension, where we're training models for genetic engineering. Previously, I've done RL at Tzafon and built evals at Endeavor. I studied computer science and math at NC State, which I left senior year to move to San Francisco. I post here occassionally to share thoughts and projects.
Research Interests
Evals. If what gets measured gets improved, then the starting point for developing any new capability is to make a benchmark for it. I believe this is one of the highest-leverage ways to drive progress.
I also maintain a capabilities index composed of benchmarks which I think are valuable.
Harnesses. Models are only useful if you can apply them. There's a lot of low-hanging fruit here.
A few of my favorites:
- RuneBench. AI writes code to play RuneScape
- Remote Labor Index. CUA harness and benchmark for measuring progress towards full automation of online freelance work. Still very unsaturated, no model gets above 16%
- Kosmos. This one helped with research for the genetic engineering project below
Reinforcement learning. A combination of the above. Make a harness for a task, an evaluation to measure the output quality, and now you have an RL environment.
Data quality. If your data is flawed, your model will be too. Garbage in, garbage out. LOOK AT YOUR DATA. Not enough people do this.
Genetic engineering. One of the projects we are working on at Recension is creating a real-life Black Lotus. If this interests you, reach out here.
MAY UPDATE: the plants survived the transformation and the pigment doesn't appear to be leaking into the wrong tissues. The next test will be if they flower and show color.
Gallery




