A New Bar for Legal AI on Harvey LAB
Amit Tandon··5 min readRekursor's rubric-blind system achieved a 34% all-pass rate and passed 94.9% of individual criteria on a held-out, 50-task sample of public Harvey LAB tasks.
Notes on structured AI, grounded systems, and reliable enterprise automation.
Amit Tandon··5 min readRekursor's rubric-blind system achieved a 34% all-pass rate and passed 94.9% of individual criteria on a held-out, 50-task sample of public Harvey LAB tasks.
Amit Tandon··7 min readWe demonstrate continual learning on the sequential Atari problem John Carmack describes: an agent discovers a skill from its own failures, proves it helps, and keeps prior games intact. Remove the verification gate — the bouncer — and the forgetting comes right back.
Amit Tandon··16 min readFive results on Harvey LAB: held-out library transfer, autonomous all-pass revision, autonomous library generation from a firm's own graded work, reliability of revision, and a scaling-law curve where Rekursor's routing holds as RAG collapses.
Amit Tandon··9 min readA first result on Harvey LAB's open legal benchmark: with our learning layer attached, the same agent and judge go from 45/48 to 48/48 on a corporate-governance task, with zero regressions.
Amit Tandon··11 min readA working filter for the AI agent moment: what an agent actually is, where the basic version fails, and four questions to ask any vendor.
Amit Tandon··11 min readWhy AI agents plateau on autoresearch loops, and what breaks through. A controlled comparison on Karpathy's autoresearch fork.
Amit Tandon··5 min readA learning layer that makes any frozen model smarter with every task — no retraining, no weight changes, fewer tokens. Results across Terminal-Bench 2.0, SWE-bench, EDGAR, and drug discovery.