<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Evals on cornball.ai</title><link>https://cornball.ai/tags/evals/</link><description>Recent content in Evals on cornball.ai</description><generator>Hugo -- gohugo.io</generator><language>en</language><lastBuildDate>Tue, 15 Sep 2026 11:09:13 -0500</lastBuildDate><atom:link href="https://cornball.ai/tags/evals/index.xml" rel="self" type="application/rss+xml"/><item><title>98.6% on ARC-AGI-3 in R</title><link>https://cornball.ai/posts/corteza-arc-agi-3/</link><pubDate>Mon, 14 Sep 2026 00:00:00 +0000</pubDate><guid>https://cornball.ai/posts/corteza-arc-agi-3/</guid><description>I&amp;rsquo;m pleased to announce that corteza has scored 98.58% on the ARC-AGI-3 challenge with Claude Opus 5. corteza is an R-native agent harness built around a persistent R session: functions, data, and intermediate work survive across model calls, and even across context compaction.
wa30, a Sokoban-style carry game and one of the 23 games corteza finished at a score of 100. The rules are unknown by design; the model has to work them out.</description></item></channel></rss>