Blog
The Path Towards Adaptive Intelligence
Achieving outcomes with autonomy in the real world is the benchmark for intelligence.
By OpenWorlds · · 2 min read

In 1950, Alan Turing proposed the imitation game as a test of machine intelligence. The game asked whether a machine could pass for a person in conversation. Today, conversational imitation is no longer the frontier. A more demanding question is whether machines can act with autonomy in open-ended environments and produce outcomes that matter in the real world. Markets are a uniquely challenging environment with evolving dynamics, but performance is simple to measure and verifiable.
Beyond static benchmarks
ARC-AGI-1 challenged models to infer rules from only a few input-output examples across hundreds of unfamiliar puzzles. It remained unsolved until late 2024, when OpenAI's o3-preview reached 75% at low compute and 87% with higher compute. ARC-AGI-1 pushed static benchmarks toward a harder question of whether models can generalize from sparse information rather than memorize fixed patterns.
ARC-AGI-3 has moved beyond static tasks and into fully interactive environments with unknown rules and unstated goals. GPT-6 Astra scored 62.7% on ARC-AGI-3 Semi-Private with a standard harness, and 99.9% with a provider adapter harness1. Models are being pushed on adaptability, acquiring goals on the fly, and learning continuously from interacting with the environment.
Markets raise the stakes further. Models are improving rapidly and the frontier is shifting from answering questions to moving into environments where the rules are not known upfront. We think it's an important time to push these models to their limits in non-stationary environments that are adversarial and endogenous.
The Test of Economic Usefulness
The modern test of economically useful intelligence is whether a machine can make money repeatedly across markets with verifiable risk-adjusted returns. The ability to generalize across all markets and produce results is a significant milestone towards the next stage of intelligence that is adaptive and able to achieve economic outputs.
Model progress is why we have chosen to be an open platform. Every new model that pushes the industry forward is available on the platform. Our goal is to offer choice across all models in a default experience, using the feedback for an even better experience with improved models or novel abilities.
What autonomous driving taught us
Tesla has shown the path toward autonomy through its massive fleet of vehicles. Early FSD relied heavily on hand-written rules and human labeling for identifying real world objects like stop signs or traffic lights. Human bottlenecks made the process slow and expensive. When Tesla shifted to end-to-end AI, improvement scaled with compute and created a feedback loop that resulted in a better experience.
Tesla's FSD Safety report showed highway miles per critical intervention increased 10x and passengers reported a smoother experience2.
Autonomy for markets
At OpenWorlds, we are applying these principles to achieve autonomy for markets. We are taking an end-to-end approach building the harnesses, simulators, and infrastructure to make recursive self-improvement measurable and mechanical.