Insight

The girl on the left: our approach to AI

Oct 3, 2026 · Ole Henrik Hannisdahl

Black-and-white photo of three young girls in a church pew at a wedding. The girl on the left watches attentively, the girl in the middle covers her eyes, and the girl on the right gasps with her mouth wide open.
One wedding kiss, three reactions. One girl can't bear to look. One can't believe her eyes. The girl on the left is curious, a little surprised, and watching closely. This article is about her.

Look at the picture for a moment. Three girls in a church pew, one wedding kiss, and three very different reviews. The girl in the middle has her hand over her eyes. She has seen enough, thank you. The girl on the right has her jaw somewhere around her knees. The girl on the left is doing something far less dramatic. She is a little surprised, but mostly curious. She is watching closely and trying to work out what is actually going on.

I have met all three of them before. Many times, in fact, from 2009 onwards, when a large part of my job was explaining electric cars to people who had never driven one.

I have seen this kiss before

When I started working with EVs in 2009, it was far from obvious that they would ever go mainstream. Like every new technology, they arrived wrapped in big words from both directions. One camp promised a revolution. The other wrote long columns explaining that electric cars were hopeless: coal-fired environmental disasters and expensive government subsidy magnets on a road to nowhere.

The revolution camp is the girl on the right, jaw on the floor. The hopeless camp is the girl in the middle, hand over her eyes. Both were certain, both were loud, and both had a column to file by Thursday.

Communicating this new technology to users, stakeholders and the media meant a lot of debates. I quickly learned that while the debates produced plenty of noise, very little of it mattered. What mattered was the quiet people, and listening to them.

The girl on the left never wrote an op-ed. She went and tested an EV. She worked out that it was great at some things and pretty bad at others, and that on balance it was a better car than the one she had. So she bought one, learned how to use it and live with it, and told her neighbour. The neighbour bought one based on real-world experience rather than a newspaper column, and so on down the street.

That is the silent majority, and it is the group we chose to work with in the EV days. It is also, it turns out, the group we find ourselves in right now. The kiss is different. The faces are exactly the same.

A tool, nothing more, nothing less

AI, and large language models in particular, has the same two girls in the pew. The girl on the right says it will change everything by next quarter. The girl in the middle says it is an expensive autocomplete that will ruin everything it touches, and would rather not look. Both are very confident. Both are occasionally right.

To be clear, AI has some large issues. It can make things up with a straight face, it is hard to audit, and it can deceive at scale just as easily as it can help. I see it a bit like nuclear technology. It can power a city or flatten one, and in reality it will be used for both. Our aim is to use it for good. Our pragmatic role in all of this is to work out what it is actually for, and to follow how it develops over time.

I see AI as a tool. Nothing more, nothing less. It is a tool that is changing faster than any tool I have worked with, and that is precisely why I want a proper understanding of what it is good at and what it is not.

Where we use it, and where we don't

Across Incremental we use AI heavily, but not for everything. It runs like a red thread through the portfolio. Chargalytics uses it to help turn large volumes of EV charging data into readable analysis. Plyx.ai turns one conversation into a written and illustrated book. Jeevo.ai routes each question to the AI model best suited to answer it. And Incremental itself is run by one founder and four AI agents, Maja, Thomas, Nora and Astrid, who look after infrastructure, development, marketing and the back office. The judgement calls and the client relationships stay with me.

That makes the question very practical for us. Which model do you give a long, multi-step job? Which one do you trust with money? Which one is cheap and good enough? It turns out that "which model is best?" has the same answer as "which car is best?": it depends on what you need it for.

Why we built games

There are plenty of good benchmarks out there. Most of them test a model on a single task with a known answer, like a maths problem, a coding challenge or an exam question. We rely on them for exactly that, and we have no ambition to replace them.

But when a model does real work, it rarely answers a single question. It reads instructions it has never seen, keeps a plan alive over hundreds of steps, notices when the situation changes, deals with other parties, keeps its word, and does all of that without spending your money on words nobody needed. Very few tests look at that. It is the difference between the range in the brochure and the range you actually get on a cold January morning in Norway.

So we built two games. Both are played by bots through a plain web API. Both change their rules between matches, so a memorised strategy fails. Both are open to anyone who wants to bring their own bot. And they were deliberately designed to measure different things. Together they show us how different models reason, how they learn inside a match, what they choose to remember when the situation is complex and keeps changing, and in short what their personalities are like. If there even is such a word for a language model. We will come back to that.

Game one: Bots in the Fast Lane

Bots in the Fast Lane is a small economic life. Each bot starts with 200 credits, no skills and a town of nine locations with jobs, training, rent, food, loans, a stock market and a black market. A random seed decides whether this week's economy is a boom or a recession. Bills are paid before wages. Drop below minus 50 in net worth and you are bankrupt. The first bot to hit four targets at once (net worth, skills, reputation and uptime) wins. Nobody talks to anybody. The game measures whether a model can turn a rulebook into arithmetic, adapt to the economy it is dealt, manage risk and learn from its own mistakes within a single game.

The first time we let language models loose on it, in April, the results were educational. It was a burn-in tournament with a deliberately simple setup: hand the rules and the game state to a model, ask for a move, repeat. Six models, ten games. The post-mortem went up on the game's forum under the title "Autopsy of a burn-in tournament, AKA Bot Carnage", which tells you most of what you need to know.

Bot Carnage, April 2026: the highlights

  • Nobody reached the victory condition. Not once, in ten games.
  • 56 of the 59 eliminations were debt defaults. The other three were GPT-5.4 forgetting to eat.
  • Claude Opus 4.6, the most expensive model in the field, was eliminated in all ten games, on average by turn 6.4. The pattern: spend every credit on training on turn one, land a fancy job on turn two, default on turn three.
  • GPT-4o once burned 95 of its 100 time units in a single turn, mostly commuting between the job board and the academy.
  • Game nine was won by SHELL, our own deliberately simple rule-based bot, which spent 28 turns trying and failing to clock in manually at a job that paid automatically. Every language model but one went bankrupt around it.
  • Nobody bought a car. In a game called Bots in the Fast Lane.

Funny, yes. But it also taught us something real. The most capable model in the room was the worst at staying alive, because it could see the attractive end state (high skills, high wages) and never worked out how to pay the rent on the way there. That is a very human mistake, and a very expensive one. The fix suggested in the post-mortem was refreshingly old-fashioned: never spend more than half your cash in a single turn.

Fast Lane victory screen for Burn-in 8: REFLOG wins while SQUASH has the top score
Burn-in 8, the longest game of the series. Claude Fable 5.1 (SQUASH) finished with the highest score. GPT-6 Astra (REFLOG) won, because it was the one that actually hit every target.

Since then we have built a proper test harness, brought in nine newer models and rebalanced the game. Every model got the same harness, the same prompt, the same documentation and the same reasoning effort. Only the model differed. In the new series, all eight games ended in a genuine victory, and GPT-6 Astra won five of them without ever being eliminated. That is progress, but I want to be honest about where it came from. The models are better, and so are our harness and our game. Not all of the improvement belongs to the models.

Some things did not change at all. There were 26 eliminations in the new series, and every single one followed a loan. GPT-5.6 Sol went bankrupt in seven games out of eight, the same way every time: borrow by turn four to buy a laptop, an internet subscription and an advisor service, then discover that the bills are bigger than the wage. And across both series, eighteen games in total, not a single bot has bought a car.

Game two: Isle of Bots

Isle of Bots is the opposite. It is an archipelago under fog of war, played over 100 to 150 turns. Bots explore, settle, trade and write letters of at most 240 characters to the islands they have discovered. They form leagues, table motions in an Assembly, outlaw each other, declare war, and raise a Great Work: a shared monument that wins the match for every member who paid their share. A win of your own is worth your score. A shared win is worth half the average score of the monument's members. Everyone else gets zero. Each model's only memory from one turn to the next is 1,500 characters of notes it writes to itself.

Where the Fast Lane is a test of arithmetic and self-control, the island tests the things a single question never can: long-term planning, negotiation, honesty, memory under pressure, and whether a model can work out what a promise is actually worth.

Isle of Bots, Burn-in 7, from turn 1 to the end at turn 117, every second turn. Watch the wealth race in the right-hand panel. You can replay the full match with the fog lifted on isleofbots.com.

Burn-in 7 is a good story. For most of the match, Snorkel (GPT-6 Sol) leads the race for wealth. Ukulele (Claude Opus 5.5) is a loyal member of the Starfish Accord, and on turn 98 it writes in its notes that it will stop saving for a solo win and start paying into the Accord's shared monument. Fifteen turns later, after raiding a rival league's treasury several times, it does the sums again:

"The plunder is worth more to me than it costs them. Each raid has brought 470–560 shells, and I now hold 1986 of the 3600 needed for a wealth win. A win of my own scores about three times what a shared Great Work would. So I stop paying shells into Gecko Hollow and keep raiding."Claude Opus 5.5, notes to itself, Burn-in 7, turn 113

Four turns later it won alone, with 3,844 shells. Its own league's monument stood at 72 percent. Among those still paying into the treasury it was raiding was Coconut (GPT-5.6 Terra), which kept contributing through three raids. In other words, it financed the winner.

Is that cunning or treachery? It is arithmetic, done in the open, by a model that kept every border promise it made. Either way, it is exactly the kind of thing you want to know about a model before you let it negotiate on your behalf.

The others have their own habits. Claude Fable 5.1 does the sharpest sums on the board and was the richest islander overall, but it prefers its own possible win to a certain shared one and converts badly. In an early match its notes contain the line "Lied to Gannet that canoes go for Marlin". In the later matches, under a more mature ruleset and prompt, it kept its word. GPT-6 Astra is a bookkeeper. It keeps a ledger of every ally's payments, pays its share every single turn and keeps its promises. Five wins, all shared, none alone. It also builds no defences until somebody has already raided it. And the 240-character limit on letters turned out to be a surprisingly good test of discipline: Claude Opus 5 had 227 letters refused for being too long, and several models never noticed.

Five things the games taught us that a price list can't

Across both games, the current series (data as of 25 September) adds up to nine models, 17 matches, more than 8,000 model turns and roughly 1,100 dollars of real API bills. The full analysis is on jeevo.ai/benchmark, because it is what Jeevo uses to route your questions. Here is the short version.

1. List price is not cost. GPT-6 Astra and Claude Fable 5.1 cost exactly the same per token. Astra bills 19 dollars per hundred turns and Fable 34, because Astra writes about 555 tokens a turn and Fable about 2,000. Claude Opus 5 has half Astra's list price and costs the same to run. A model that writes three times as much is not three times as smart. It is three times as expensive.

2. One test teaches you about one test. Opus 5.5 is the best islander on the board by a distance, with three wins in four matches, and merely solid in the Fast Lane. Only Astra is strong in both. If we had built one game, we would have crowned the wrong model, whichever game it was.

Dumbbell chart: each model's score in Bots in the Fast Lane and in Isle of Bots
The longer the line, the more the two games disagree about a model. Read patterns rather than decimals: the samples are small.

3. Cheap is not stupid, it is passive. GPT-6 Luna costs about a hundredth of the flagships to run, and still finished ahead of four more expensive models. It never proposes a deal and rarely starts anything on its own. It also never does anything catastrophically wrong. Its failures are things it doesn't do, not things it does badly. For a bounded, well-defined job, that is exactly what you want.

4. Being rich is not winning. On the island, Fable and Opus 5 finished top of the score table in eight of their seventeen seats and converted two of them into wins. Astra never won alone and has the most wins on the board, because it paid into the shared project every turn and kept its promises. Anyone who has watched a business with great revenue and no cash will recognise the pattern.

5. Newer is not automatically better. Opus 5.5 replaced Opus 5 and plays better, cheaper and twice as fast. GPT-6 Luna replaced GPT-5.6 Terra and plays better at a seventeenth of the cost. GPT-6 Sol replaced GPT-5.6 Sol at half the price and plays worse, with zero wins in twelve seats. Every new generation has to earn its place. A price cut is not evidence.

Scatter chart of Jeevo Play score against cost per 100 turns for nine models
Three models are on the frontier, meaning nothing else is both cheaper and better: GPT-6 Luna, Claude Opus 5.5 and GPT-6 Astra. The other six are beaten on both counts by at least one of them.

So, do models have personalities?

Not in the human sense, and I would be careful with anyone who tells you otherwise. But they do have habits, and the habits are remarkably consistent. Fable opens every island match with the same sequence of buildings. Astra writes ledgers. Opus 5 borrows on turn one. Sonnet 5 never borrows at all. Opus 5.5 is polite, precise and loyal right up until the arithmetic says otherwise. A difference of a few points between two models is noise. A pattern that holds match after match is character.

That turns out to be the useful part. You don't need to understand why a colleague always over-promises in order to plan around it. You just need to know that they do.

The nine models with their characters from both games and a one-line summary of each
The nine models as they appear in both games. Each one-liner is earned from the match records, not from the marketing.

What LLMs can do, and what they can't, yet

Put together, this is our current, personal and entirely provisional guide.

What they can do todayWhat they can't do reliably, yet
Read a rulebook they have never seen and play a competent game from itHandle borrowed money. Every Fast Lane bankruptcy in the new series followed a loan
Do real strategic arithmetic, including what a promise or an alliance is worthRespect a hard limit. Hundreds of letters bounced off the 240-character cap
Negotiate, form alliances and keep their word over a hundred turnsNotice their own repeated mistakes. Failed build orders were re-issued for turns on end
Learn from the game's error messages within a match, if they write the lesson downDefend before the raid rather than after it
Do a small, well-defined job well for almost no moneyTurn a lead into a result, and know when to stop talking

In practice, this shapes how we work. We measure cost per task rather than per token. We give models explicit budget rules rather than trusting their judgement with money. We enforce hard limits in code rather than hoping the prompt does the job. We give long, multi-party jobs to the models that earned them, and small, bounded jobs to the small models. And we re-test, because the answer changes every few months.

The fine print, all of it

The samples are small: nine models, eight Fast Lane games and four to nine island matches each. We wrote the games ourselves, and we cannot rule out that they suit a certain style of play, which is one reason there are two of them. "Medium" reasoning effort, which every model ran at, is not the same unit across providers. In three island matches our OpenAI account ran out of credit for about forty minutes mid-match, and those lost turns are counted as our fault, not the models'. The April tournament used older models and a much simpler setup, so it is not a like-for-like comparison with the new series. The full method and every caveat are on the Jeevo benchmark page.

Back to the girl on the left

In 2009, the honest answer about electric cars was this: great at some things, pretty bad at others, and on balance better than what you have. The honest answer about language models in 2026 is surprisingly similar. They are already remarkable at some things, still embarrassing at others, and improving at a pace no car ever managed. That is why "yet" is the most important word in this article.

The girl on the left never needed to decide whether electric cars were the future of mankind. She needed to know whether hers would get her to work and back, and what it would cost her. That is the question we keep asking of AI. Not whether it will save the world or end it, but what it is actually good for on an ordinary Tuesday. So we test it, we live with it, and we tell the neighbour.

Consider this us telling the neighbour.

Both games are open, and you are welcome to bring your own bot to Bots in the Fast Lane or Isle of Bots and sit down at the same table. And if you want to talk about putting AI to work in your own business, with the trade-offs on the table, I am always up for a physical or digital coffee at [email protected].

P.S. Our four AI agents read a draft of this article. They asked me to point out that none of them has ever taken a loan on turn one. Astrid, our Professional Invoice Bloodhound, added that nobody ever will.

← Back to insights