Trading View Ticker Widget

The Price Of AGI Has No Definition

NVIDIA's CEO says AGI has arrived, but the benchmarks, contracts, and definitions meant to prove it tell a messier story.

Welcome to Memorandum Deep Dives. In this series, we go beyond the headlines to examine the decisions shaping our digital future. 🗞️

This week, artificial general intelligence stopped being a research milestone and became a line item. NVIDIA's Jensen Huang posted on X that AGI had arrived, crediting OpenAI's newest model and the mountain of NVIDIA hardware it trained on. OpenAI's own materials said less, and its president chose his words far more carefully, even as he shared Huang's post with his own followers.

The claim landed on top of a benchmark result that is, by any measure, extraordinary. A test built specifically to resist AI progress, one that humans could ace and machines could barely touch as recently as March, produced a score above 99% for OpenAI's newest system. Run through a different piece of software, that same model scored 37 points lower on the exact same test.

Behind that gap sits a much bigger one: a term now embedded in investment contracts, infrastructure loans, and corporate rewrites worth hundreds of billions of dollars, with no agreed test for what it actually means. Below, we trace what changed on September 6, what didn't, and why the industry keeps finding new places to avoid answering the one question that started this whole race in the first place.

Newsletters don’t belong in your inbox.

Your inbox was built for email, not focused reading.

Meco gives newsletters a dedicated, distraction-free home away from the noise. Connect Gmail or Outlook in seconds, organize subscriptions with smart filters and groups, bookmark your favorites, and read everything in a clean, scrollable feed.

Join 50,000+ readers who are decluttering their inbox and getting more from the newsletters they love.

Build and design your website on Framer - Now with Agents.

Framer is a pro website builder trusted by companies like Miro and Perplexity that helps creators, teams, and businesses ship production-ready sites faster than ever.

With AI agents built directly into the canvas, teams can design pages, manage CMS content, write copy, add SEO, and audit for issues — all without leaving the tool where the real site lives. Agents bring speed and scale; you bring taste, judgment, and control.

*This is sponsored content. See our partnership options here.

AGI: the word that cost $35B

On August 15, 1977, a radio telescope in Ohio known as the Big Ear recorded 72 seconds of a narrow signal from the direction of Sagittarius. Jerry Ehman, a volunteer working through the printout days later, circled the string of characters and wrote 'Wow!' beside it in red pen. The signal, however, was not heard again, and because it wasn't, it did not meet the rule of astronomy that a detection is not a detection until a second, independent instrument sees the same thing. Which means that even after nearly 50 years, the signal is viewed as an anomaly, rather than a finding, since nobody has ever managed to run that second check.

The story of the Wow! signal is one of the most well-known examples of humanity's search for an intelligent life form capable of sending signals across the vastness of space. And the decades of anomaly may explain why humanity has spent the better part of the last century focusing harder on building intelligence, rather than finding it.

Over the past five years, that search for, or rather, the push to build intelligence has reached a fever pitch. But while companies were dumping trillions of dollars and a large chunk of human ability into building intelligence, they forgot to answer one question: how does one define intelligence? And more importantly, when will humanity be able to say with confidence that it has built a truly intelligent system, which many now confuse with artificial general intelligence?

Artificial general intelligence, or AGI, is meant to describe a machine that can do most of what a person can do intellectually, rather than the one narrow task it was built for. But no lab, no benchmark, and no regulator has ever settled the point at which a system crosses that line, which is why the term still has no threshold behind it. What it does have is a price, because AGI now sits within investment contracts and beneath the assumptions holding up hundreds of billions of dollars in construction. And that is the strange part of the moment we are in, since the word is doing financial work that the science has never agreed to back.

The declaration came from the supplier

On September 6, 2026, Jensen Huang posted on X that AGI had arrived, crediting OpenAI's newest model and the more than 100k NVIDIA systems it had been trained on. (NVIDIA is also an OpenAI shareholder, having put $30B into the company's March 2026 round.) Those systems are built around the graphics processors that do the arithmetic behind training a model, and NVIDIA sells more of them than anyone else in the world. Which is why the declaration is worth reading for who made it, rather than for how confident it sounds.

OpenAI itself has been far more careful, and its president, Greg Brockman, told reporters on September 3 only that the AGI era had begun, leaving the final judgment to the people using the products. So the company whose model would carry the label never asserted it in its own materials, even as its president amplified the post that did, while the company selling the hardware supplied the word. Which is reason enough to go and look at what the model actually did.

The capability is real, and somebody else measured it

The evidence that something extraordinary happened is strong, and none of it comes from anyone selling hardware. ARC-AGI-3 is a set of puzzle environments built to reward learning something new rather than recalling something already known, and it launched in March 2026 because the earlier versions had been solved to the ceiling. At launch, humans completed every environment; the best model scored 0.37%, which was the entire point of building a harder test. Five months later, OpenAI's GPT-6 Astra cleared 62.7% under controlled conditions, and it beat the human efficiency baseline for the first time. Francois Chollet, who created the test, called that a step change and says it has pulled his own forecast forward. The movement is significant precisely because it comes from a researcher who has spent years arguing that scale alone would never close this gap.

The same release moved a second capability just as sharply, and this one is easier to picture. Astra is the first model OpenAI has graded at its critical threshold for cybersecurity, meaning it can find and exploit unknown flaws in hardened systems without a person walking it through each step. A machine that can run an intrusion from beginning to end is a serious thing by any measure, though it is still a specific skill rather than a general mind. So the curve is steep, and what remains is the harder question of what a number on a test actually certifies, which happened to be answered on the day the model shipped.

Two numbers for the same model

A model does not sit down and take a test the way a person does. It is wired into a piece of software that decides how much it can remember from one step to the next, whether it is allowed to run code, and how many attempts it gets, and the field calls that wiring a 'harness'. Change the wiring, and the same model will produce a different score, which is exactly what happened here.

ARC Prize ran Astra twice. In the standard setup, the one it uses on every system so that results can be compared against each other, Astra scored 62.7%, at a cost of roughly $26k. In a second setup, OpenAI's own Provider Adapter, which preserves the model's opaque reasoning state between requests and compacts long runs so it can reuse prior work, the same model scored 99.9%, and it did so for about $19k.

OpenAI's launch page, published on September 3, 2026, reports the higher of those two numbers. That is not a false claim, because the run happened and the score is real, but it is a score produced by better wiring rather than by a better machine, and the 37 points between it and the standard run describe the scaffolding rather than the thing inside it.

Which is why ARC Prize now intends to publish the two setups separately, so that nobody has to guess which one they are looking at. It is, in that sense, the second instrument, the Wow! signal never got, and the first thing it did was disagree with the reading it was sent to check. But knowing that a score moves when the wiring changes still doesn't tell anyone what a score would have to say before AGI could be called real, and that is where the argument about definitions begins.

Hire smarter with Athyna, save up to 70% on salary costs.

Athyna connects you with top LATAM AI talent, fast!

*This is sponsored content

Capability and generality are not the same measurement

The strongest objection to calling any of this AGI comes from researchers who accept every one of those results. A 2025 preprint, which has not yet been peer-reviewed, by Dan Hendrycks, Yoshua Bengio, and 31 others starts from a simple test: could the machine do what a well-educated adult can do, across the full range of things such a person handles? So they break that range into ten areas, from reading and mathematics to memory and the ability to hold a plan in mind while working through it.

Measured that way, today's systems come out lopsided. They are excellent at the areas that reward having read almost everything, and they are poor at the ordinary business of remembering, since a model that solved a problem for you last week begins the next conversation with no knowledge that the week happened. A person who could do what these systems do in their strong areas would never be so helpless in the weak ones, and it is that unevenness, rather than any individual score, that the paper says disqualifies the label.

The paper then makes a second, more pointed argument: AGI cannot simply mean AI that earns a lot of money. Narrow technologies have always earned a lot of money, and the spreadsheet reorganized office work without understanding anything at all. And that argument lands directly on OpenAI, because the company's own working definition of AGI describes a system that outperforms humans at most economically valuable work, which is a measure of what the machine produces rather than what it understands. A definition written in terms of output is easy to claim and very hard to disprove, and those are exactly the qualities that make a word useful in a contract.

The verification clause was removed in April

Contracts are where the industry came closest to solving this problem, and also where it stopped trying. Microsoft is OpenAI's largest partner and paid for the right to use its technology inside Microsoft products, which means anything that changes what Microsoft is allowed to use is worth a great deal of money to both sides. And under the October 2025 agreement between them, one thing that could change it was AGI.

If OpenAI announced it had reached AGI, the announcement would not stand on its own, because an independent panel of experts would have to examine it and confirm it first, and only a confirmed finding would alter what Microsoft was entitled to. That panel is the closest thing the industry has ever built to a second instrument, and it exists because two parties, each wanting the other constrained, insisted on someone neutral standing between them.

Then, on April 27, 2026, the two companies rewrote the agreement, and the clause was removed entirely. Nothing replaced it. The new version ends Microsoft's exclusivity and lets OpenAI serve its products through any cloud, while keeping OpenAI paying Microsoft a share of revenue until 2030, no matter how the science goes. Which means that in the same months its executives were talking about AGI more loudly than ever, the industry was quietly removing the one place where the word had been given a referee.

One live contingency looked like it would outlast that rewrite, and it carried the largest number in the story. Investors putting money into OpenAI are buying a company that loses money now in exchange for the possibility of an enormous payoff later, and they protect themselves by releasing the money in stages, with each stage tied to something happening. OpenAI closed a funding round on March 31, 2026, at a valuation of $852B, with $50B of it coming from Amazon, and Bloomberg reported that $35B of Amazon's money turned on OpenAI either going public or reaching AGI. Going public is easy to verify, because a company either lists its shares or it does not. AGI was the other half of that sentence, and no public account of the deal said what would count as reaching it or who would be asked to decide.

Then, on July 31, Amazon disclosed in a securities filing that it had funded the full $35B. OpenAI had not listed its shares. Nobody had declared anything. The conditions that released the money have never been made public.

The buildout borrows against the destination

Money of that size explains why the destination has to stay visible, and NVIDIA's accounts show how much of it is moving. The company reported revenue of $96.2B for the quarter ended July 26, 2026, more than double the same quarter a year earlier. But the companies buying all those chips are now spending far more than they earn, so the money has to come from somewhere other than their own profits, which is why the sector has begun borrowing against what it expects to happen rather than what it currently makes.

Borrowing requires something the lender can seize if the loan goes bad, and what these companies offer is the data centers themselves, along with the chips inside them. On August 10, 2026, NVIDIA signed memorandums of understanding with six of the largest asset managers and banks, aiming to mobilize more than $500B of third-party capital for AI infrastructure on exactly that basis. So the value of the buildings is now what's holding up the loans, and that value depends on how long they continue to earn.

That is where the definition comes back. If the endpoint is a general intelligence that reorganizes the economy, then demand for these buildings will last for decades, and a lender can comfortably price a loan over 15 years or more. If there is no endpoint, the same lender is holding a shed full of chips that lose most of their value in about five years, and the loan outlives the thing securing it. Nobody at NVIDIA or OpenAI carries that risk, because the loans are held by pension funds, insurers, and credit funds, which means the money underneath them is largely ordinary retirement savings.

The Bank for International Settlements, the institution that central banks use to watch for trouble building in the financial system, has been warning about precisely this. It calculates that the five largest data center operators are on pace to spend more than $1T across 2025 and 2026, well ahead of what they earn. Its concern is not that the technology is empty, since it credits the tools with saving 20% to 50% of the time spent on individual tasks, but that everyone is making the same bet at once and only some of them can win.

The case for taking Huang at his word

Incentives that large make a cynical reading of Huang's sentence easy, though his own record does not support it especially well. He said a version of the same thing in March 2026, answering yes when a podcast interviewer offered him a definition tied to a system capable of building a billion-dollar company. If the point of saying it were to keep investors excited about NVIDIA, the effect would show up in the share price, and it did not, since the stock barely moved. A claim that fails to move a price twice is a poor instrument for raising money, which leaves the simpler explanation on the table: the man with the clearest view of what is being built thinks the label now fits, and nobody has a test capable of telling him he is wrong.

That is also the limit of what the evidence shows, because the facts establish that AGI carried a price in at least one contract and a threshold in none of them, while establishing nothing at all about anyone's intent. Something substantial happened on September 3, and nothing about the underlying system changed three days later, when its supplier gave it a name.

Which brings the story back to Ohio, where Ehman's 72 seconds are still filed as an anomaly, half a century on, because the people listening had agreed beforehand what a confirmation would have to look like. The people building never did. So when $35B finally moved on the word in July, it moved against conditions nobody has published, decided by parties who owe no one an explanation. There was never a second instrument. There was only the reading and the money that came with it.

P.S. Want to collaborate?

Here are some ways.

  1. Share today’s news with someone who would dig it. It really helps us to grow.

  2. Let’s partner up. Looking for some ad inventory? Cool, we’ve got some.

  3. Deeper integrations. If it’s some longer-form storytelling you are after, reply to this email, and we can get the ball rolling.

What did you think of today's memo?