By Jen Saarbach & Kristen Kelly, Co-Founders of The Wall Street Skinny
![]()
The Best Explanation of AI Model Companies, How The Models Work, and Who Pays for What
Anthropic and OpenAI are gearing up to go public. In anticipation of what’s expected to be (in Anthropic’s case) one of the largest IPOs of all time [see HERE], we are doing a deep dive into the business models of these frontier lab companies.
- We’re going to break down what it takes to make and sell the product
(training and inference) - Who the competitors are (open vs. closed weight model companies specifically)
- What the on vs. off balance sheet spending looks like (meaning, agreements with neoclouds and hyperscalers for compute vs. building their own data centers),
- and ultimately the key business risks.
…Minus the much-discussed chance these models destroy all of humanity.
This is part 4 of our finance of AI.

To start, let’s review the AI ecosystem. NVIDIA’s CEO Jensen Huang has described the AI ecosystem as a five layer cake. Our focus today will be layer 4: the models. OpenAI and Anthropic are the two poster children for visualizing this layer.
But in order to build the models, we first need layers 1-3: the energy, the chips and the infrastructure.
We’ve done numerous videos (all under 3 minutes each) on these, which you can watch below. They help lay the foundation for what we’re getting into here.
Layer 1: Energy
Layer 2: Chips (Part 1, Part 2, Part 3)
Layer 3: Infrastructure
Layer 4: Open vs. Closed Weight AI Models
Access the final sketches HERE.
Please note there is a password to access the DropBox link and to open the file. The password is the word “industry”
So let’s start with what these companies are building in the first place.
What are these AI Models?
The models, by definition, are Large Language Models or “LLMs”, which is basically just a text prediction machine. Text goes in, and the model predicts what word comes next. You can kinda think of it like “autocomplete”, but that dramatically undersells what is actually happening. Why?
Two primary reasons:
a.) getting the next word right forces the model to build machinery that tracks tone, the arc of an argument, and human syntax (beautiful fall day vs. that was a nasty fall down the stairs), and
b.) we can ask these machines questions, they reason through problems, and they can take actions in the world such as search the web, run code or read files.
The mechanics of how it works are as follows. The LLM takes in as much text as the model makers can get their hands on (think all the text on the entire internet, plus printed books that got scanned in), and then it applies a regression or curve fit.
On a smaller scale, we can think of it similarly to how we come up with the beta for a stock. To calculate beta, we plot the return of, say, the S&P500 against the return of a single name stock and try to fit a line to it. In other words, it’s a regression.
For my finance newbies, if the line that fits is y = 0.4x + b, the slope of 0.4 is the company’s beta.

The LLM does something similar…but instead of one coefficient (like 0.4) fitting a straight line, we’re talking about hundreds of billions of them. Each coefficient is called a “parameter” inside the network.
It’s worth noting what that “network” is. It’s a neural network, a term that goes back to 1943 when a neurophysiologist named Warren McCulloch, and a logician named Walter Pitts, reduced a neuron (a.k.a., a brain cell) to a piece of arithmetic. They published a paper modeling the nervous system as a network of logical elements. The artificial neuron takes in inputs, computes a weighted sum, and only fires if that sum exceeds a certain threshold. They showed that wiring enough of these units together could compute any logical function. The connections between those units are what we’ve been calling “parameters”. For anyone who has studied neuroscience, think of the synapses between neurons rather than the neurons themselves. A brain has something like 86 billion neurons with something like trillions of synapses between them. A frontier model has hundreds of billions of parameters.
Please note that a neural network is not, in fact, how the brain works. Everything built since has moved further away from biology, but the name stuck. And this set the foundation for how LLMs are built today.
Ok So, How Do We Build These Models?
First: pretraining.
Companies take an enormous amount of text and use it to train their models. But hang on, you can’t do math on text. Anyone who’s tried to multiply = 5 x “Hello” in Excel and gotten a #VALUE! error knows this.

The words must be converted into numbers first.
So they get split into workable units called “tokens”, averaging about three quarters of a word apiece. “The dog is sleeping” is four tokens, one per word. Longer and rarer words get split into pieces.
Before training starts, the model builds a lookup table with one row for every token in its vocabulary, and each row holds a long list of numbers called a vector. A token’s ID is just which row it points to. At the outset, every number in the table is random noise. Over rounds and rounds of training, they get nudged towards something meaningful, with tokens that appear in similar contexts drifting towards similar vectors.
In fact, there was a classic demonstration from a 2013 model where they took the vector for “king”, subtracted “man”, added “woman”, and landed near the vector for “queen”. The model was never taught that kings are men or queens women, but that was a product of the math.
Now, back to training. Take the sentence “the dog is sleeping”. It becomes three separate training exercises:
- The -> dog
- The dog -> is
- The dog is -> sleeping

Exercise one: the model takes its token for “The” and predicts what word comes next. The output is a probability across every token it knows. Pretend “dog” comes in at 0.3%.
Now, we know the real answer is 100%, since we’re looking at this one sentence and know for certain that “dog” is what should follow. So after running this exercise, every parameter in the model gets nudged a fraction in the direction that would make the model less wrong. Run it again and “dog” comes out at 0.31%. Now, because the probabilities for every word in the vocabulary sum to 100%, pushing “dog” up pushes everything else down a smidge.
Exercise two: we run the same process again, only now the input is “The dog”, not just “dog”. As we can see, the input length has grown (and will continue to grow). Because the model is training on everything that came before it, it can track subjects across a paragraph.
This happens trillions of times.
So what actually gets saved and what doesn’t? The only thing that persists is the weights. The probabilities — e.g., the 0.31% for the word “dog” — are outputs of the calculation and get re-calculated from scratch every time.
When pretraining is done, the model has become a nice autocomplete machine. Pretraining is also what consumes the bulk of the compute (more on that later).
BUT our model isn’t particularly useful yet. If you ask a raw, pre-trained model “what is your first name?”, it might answer “what is your last name?”, because it trained on tons of blank forms. While that might technically be the most likely continuation of that text (meaning the model is doing exactly what it was trained to do), it might not be helpful for your purposes.
So there are two more phases that fix that.
The first is instruction tuning, which trains the model on human written question & answer pairs.
The second is reinforcement learning, which scores the output (sometimes done by humans, but more regularly by automatic graders), and nudges it towards what scores well.
Once the model has gone through those final two phases, the weights in the model are all frozen.
The final output is a file that contains hundreds of billions of parameters.
Now this was all extremely difficult — and expensive — to achieve. So you’d want to keep it a secret all for yourself, right?
Open vs. Closed Weight Models
Well, some companies decide to let anyone and everyone see their output. That’s called open weights. It means the file itself can be downloaded and run by anyone.
Worth flagging that open weights is not the same as open source. Open source would include the training data and training code, so releasing the whole method. Open weights just means the finished product can be viewed and used by anyone.
However, most of the companies currently working on cutting edge models have closed weight models. Anthropic has never released an open weight model; OpenAI did two in 2025, but currently keeps its frontier line closed.
The labs that release their new open weight models tend to mostly be in China: Alibaba’s Qwen, Moonshot’s Kimi, DeepSeek, and Zhipu’s GLM. On the American side, Meta’s Llama line is the most widely used, OpenAI released its two lower tier open weight models in August 2025, and Nvidia has the Nemotron line. These are much cheaper to use.
We can quantify that difference in price, per million tokens, input and output:
- Anthropic’s Claude:
- Fable: $10 (input) and $50 output / Opus 5: $5 and $25 / Sonnet 5: $2 and $10/ Haiku 4.5: $1 and $5
- OpenAI’s ChatGPT
- GPT-6 Astra: $10 (input) and $50 (output) / GPT-5.6 Sol: $4 and $20 / GPT-5.6 Terra $2 and $12 / GPT-5.6 Luna $0.20 and $1.20
- Hosted open models span a far wider band.
- DeepSeek’s v4 tiers run $0.22 to $0.66 on input and $0.66 to $2 on output
- Moonshot’s Kimi K3 runs $3 and $15 on input / output.
So the frontier companies charge roughly the same prices, whereas the open source models, Deepseek in particular is significantly cheaper, about 25x below the frontier.
The next question becomes, what does getting access to an open weight file actually do for you? Is there an “open weight app” you can download from the app store? Not quite.
Turning a file with hundreds of billions of numbers with code to run them into something useful requires the following:
- The weights, which can be downloaded free from Hugging Face which is essentially the file repository for the industry.
- Hugging Face (named after this emoji 🤗) is that same company that OpenAI’s own models broke into this past summer and the company Nvidia is now reportedly acquiring for around $13 billion.
- Compute: in some cases, smaller models can be run locally on your personal device. For larger models, compute needs to be rented from Amazon’s AWS, Google, Microsoft’s Azure, CoreWeave, or similar.
- Serving software, such as vLLM to run the file
- A harness, meaning” the tools, memory, guardrails, and monitoring, which you can build yourself if you want.
So what does this mean practically speaking?
Well step 1 is easy…download the file for free!
Steps 2 and 3 are often outsourced to companies like Together, Fireworks, or Amazon’s Bedrock. They put the model weights on their servers and give you a web address to send text to. That address, plus the rules for how to talk to it, is the Application Programming Interface (“API”). An API is the thing that lets your software interact with the model. You pay per token, similar to how you pay for Anthropic.
Step 4, the harness, has its own off the shelf options too. You could build the harness yourself using developer toolkits like LangGraph or install a finished chat app like Open WebUI. Anthropic sells a harness too — Claude Agent SDK — although it only points at Claude models.
Now quick aside for those who asked us to explain how Perplexity and Microsoft Copilot fit in. Both are examples of the harness layer, products built around models rather than the model themselves.
Perplexity wraps a web index and citation engine around a mix of its own models (Sonar) and rented frontier models from OpenAI and Anthopic. Sonar is also a good illustration of what having the file actually gets you. Perplexity took Meta’s open weights (from Llama) and retrained them into its own in-house model, tuned for the type of search and site answers its product needs. Copilot on the other hand is in the harness Microsoft built into Office and Windows, letting the model reach into your documents and calendar.
Ok, so almost none of these have to be built from scratch. But if it’s this easy to use an open source model, doesn’t that pose a huge risk to the business case for Anthropic and OpenAI?
The answer is “maybe”, but:
a). It does require a bit of technical savvy that many consumers lack. I can promise my 80 year old mother isn’t going to Hugging Face and downloading an open source model but she is using Claude and ChatGPT.
b). Open source models are almost always a level behind, because many rely on “distillation” to create their models. Distillation means companies aren’t training their models from scratch. Instead, they’re taking someone else’s fitted model and generating predicted values from THAT, then regressing based on those results. Their coefficients approximate the frontier labs at a fraction of the cost.
Additionally, while your average consumer likely does not need the frontier models, many companies (and governments) that are using models to do cutting edge research at the enterprise level do.
There’s a moat, for now.
But that moat hinges upon these labs’ ability to develop and unleash the next / better version. It at least partially explains the race to build new models faster and faster, even when that race might be at odds with what is best for humanity.
Now there are two other important points. First is on the revenue side at companies like OpenAI and Anthropic. What they charge via subscriptions, like a $200 monthly Pro plan, can cost the company more than it takes in from heavy users. Forbes reported that someone on a $200 a month plan could consume $5,000 in compute.
So for a heavy user, Claude may be the cheaper option right now. Buying that same usage metered through the API could cost many multiples of $200, and the open weight alternative still requires you to pay someone for every token you run. The subscription is underpriced/being subsidized, which is why it loses money. So right now the business case seems to be get users hooked on Claude, build the surrounding tooling, and figure out how to raise prices later once people are locked in.
But let’s talk about the expense side as it relates to chips, compute, and the rest for these frontier model companies..
AI Logic Chip Requirements
We can’t have a conversation about CapEx without talking about the chips that run these models — specifically, the logic chips.
When training a model, switching between different companies’ chips is extremely costly.
In other words, a lab that has trained its model on NVIDIA’s GPU’s can’t easily swap in Google’s TPU. The model itself is just a file of numbers and the math is just multiplication, but each chip family has its own software stack for turning model code into instructions the chip can execute. Getting a frontier model to run efficiently on a new stack takes months of specialist engineering. That is why CUDA, Nvidia’s programming layer is such a moat, along with the networking and developer tools built around it. Nvidia is entrenched with these AI companies the way Microsoft Excel is entrenched on Wall Street — bankers already know Excel shortcuts, so while Google Sheets might work similarly, it would require an investment in retraining.
With that said, companies have started to diversify. Anthropic has deliberately built its models to run not just on Nvidia’s GPUs with CUDA, but also on Google’s TPUs with XLA and AWS Trainium with Neuron. OpenAI still mostly uses Nvidia GPUs but has committed to 6GW of AMD GPUs and recently announced its own custom chip called “Jalapeño”, built by Broadcom, and designed specifically for inference.
Which brings us to training and inference.
Training is the more demanding workload of the two. Thousands of chips have to act as a single machine, all working the same job in lockstep, so the interconnection between them is just as critical as the chips themselves.
Inference — the fully built model now performing tasks in response to prompts from customers — splits into many small independent jobs, which is far easier to spread across cheaper and older chips.
This connects back to the Michael Burry depreciation debate [see HERE] where the whole fight was over how long these chips stay useful. Burry’s argument was that the hyperscalers (Meta, Google, Amazon, Oracle) and neoclouds (Nebius, Coreweave etc.) are booking 5 to 6 year useful lives when the real economic life is closer to 2 to 3 years. But there does need to be an acknowledgement of the training versus inference split. A chip that’s no longer sufficient for training a frontier model can still do years of useful work serving inference (at least for now), which is the strongest case that these assets can get repurposed and used to generate revenue vs. retired.
But notice that debate isn’t necessarily about Anthropic or OpenAI. The question of the real useful life and value of a chip is a critical one for whoever owns the hardware. For the most part that isn’t the frontier labs. Their business is running models, not owning infrastructure.
The vast majority of their compute comes through lea