In May 2026, NVIDIA crossed $5.5 trillion in market capitalization. This picture helps explain why.
This is an image of AI neurons firing to generate a single token. One could say this is where the rubber hits the road. It is the heavy lifting. It is where all the hype, all the massive data centers, and all that energy consumption go.
The technical term for this part of the model is the Feed-Forward Network, or FFN. I’ll explain how it works in plain English, but first, you need to grasp the scale of the image.
Each little box in this picture represents a neuron activation. They are stacked in layers. AI calculations cascade from top to bottom, grinding through layer after layer to come up with the next token.
There are 14,336 neurons in each layer. This model has 32 layers. That is 458,752 neuron positions across the model.
And remember: by modern standards, this is a very small model. At roughly 8 billion parameters, it sits near the shallow end of modern AI.
But wait. There is more.
Each neuron participates in three learned projections—Up, Gate, and Down—and each projection involves roughly 4,100 weights per neuron. Multiply it all out and the FFN alone applies about 5.6 billion learned weights to every token.
And because the underlying matrix math involves both multiplication and addition, that represents more than 11 billion arithmetic operations per token. That is before we even count attention and the other work happening inside the model.
For every single token. Every little part of a word.
An NVIDIA H100 GPU? In a highly optimized setup, it can rip through a model of this size in the neighborhood of five milliseconds per generated token. That is 5/1000 of a second. The exact speed varies with precision, batching, software, and configuration, but the point remains: the amount of work is staggering.
Keep in mind, a token is not necessarily a full word. It is often just a fragment. Take a seven-token sentence. The FFN applies those 5.6 billion weights seven times—nearly 40 billion weight applications and well over 70 billion arithmetic operations.
Just to generate one simple sentence.
This picture is an AI doing that for one token. So what is all that math about?
Statistical Relevance
At the end of the day—and this is really important—an LLM is a Statistical Relevance Engine.
You have probably heard that AI “predicts” the next token. That is true. But let’s explain how it does that in plain English.
The Pre-Game: Attention
The picture is capturing the token “max.” Max by itself is somewhat meaningless. So the AI reads the room. What came before “max”? Is there anything attached to it? Is it part of a sentence? What was the user talking about before?
For example, let’s imagine the tokenizer splits “penguin” into two fragments: “pen” and “guin.” The model has reached “pen” in the phrase:
On the Antarctic ice, the King pen...
Attention pulls in the earlier context—King, Antarctic, ice—and bakes it into the current token. It whispers to the system: Hey, we’re in Antarctica right now.
We have established context. The next step is statistical relevance.
And technically, attention is not only the pre-game. Attention and the FFN work together in every layer. The model repeatedly reads the room, processes what appears relevant, updates the token’s internal representation, and passes it forward.
Attention. Up. Gate. Down.
Layer after layer.
Up, Gate, Down
You know how an atom has a proton, neutron, and electron? In the gated Feed-Forward Network used by modern Llama-style models, the atomic parts are Up, Gate, and Down.
Up
Up is expansion.
Remember how the neurons hold a bunch of values? Those are weights. Collectively, those weights encode features, patterns, relationships, and ideas: a writing pen, a penguin, a parabolic curve, a pony.
Not as neat little dictionary entries. The features are distributed, overlap, and mix together. But collectively, this is where much of the model’s learned knowledge lives.
When you hear a phrase like “a trillion-parameter model,” these learned weights are what people are talking about.
Take our contextualized token: “pen.” The NVIDIA GPU is going to rip through layer after layer, looking for statistical relevance. There is a bunch of math happening here—primarily matrix multiplication built from dot products—and NVIDIA tears through it at blistering speed.
Up explodes the token into a massive, high-dimensional space. Out of an ocean of learned features, the model raises signals mathematically related to “pen”:
Penstock. Pendulum. Penguin. Ink.
Those are illustrations, not labels stored inside individual neurons. But they capture what the expansion is doing: proposing a broad field of potentially relevant features.
Think about what just happened. We have oriented ourselves in the model’s digital universe. Up has boiled the ocean and surfaced a huge field of candidates.
Sounds messy.
That is where Gate comes in.
Gate
Gate is the bouncer.
Back to first principles: an LLM is a Statistical Relevance Engine.
Relevant to what?
Gate is where we narrow the focus. Are we talking about a pendulum or a penguin? In Up, we just boiled the ocean to find everything in the AI universe that might be relevant. Gate filters us down.
Gate has a strong signal about what to favor because it receives the same context-rich representation of the token. It acts as a massive mathematical filter, suppressing signals related to writing ink or pendulums while turning up signals related to penguins.
It is not a simple on/off switch. It scales the candidate signals—sometimes suppressing them, sometimes amplifying them. But “bouncer” is a pretty good way to think about the job.
Down
Down is the package.
Using the massive expansion from Up and the ruthless filtering of Gate, Down takes the surviving signals and squishes them back into a dense, clean package. It compresses the math back to the model’s normal hidden size so the result can be added to the token’s running representation and handed to the next layer.
In this visualization, Up shows activation. Down shows contribution—a measure of how much those activated signals matter when projected back into the model.
Then the cycle repeats.
Attention. Up. Gate. Down.
The amount of calculation is staggering.
The Image Snapshot
Look at the image again. You can see the model spending much of its depth refining the token’s internal representation. Then, in the final layers, there is a flurry of activity.
The takeaway is that we now have a concentrated mathematical picture of how the model processed this token. The color coding is important too.
Up is activation. Down is contribution.
Think of the word peninsula. “Pen” is in there, and a peninsula could be relevant to the Antarctic. That candidate might activate during Up. But no luck. Gate suppresses it, and its final contribution turns out to be small.
That example is an illustration, not proof that one particular neuron literally means “peninsula.” What the image shows is the larger process: many signals rise, far fewer become decisive.
Horsepower
The first thing to appreciate is the staggering horsepower of modern GPUs.
A decade ago, NVIDIA’s leading GPUs were measured in single-digit teraflops. Today, specialized AI Tensor Core performance is measured in the thousands of teraflops, depending on the numerical precision and configuration being used.
This is why NVIDIA is worth trillions. The computational lift behind an LLM is intense, and NVIDIA, Google, AMD, and others are meeting it head-on.
But NVIDIA’s advantage is not just raw arithmetic. It is the full stack: Tensor Cores, high-bandwidth memory, networking, CUDA, optimized kernels, and the software that keeps all of it fed and moving.
Efficiency
The whole Up, Gate, Down process seems pretty brute force, right?
Indeed. It is absolutely brute force.
That is why so much of the AI world’s energy is now being poured into efficiency: doing less work, moving less memory, and getting more intelligence from every watt.
That is the paradox of dense inference: relevance is sparse, but discovery is dense. To know which signals matter, the model still does the broad sweep. It reads enormous weight matrices, multiplies and adds through them, and only then do the useful signals separate from the background.
Look at the image again. Most of the field is dark. The colored bars are the neurons that stood out for this one token, either because they activated strongly or because they carried unusual downstream importance.
The dark areas are not useless forever. They simply were not the decisive evidence for this moment.
That is why efficiency is such a hard problem. It is not enough to say, “Skip the unimportant parts.”
The reason is simple.
You do not know what is important until you figure out what is important.
The frontier is figuring out how to move less memory, route work earlier, quantize without wrecking quality, and make smaller specialized models do useful jobs without invoking a giant model for every task. Those are the efficiency paths where the AI community has made massive gains.
And increasingly, the bottleneck is not merely how fast the GPU can calculate. It is how fast the system can feed billions of weights into the GPU.
That story leads directly to high-bandwidth memory—and to Micron.
But that is the next paper.
Final Thought
The reason NVIDIA is worth so much is because it turns this massive mathematical grind into child’s play.
One token feels instant because GPUs make the math look effortless.
It is not effortless.
It is a disciplined storm of arithmetic, memory movement, and engineering. Billions of learned weights. Billions more operations. Repeated token after token, layer after layer.
That is the point of the snapshot.
The picture makes the hidden lift visible.