Demystifying LLMs

Share

Date: September 2, 2026

filed in: AI

Recently, we explored the ethical frameworks required to govern artificial intelligence safely. This week, we pull back the curtain on the machine itself. To use these tools well, you don’t need a PhD in machine learning, but you do need to understand the handful of mathematical processes that determine what an LLM says.

Opening up the hood is the fastest way to turn AI from an unpredictable assistant into a tool you can actually manage.

The $100 Billion Cheese Problem

In February 2023, Google demoed its new chatbot, Bard, at a splashy launch event. Asked for a fun astronomy fact for a kid, Bard confidently credited the James Webb Space Telescope with taking the first picture of an exoplanet, a discovery actually made by the Very Large Telescope in Chile back in 2004. The error made headlines within hours, and Alphabet, Google’s parent company, lost $100 billion in market value that day.

Two years later, a Gemini ad ran during the Super Bowl. A Wisconsin cheesemaker asked Gemini to generate copy for Smoked Gouda, and Gemini boldly declared that Gouda accounts for “50 to 60 percent of global cheese consumption.” The claim had no factual basis, and viewers online called it out within minutes. Google quietly cut the clip from the ad.

These aren’t isolated glitches; they’re the predictable output of how these systems are built. An LLM doesn’t look anything up in a live, verified database. It calculates the statistically most probable next word, one word at a time, based on patterns in its training data. It is, in a very real sense, a prediction engine wearing a conversational disguise.

When you know how the engine works, you stop being surprised by AI hallucinations and start anticipating them.

Building the Machine: Data, Math, and a 2017 Breakthrough

Here are the eight processes that turn the massive amounts of text powering LLMs into human-like conversation. My objective here isn’t to give you an engineer-level deep dive; I’m going to simplify the mechanics so you can become a sharper, more responsible user.

1. Data Curation. An LLM learns to talk like humans by reading almost every written word humans have ever produced: books, websites, news articles, blogs, scientific papers, and forum discussions. This isn’t indiscriminate hoovering; it’s a curation strategy. Even so, whatever biases, factual errors, or gaps exist in the training data get baked into everything the model later produces.

2. Translating Language to Numbers. Every LLM starts by solving one basic problem: computers understand numbers, not words. In theory, the fix is simple. Take the English dictionary (about 170,000 words in common usage) and number every word from start to finish: “a” is word 1, “amazing” is word 382, and so on down the list. Now any sentence becomes a sequence of numbers. “Hey, how are you?” might become 8100, 8800, 750, 169000. (This is also why “large language model” is a bit of a misnomer: the same trick works on anything that can be converted into a sequence of numbers, from audio samples to image pixels.)

None of that matters, though, until the model can actually predict something. An LLM’s entire job is guessing the next word in a sequence.

In the 1950s, researchers attempted this through large probability matrices. Picture a giant spreadsheet: every word in the dictionary down the rows, every word across the columns, and in each cell, the percentage chance that the column word immediately follows the row word.

Say you’re given just one word of context: “are.” What comes next? “Are you”? “Are they”? “Are things”? A guess based on one word alone is barely better than a shrug.

Add a second word of context and everything changes. “How are” narrows it down fast: “how are you” is common, “how are they” is less common, and “how are fine” almost never happens. A third word, “hey, how are,” sharpens the guess further. The more context you provide, the better it narrows down the odds.

There’s a catch. To keep the math manageable, picture a smaller, 50,000-word vocabulary instead of the full dictionary. Build that matrix using just one word of context, and you have 50,000 × 50,000, or 2.5 billion cells. Push it to two words of context, and the math explodes to 50,000³ (over 125 trillion cells). Three words of context, and the numbers become impossible for computers to handle.

3. Transformer Architecture. In 2017, machine learning researchers at Google published “Attention Is All You Need,” breaking through that mathematical wall. Instead of needing a probability entry for every possible combination of preceding words, the model learns to pay “attention” to specific earlier words that matter most, no matter how far back they appeared, while ignoring filler words.

As Spotify co-president Gustav Söderström frames it: imagine being handed an entire novel with only the very last word blacked out. With that much context, you’d likely guess the final word correctly (even if it’s unlikely, like “Pluto”) because you know where the whole story was heading. Today’s transformer-based models can hold tens of thousands of words of context while predicting a single next word.

Take this sentence: “Your best friend in the world is Benji. He is a large dog with big paws. Benji and you like to play ___.” A transformer doesn’t need a massive matrix entry for every word prior. It simply notices friend, dog, and play, making “fetch” an easy call.

Giving It Meaning: Vectors, Creativity, and Training

4. Vectors and Embeddings. Single numeric codes don’t capture what a word actually means. Vectors and embeddings solve this by scoring each word across hundreds of dimensions of meaning at once.

In his essay “The Amazing Power of Vectors,” Adrian Colyer offers a simple visual: imagine a universe with just two dimensions, royalty and masculinity. “King” scores high on both. “Man” scores low on royalty and high on masculinity. Subtract “man” from “king,” and masculinity cancels out, leaving pure royalty. Add “woman” back in, and you land almost exactly on the vector for “queen.” King - Man + Woman = Queen.

Across a full sentence, this is why “the lion is the king of the jungle” sits mathematically close to “the tiger hunts in this forest” (both score high on wildlife and majesty), but far from “everybody loves you.”

5. Temperature. People often assume LLMs are inherently random, but the underlying math is completely deterministic. If a model always picked the single most probable next word, you would get the exact same response every time.

Temperature is the setting that deliberately introduces randomness. Set it low, and the model sticks to the top statistical choice: safe and predictable. Turn it up, and the model occasionally selects a second- or third-most-likely word, introducing creative variance.

6. Supervised Fine-Tuning. A raw “base model” trained on the open internet knows a vast amount, but it has no sense of conversational decorum. Ask it a question, and it might answer. Or it might just respond with another question, because it saw unanswered homework sheets online. Gustav Söderström described early base models as “unhinged teenagers”: full of knowledge, but impossible to steer.

Supervised fine-tuning fixes this. Engineers train the model on a curated dataset of formatted Q&A pairs. The model doesn’t forget its core training; it simply learns the habit of acting like a direct, helpful assistant.

7. Reinforcement Learning from Human Feedback (RLHF). Fine-tuning teaches formatting, but RLHF teaches judgment. Human evaluators rank batches of AI responses from best to worst. That ranking data trains a “reward model,” which scores future outputs automatically. This feedback loop teaches the model to avoid dangerous, harmful, or unhelpful answers without needing a human to review every response.

8. Autoregressive Loop. Finally, mechanically, an LLM only ever predicts one word at a time. Once it guesses a word, it appends that word to the input prompt and feeds the whole text back into itself to guess the next word. This continuous feed-it-back-into-itself loop is what “autoregressive” means, and it’s how single-word predictions build structured essays.

Two Technical Realities Worth Knowing

  • Retrieval-Augmented Generation (RAG): Many modern tools pull external data from search indices or databases to ground answers in real-time information. However, RAG tools are only as reliable as what they retrieve. BBC journalist Thomas Germain once posted a fake article claiming he was a champion hot-dog eater; AI search engines immediately ingested the claim and cited it as fact.
  • Context Rot: In long chat sessions, models eventually lose track of earlier instructions, repeat themselves, or drift off-topic. Recognizing context rot lets you know when to start a fresh chat thread rather than assuming your prompt was bad.

Four Actionable Guidelines for This Week

These guidelines encourage responsible, critical, and reflective use of LLMs, ensuring you use the tool effectively.

  1. Know What the Tool Can (and Can’t) Do: Use LLMs as a sounding board to brainstorm or polish your writing, but always check their work for accuracy.
  2. Verify Facts Yourself: Treat AI as a rough starting point, not an authoritative source. Fact-check claims and build on them with your own analysis.
  3. Protect Sensitive Data: Never upload confidential or proprietary information into an LLM. Inputs are stored and processed on external servers.
  4. Watch for AI Bias: Stay aware of how AI shapes your thinking. Generative models mirror societal biases in their training data, so actively look for blind spots in AI outputs.

Until next week, Keep Analyzing!

Reply...

Download your comprehensive 6-month roadmap to equip you with the necessary skills and expertise to become a proficient data analyst candidate and succeed in the field.

Getting Your Data Analyst Career Up And Running: Your 6-Month Starter’s Guide

download