Podcasts by Category

NVIDIA Generative AI

NVIDIA Generative AI

Cloudadorn Academy

What is actually inside the phrase generative AI, and why does every answer end up being a cost decision? NVIDIA Generative AI is a twenty-three-episode series from Cloudadorn Academy for people who want the real picture and are not already specialists. Shelley asks the questions a curious adult would actually ask. Rob answers with something you can picture first and the name of the idea second: nesting dolls for the AI hierarchy, a group chat for attention, sand and statues for diffusion, a case conference for sensor fusion. Season one walks four acts. What these things are, from the hierarchy and the learning paradigms through attention, training, and the architecture zoo. How you make a model yours for the least money that works, from prompting to retrieval to LoRA to a full fine-tune, and which number tells you it worked. How a model gets eyes and ears, through shared embedding spaces, patches as tokens, diffusion, segmentation, and sensor fusion. Then the stack itself: precision formats and memory bandwidth, serving and packaging, the training pipeline, agents, guardrails, robots and world models, and why the ecosystem is hard to leave. NVIDIA is the through-line because it sells at every one of those layers, so its product map doubles as a map of the field. Where a vendor number cannot be confirmed against a primary source, the hosts hedge it or drop it and say which one they did. This series is narrated using AI voice technology. The content and scripts are original. Shelley and Rob are original hosts, not impersonations of real people. If the field finally holds still long enough to make sense, subscribe, leave a review where you listen, and visit cloudadorn.com.

23 - S1E01. Nesting Dolls: What Generative AI Actually Contains
0:00 / 0:00
1x
  • 23 - S1E01. Nesting Dolls: What Generative AI Actually Contains

    This episode is narrated using AI voice technology. The content and script are original. Artificial intelligence, machine learning, deep learning, generative AI. Four phrases, used interchangeably, by the same person, about the same product. They are not synonyms. They are nested, and every step inward is a stronger claim about what is actually happening. Shelley makes Rob open the dolls one at a time, biggest to smallest. You will learn why the nesting only runs one way, so a chess program built from hand-written rules is AI and is not machine learning. Why foundation model is not a fifth doll at all: it describes a model trained wide enough to be reused, which is not the same as a model that makes things, and CLIP is the counterexample that proves it — broad, reusable, and it makes nothing at all. NVIDIA's Cosmos family is the other half of the point: world foundation models that generate video of the physical world, also not language models. The difference between learning the boundary between categories and learning the shape of the data, which is the whole reason a model can make something new instead of only sorting what exists. And the four ways a model learns, including the one people get backwards: pretraining a large language model is self-supervised, not unsupervised, because hiding the next word turns the sentence into its own answer key. Then the arc. 2012, when a deep network won the ImageNet competition by a distance. 2017 and the paper that threw out recurrence. 2020 and a model doing new tasks from examples in the prompt. 2021, when the same architecture came for pictures. Underneath all of it, multiplying big grids of numbers on a chip that was built for video games. Takeaway: when a product says AI, it has told you almost nothing. Ask which doll. Subscribe for the rest of the season, and visit cloudadorn.com.

    Tue, 25 Aug 2026
  • 22 - S1E02. Attention: How a Model Reads a Sentence

    This episode is narrated using AI voice technology. The content and script are original. You type a sentence. Something types back. This is what happens to your words in between, in order, with the clever bit slowed right down. Rob walks Shelley through the five steps every one of these models runs: the sentence is chopped into pieces smaller than words, each piece becomes a location in a space, position gets added on purpose, the stack of blocks runs, and one word comes out. Then it all runs again. The step that sounds like housekeeping turns out to be the one worth stopping on: attention is blind to order. It sees a heap, not a line. Leave that step out and "the dog bit the man" and "the man bit the dog" are the same input to the model. Then attention itself, which is smaller than its reputation. Every token asks a question, advertises what it knows, and offers something to contribute. Everything gets weighed against everything else, and nothing is ever ignored, only weighted near zero. That last part comes with a bill: ten pieces is a hundred comparisons, twenty is four hundred. Double the length, quadruple the work. That single curve is why long context is expensive, and every trick in the back half of this episode is somebody trying to get out from under it. Also: why a blindfold is the reason text models are built the way they are, how to tell a model's shape from the job it does, and why one model honestly has two parameter counts ten times apart. NVIDIA's Nemotron 3 Nano has 31.6 billion parameters and runs 3.2 billion of them per token. When somebody quotes you a parameter count, ask which one.

    Tue, 25 Aug 2026
  • 21 - S1E03. How a Model Learns: Forward Pass, Loss, and Backpropagation

    This episode is narrated using AI voice technology. The content and script are original. Training a model is four steps in a loop: guess, score the guess, work out who is to blame, nudge everybody, then do it again. When you type into a finished model and it answers, that is the forward pass, and the forward pass changes nothing. Learning is the backward pass plus an optimizer step. A model out in the world is frozen. Shelley makes Rob slow the blame step down. One number says how wrong the answer was. Billions of numbers inside contributed to it. Backpropagation works out, for every one of them, how much of that error was its fault, by going backwards a layer at a time. The loss you pick is you telling the machine what you care about: next-word prediction is classification over the whole vocabulary. The crude activation beat the sophisticated ones because it does not destroy the correction on the way back. Residual connections are a megaphone so the message is not whispered through sixty handovers. Six knobs, each of which breaks training in its own way, and the distinction that is the spine of the last third: parameters are learned. Hyperparameters are chosen. Takeaway: nothing is learned on the way forwards. Subscribe for the rest of the season, and visit cloudadorn.com.

    Tue, 25 Aug 2026
  • 20 - S1E04. When Training Goes Wrong: Overfitting, Bad Data, and Metrics That Lie

    This episode is narrated using AI voice technology. The content and script are original. There are two ways for training to fail, they are opposites, and the fix for one makes the other worse. Overfitting is the driver who learned one route to work perfectly and is helpless on any other street. Underfitting is the driver who had one lesson and stopped. The only way to tell which you have is a slice of data locked in a drawer before you start. Rob walks Shelley through the pair of numbers that is the entire diagnosis, early stopping done honestly on a bumpy curve, and why looking at the data comes first. Label errors in the collections this industry measures itself against. Five data failures, each with its own symptom: lopsided categories, gaps, extremes, leakage, duplicates. Accuracy is the number that lies most often when the rare thing is the whole job. Precision and recall point at two different mistakes, and each can be gamed alone. Takeaway: the most expensive mistake is treating the wrong failure. Subscribe for the rest of the season, and visit cloudadorn.com.

    Tue, 25 Aug 2026
  • 19 - S1E05. The Architecture Zoo: CNN, RNN, Transformer, GAN, VAE, Diffusion

    This episode is narrated using AI voice technology. The content and script are original. These are not brands. They are shapes. Somebody looked hard at one kind of problem, worked out what shape of machine would suit it, and the shape got a name. Every name on the shelf is superb at one thing sitting directly next to a thing it is hopeless at. Rob names them that way, and no deeper: the CNN sliding a window over a grid, the RNN walking a sequence and forgetting, LSTM holding on longer, the transformer looking at every position at once and paying quadratic rent for it, Mamba as the young challenger whose work grows with length instead of length squared. Then the makers: a GAN is two networks set against each other, and mode collapse is the word that belongs to that family only. The autoencoder squeezes a thing down to a recipe. The VAE lets you walk around inside that recipe. Diffusion gets named and handed forward. Latent space is the idea worth taking with you. Takeaway: ask what the shape is for, and what it cannot do. Subscribe for the rest of the season, and visit cloudadorn.com.

    Tue, 25 Aug 2026
Show More Episodes