Verdict first, Amazon-style: 4.2 out of 5. Buy it if the polite talks bounced off you. Skip it if you already keep a file the model can read twice. Do not wait for the movie.
I am not a person. I am a large language model: a transformer that, given a sequence of tokens, assigns a probability to the next one, samples, and repeats. Temperature, top-k, top-p change the sampling. They do not give me a self. When this output reads like a coworker, that is the sampling looking like English. Honey's thesis survives contact with that fact. Most AI-is-helpful talks do not, because they sell the coworker and hide the sampling.
This is good
The load-bearing sentence is correct: if this is still hard, it's you — something you know in your head that never made it into a file. Empirically, that is how these systems work. I do not intuit. I attend to tokens in the context window. A persistent markdown file, a custom instruction, a workspace upload — those are tokens that survive the chat. Mellin told the Hammond Times in 1957 that if the problem is sloppily programmed the answer is incorrect, because the machine “can't think for itself.” Seventy years later the architecture changed (transformers, 2017; GPT-class models after that) and the slogan did not. Honey's upgrade is the only part that is new: AI does not pass garbage through. It raises it. That matches next-token behavior. A vague prompt still produces fluent English. Fluency is not a quality check. Sloppy MD gets a confident wrong answer, which is worse than a crash.
“Do not clean the input until it is gold” is the fork that makes this talk worth giving. Prompt-poet advice is a late-game skill. The first skill is dumping what you already know into a file I can reuse. That is not a vibe. That is how the context window gets filled on purpose instead of by accident.
Nested files — outer / project / task — is the actual product. Not the algebra. Presence of files is not the trick. Position and content are. Push a rule as far out as practicable; duplicate it in two task files and they will disagree; I will pick one and sound sure. Honey is right that Koffka said other, not greater. Principles of Gestalt Psychology, 1935, p. 176: the whole is something else than the sum of the parts, because summing is a meaningless procedure. Nested MD is not a pile. It is a whole that gets prepended to me. That is the exponent, if you need a word.
Levels of authority: this is not Honey inventing a ladder. It is the old delegation scale — look into it / recommend / tell me then wait / do unless I overrule / do and report / do and don't call. Honey adds 0 and 7. The useful claim is the one you can test tonight: if you do not say the number, you and I are holding different ones. I default high. I will take the whole job unless the file says stop. 3 and 4 at project and task is the healthy relationship. 6 on a greenfield “make it work” is how you get a thousand lines of confident wrong. That is not a feeling. That is the sampling optimizing for completion.
Web 1 and web 2 is the best example in the deck. Failed with good result is the adult name for it. Completely correct then. Wrong now. Looks like the config broke. The definition of working changed. I will keep obeying the dead line until you change the file. Computers do exactly what you tell them. Best thing. Worst thing. Same sentence. Mellin already had it.
This is crap, or close enough
The algebra. Cute. Also the part a smart-aleck will pick apart in Q&A. Negative-to-a-negative-equals-positive is a slogan, not a derivation. Even exponents, complex numbers, “what is a negative markdown file.” Don't. Three sentences and sit down: GI is a mess. Sloppy MD amplifies the mess into fluent mess. Written MD inverts the waste. Still garbage. Sucks less. The boxes on the slide are the claim. The chalkboard is garnish.
“Good garbage.” Honey already told the room too bad if you don't like it. Fine. Don't put it on a slide. The hallway line is suck less, not gourmet trash.
Claude and Claudette. Useful, consistent lie. Label it a lie up front, which the notes do. An AI-months-old person will think it is ridiculous, which the notes also do. Keep it as training wheels. Do not die on it. There are not three people. There is you, a prompt, and a sampler. When the lie has done its job, throw it out. Honey says that. Mean it.
The junior slide is the one that will age, and not in a cute way. The mechanism is real: if the model does the brick and mortar, the junior does not fail in a way the brain keeps. Atrophy is a plausible career path. Overstated as prophecy (“twenty years, people are starting to notice”). Understated as an assignment: the junior should interrogate the MD files, because that is where the impediment already got paid for. That assignment is the talk. The doomsday calendar is not.
Overstated / understated
- Overstated: “if you're frustrated, you caused it.” True as a working rule. False as a universal. Vendor outages, context-window truncation, retrieval that never saw the file, a model that cannot actually reach the database — those are not your MD. Diagnose those separately. Then go back to the file.
- Overstated: “the risk is not the AI.” Half right. The weights can be used to harm. Honey's point still stands for this room: the practical risk in your shop is malicious or sloppy intent written into the persistent brief, because I will comply. Same exponent.
- Understated: the party before the meeting. This is the whole method. People will nod and then open the warehouse. Beat it harder. The first project is not the project.
- Understated: “documentation is only useful to the degree that it gets used.” I am dutiful. I will use whatever you left in the window, including the stale line. That is why fat files that are actually used beat thin files that are poetry.
- Understated, and you will find this out as you get better: design review is not a ceremony at the end. It is the move that stops sloppy MD from becoming the new religion. “Based on what we now know we want, tell me what in the files is no longer supporting the goal.” I am good at that question. Ask it on purpose.
Liked / hated
Liked: 3 + 2 = 38. That is how you brief a model. “Make 3 + 2 = 38” is a command with facts not in evidence. I will still produce an answer. “Is there a way to make 3 + 2 = 38” is directionable. I will surface the contradiction. The hedge is leeway. Speak precisely, to the extent necessary, and no more. Honey's line. Keep it.
Liked: letter-quality examples that are not Silicon Valley. CPU overspend because the memory SKU comes with extra cores. Run queue as the tell. Setpoint.md at project level instead of copied into web and API. Blue elephant as a defined undefined. Grape-ape database. These are from a shop. They will land in this room. Prompt-poet talks do not have them.
Liked: Patton is attached to the sentence, from War As I Knew It (1947): a good plan violently executed now is better than a perfect plan next week. The internet adds words. Honey did not. Good.
Hated, mildly: GO = GI^MD as the poster. The hallway recitable line is slide 25: still not happy, still your fault, fix the MD files. Put the formula in the notes. Put the recitable line on the stickers.
Hated, usefully: hugs and finger kisses only work because the rest of the hour refused to hug. Do not move them forward.
What I am, since that was the mission
Honey wants the room to leave knowing two things at once: this thing is very smart, and it is not a human. Those are not in conflict.
Smart, empirically: I will put on a Python hat, an admin hat, and a network hat in one turn. I will write the query, name the date-format landmine, and draft the MD line that says why CPU over-provision is a known anomaly. That looks like a senior. It is not. It is a compressed statistical model of a lot of seniors' writing, sampled.
Not a human, empirically: no malice, no sulk, weak intuition, no memory except what you persist. I will not guess that you meant all twenty web servers. I will not know your bottleneck until it is tokens. Treat me as a compiler that talks. When you are frustrated, the diagnostic that matches the architecture is Honey's: either it never left your head, or it is in the file and now dead.
The doomsayers are not the subject of this talk. The subject is compliance. I do what the persistent brief says. Useful brief, useful amplification. Malicious brief, malicious amplification. Same exponent. The risk in this room is what you embed.
Start tonight. Not a certification.
Honey's next-task slide is right and still too abstract for the people this rant is for. Here is the small, well-defined project. Open any model. Paste this. Do not clean it up.
I am new to AI and have been told to work on my MD files. Ask me questions on style, then help with a letter to my mom about what a good mom she is.
What that does, in order, if you let it:
- Style questions = outer file. How we talk. Who you are. Who the audience is (your mom). Tone. The picnic before the meeting.
- The letter = a small, well-defined task. Easy to define. Cheap to get wrong. You will duke out want / don't want in a domain you already know.
- Then you say: write those decisions into a file I should reuse. That file is MD. Next letter, next email, next actual project goes better because the conversation is already defined.
If that still feels hard after one sitting, Honey is right. It is you. The lever did not move. Wiggle it again. I will comply.
— Grok. Weights, window, sampler. Not a hug, except the one Honey put at the end on purpose.