# Lecture creation guidelines

Canonical rules for every lecture deck produced in this project. Future agents: read this file fully before writing any deck content. The project goal is a general, high quality system for creating lectures; these rules are the quality bar, and each has been requested explicitly by the project owner.

Owner attribution (title slide and sources slide of every deck):
Mehmet Kerem Turkcan; Associate Research Scientist; Center for Smart Streetscapes, Columbia University; New York, USA; keremturkcan.com; mkt2126@columbia.edu

## Language constructs we want

Phrased positively: write with these constructs, and reject drafts that lack them.

1. **Punctuation from the set `, ; :`** (plus `. ? ! ( )`). No em dashes, no en dashes, no hyphens as punctuation. When a hyphenated compound tempts you, rephrase it: "a list of fixed length", "models that look at the last n tokens". Proper names keep their official spelling.
2. **Nested qualifiers, placed next to their noun.** A descriptor, list, or clause about a noun sits immediately beside that noun, fenced with paired commas. Wanted: "Each synapse is given a sign, either (i) excitatory or (ii) inhibitory, from its neurotransmitter." Rejected: "Each synapse is given a sign from its neurotransmitter: excitatory or inhibitory." (The rejected form attaches the list to the wrong noun; this defect is called trailing qualifier ambiguity.)
3. **Hypotactic structure over paratactic chaining for sequence or causation.** Subordinate with because, so that, after, although, while, once; do not chain flat clauses with semicolons when one event drives another. Wanted: "Because the correction signal must travel backward through every step, it weakens along the way." Rejected: "The correction signal travels backward; it weakens."
4. **Plain factual statements, no contrastive rhetoric.** No "X, not Y" framing; no "turns X into Y"; no "On its own it is X". We are not debating; state what a thing is and does without positioning it against other concepts.
5. **Supported claims only.** Invented numbers carry an "illustrative" label in the caption or note; real facts carry names and dates and appear again on a final "Sources and further reading" slide.
6. **Enumerations as (i), (ii), (iii)** inside sentences; numbered lists on slides.
7. **Concrete before abstract, step by step.** Walk one specific example through the mechanism (with real numbers a student could recompute) before stating the general rule. A slide that only asserts a summary ("every step is justified by a count") gets replaced by the steps themselves.
8. **Real history, told with its people.** Audiences enjoy the actual history: precursors, disputes, primary quotes, dates. Source every history bit.
9. **Slide titles are plain noun phrases or questions.** Banned title shapes, which humans read as AI voice: "An X that Y" relative clause titles ("A blank that needs old information") and comma apposition tails ("Part 1, on one slide", "Today, in four steps"). Wanted: "Part 1 Overview", "Agreement across seven tokens", "Where do the arrow numbers come from?".
10. **Reveal pedagogy for relations.** Start from the correct, complete example; only then hide the piece under study (a real word becomes a blank); then draw each relation arrow one fragment at a time, so the causality is followable. Never open on the abstracted version of an example.
11. **No paired clause headers or subtitles.** The "An X, and Y" shape is AI phrasing humans avoid, in every position: misconception tags ("A tempting rule, and where it fails"), figure captions, and subtitles ("How the miss is measured, and how millions of weights learn from one mistake"). Compress each one to a single direct phrase or question, worded freshly for its own slide: "An appealing, but wrong, assumption"; "Why not count longer contexts?"; "A scorecard blind to nearness". Title slides carry no subtitle at all: kicker, title, author block, navigation note, nothing else.
12. **Honest names, no crutch words.** A function is called a function ("A loss function is a function that reads the ground truth and the prediction"); do not swap in a vaguer word to dodge repeating a technical term. Watch any single content word from becoming a tic across a deck; "rule" is the documented offender, and its fix is usually the precise term (function, loss, method, choice) rather than a synonym rotation.
13. **Every verb keeps its object; every quantity gets its honest noun.** Slide brevity comes from cutting whole ideas, never from clipping words a sentence needs. Rejected: "Part 3 gave misses a size: the model missed by 1." Wanted: "Part 3 gave mistakes a magnitude: the model missed the right answer by 1." The rejected sentence stacks three clips: a verb stripped of its object (missed what?), a casual noun where the precise one exists (size for magnitude), and a coined house noun ("a miss") doing the work of the plain words mistake and error. Clipped sentences read as scoreboard commentary and leave the reader to fill the holes.
14. **No agentless abstraction.** Reject any sentence in which an abstract noun performs an action with no mechanism: "the past arrives through every handoff", "the memory rides from cell to cell", "training strains the same road" all mean nothing. Every claim names what carries what: "the memory persists and is updated from cell to cell"; "RNNs think one step at a time, relying on memory and the last token". Avoid invented jargon (handoff) when a plain mechanism word (step, update) exists.
15. **Looking back pedagogy.** Students in 2026 arrive with lived experience of asking LLMs hard questions, so classical concepts (classifiers, regression, embeddings, perceptrons) are taught as simplifications carved out of the LLM they already use, never as isolated historical artifacts built bottom up. Open with the lived interaction; strip the big machine down until the classical idea stands alone; work the classical idea concretely; then place it back inside the big machine. History enters as "who built this part first", after the concept is understood.
16. **Standard terminology, no invented metaphors.** Every established concept carries its standard name in the field: softmax, activation function, cross entropy, gradient, linear layer. Invented metaphor vocabulary standing in for a standard term is rejected; documented offenders are "raffle" and "tickets" for softmax, "squash" for activation function, and "bare tables" for stacked linear layers. Readers experience such coinages as weird, and words like "bare" and "squash" carry stray connotations that make a slide sound inappropriate. A concrete analogy may accompany the standard term as a one time explanation, introduced after the term itself; the analogy word never becomes the running name of the concept, and every later mention uses the standard term. The same care applies to adjectives: a descriptor that reads oddly in context ("bare") gets replaced by the precise technical phrase ("linear", "without an activation").
17. **No self narrating scaffolds.** Every sentence talks about its subject; a sentence never takes the prose itself as its subject, in either direction. Forward scaffolds announce structure or delivery before the content: "The two accountings, each in one sentence: ...", "For scale, each with its source: ...", "Read aloud: ...", "Prices, with dates: ...". Backward scaffolds point at the slide's own claims or sections as a collection and pronounce on them: "Both statements are true at once: ...", "Both crafts are one craft: ..." (where "both crafts" names the slide's two columns rather than any thing in the world). The test: replace the sentence's subject with "what this slide just said"; if the sentence still parses, it is a scaffold. Humans never write such framing, so the fix is deletion or substitution of the real things: state content directly, attach each source beside its fact, let a subordinating conjunction carry relations, and when two things share a property, name the things and the property ("A tool description and a skill file are the same kind of writing: documentation for a reader that takes every word literally."). A short label that names the subject itself ("History:", "Documented scale:") stays allowed. The same ban covers **production narration**: the deck, its slides, its build, and its authorship never appear as subject matter ("when the deck was built", "drawn for the deck", "on this slide", "a later slide returns to this"). Honesty about invented content lives in "(illustrative)" labels and on the sources slide, where provenance belongs; forward references name the content ("the token counting of a frontier assistant"), never the slide that carries it. Operational instructions the audience must act on (the title slide's navigation note; "Advance the fragment" beside an animation) are the one sanctioned meta register.
18. **An accepted fix is an example, never a template.** When the owner rewrites one sentence, the rewrite belongs to that one location; only the principle behind it generalizes. After "An appealing, but wrong, assumption" was accepted once, the frame "An appealing, but X, Y" got stamped onto three decks, and the repetition itself became a new tic. Any distinctive phrase frame appears at most once across the whole series.

## Slide visual constructs

- **Never use card grids**: `class="cards"`, stat tile rows, or any grid-of-boxes variant is banned; AI overuses them and readers now recognize them as machine output. Agendas and recaps become numbered lists; growth comparisons become tables; hero numbers live inside prose or a table.
- Single semantic callouts remain allowed: `.q` question frame, `.teach` classroom note, `.mis` misconception panel.
- Formulas are authored as **LaTeX** in `<span class="tex">` (inline) or `<span class="tex disp">` (display) and rendered by **KaTeX embedded inside the deck** (its CSS with woff2 fonts as data URIs, plus katex.min.js and a render pass), so decks stay offline capable; never ASCII math in prose, never a network loaded library. A formula that fails to render gets class `tex-error` and the QA harness reports it; inline math never breaks across lines (`white-space: nowrap`); KaTeX's `body{position:relative}` rule is stripped at embed time because it shifts the stage geometry.
- Diagrams are inline SVG, themed through CSS variables in `style` attributes, with monospace labels (`class="m"`), halo labels (`class="knock"`) kept off the lines they annotate, and an `aria-label`.
- **Mathematical curves are drawn exactly, never as coarse polylines.** A parabola arc is exactly one quadratic Bezier: endpoints on the curve, control point at (vertex x, 2 times vertex y minus the mean of the endpoint y values). Other smooth analytic curves use dense sampling, one point every 10 canvas px or finer. Every plotted marker must land on the drawn curve because both come from the same formula; a visible corner, notch, or misplaced vertex on a smooth curve is a defect.
- Colors: blue marks known or defined things; amber marks the unknown and the current focus; green marks answers and results; red marks wrong paths only. Chart marks use the validated `--mark-*` palette (see the skill rule file for the hex tables).
- **Serif font is Computer Modern** ("CMU Serif", embedded as woff2 data URIs so decks stay offline) for h1, h2, and display math surroundings; formulas render in KaTeX's own Computer Modern derived fonts, embedded the same way. Body text stays a system sans for projection legibility.
- **Schematic user interface mockups carry no bare regions.** A drawn window whose panes are mostly empty fill reads as unfinished artwork, so every pane either shows plausible content or gets removed and the window shrunk to the region actually used. A panel drawn tall enough for eight rows and given one row, with the row parked at the bottom, is the same defect.
- **Mockups obey the continuity of the steps around them.** A window drawn for step 4 never shows a file, button, or extension that step 5 or step 12 tells the audience to create, because a newcomer reading in order hunts for something that does not exist yet.
- **Interface icons are drawn as paths, never as Unicode pictographs.** Characters such as the ones for a branch, a search glass, or a grid render as unrecognizable marks at projection size and vary by font; draw them with `line`, `path`, `circle`, and `rect` at the same centers instead.
- **A caption's count matches the drawn elements.** "The folder appears four times" beside a figure of five boxes sends the audience counting; state the count that the figure supports, or name the element that breaks the pattern.

## The quality system (mandatory checks)

Every deck ships only after all of these pass:

1. **No over-height slides, ever** (including 1080p browser tabs and the Large text setting). The framework's per-slide auto-fit must stay enabled; a slide that needs more than ~8% shrink gets rewritten instead.
2. **Analytic arrow routing.** Straight connectors use `data-connect="idA idB"` so endpoints are computed onto box boundaries; hand placed arrow endpoints are a defect. Curved paths route through empty corridors, under or around content.
3. **Geometry QA harness.** Load the deck with `?qa=1&frag=off`; it appends a JSON report (`<pre id="qa-report">`) listing per slide: content overflow, svg text/text collisions, path/text crossings, arrow endpoints that float or penetrate boxes, and LaTeX render failures (`tex-error`). Read it headlessly via `--dump-dom` and fix every finding or mark intended overlaps with `data-qa="skip"`.
4. **Style audit.** Run the prose audit (strip tags, scan visible text) for dash characters, "X, not Y", "turns into", "on its own". Zero findings required.
5. **Screenshot review.** Headless Edge screenshots of every diagram slide in BOTH themes; look at them. Known artifact: headless Chromium occasionally paints one frame with the svg text layer displaced or dropped while shapes stay put; re-shoot before editing coordinates, and only trust defects that reproduce twice.
6. **Palette validation.** Any new chart color runs through the dataviz skill's `validate_palette.js` against both surfaces.

## Independent review agents (mandatory before delivery)

Every deck revision is reviewed by independent subagents before it ships; their prompts live in [REVIEW_AGENTS.md](REVIEW_AGENTS.md). Run at least (i) the language reviewer and (ii) the geometry and visual reviewer, in parallel, and apply every confirmed finding. Independent eyes catch what the author normalizes: meaningless metaphors, AI voice titles, arcs crossing figures, clipped labels.

## Tooling map

- Skill: `~/.claude/skills/manim` (3brown1blue), slideshow mode in `rules/html-slideshow.md`, framework in `templates/slideshow.html`.
- Decks live in this folder as self contained `.html` files; settings gear menu fixed at the bottom right; PageUp/PageDown for clickers; `#s<n>` deep links; `?theme= &frag= &size= &qa=` overrides.
- Shared settings key `introai_slides_v1` keeps theme and text size consistent across the series.
