Here's the project as a plain-text workflow: **Project: Guess a movie's genre from its plot — two ways** **1. Setup** - 58 movies, 6 genres (Sci-Fi, Crime, Romance, Comedy, Horror, Animation) - Each movie has a title, a genre label, and a one-line plot - Both lessons below use the exact same movies — the only thing that changes is whether the genre label is used **2. Lesson 1 — Supervised Learning (learns from answers)** - Show the computer most of the movies *with* their genre labels attached - It studies the patterns: what kind of plot tends to go with each genre - Test it on the movies it hasn't seen yet, genre hidden - Result: it guessed correctly about 4 times out of 5, and correctly tagged a brand-new made-up astronaut plot as "Sci-Fi" **3. Lesson 2 — Unsupervised Learning (finds its own patterns)** - Show the computer the same movies, but *no* genre labels at all — just the plots - It looks for plots that use similar words and groups them together - It ends up with 6 groups, but has no idea what to call any of them - Result: Horror and Crime movies grouped together cleanly (shared words like "dark," "killer"); Comedy and Romance blurred together more (shared lighter, relationship-driven language) **4. The reveal** - Only now do we check the computer's groups against the real genre labels - Score: meaningful overlap, but far from a perfect match - Takeaway: it rediscovered real structure using nothing but word patterns — no labels needed **5. The core difference** | | Supervised | Unsupervised | |---|---|---| | Needs | Labeled examples | Just the raw data | | Does | Predicts a known category | Groups similar things together | | Real-world example | Spam filters, price prediction | Customer segments, topic discovery |