Here's the project as a plain-text workflow: **Project: Guess a movie's genre from its plot — two ways** **1. Setup** - 58 movies, 6 genres (Sci-Fi, Crime, Romance, Comedy, Horror, Animation) - Each movie has a title, a genre label, and a one-line plot - Both lessons below use the exact same movies — the only thing that changes is whether the genre label is used **2. Lesson 1 — Supervised Learning (learns from answers)** - Show the computer most of the movies *with* their genre labels attached - It studies the patterns: what kind of plot tends to go with each genre - Test it on the movies it hasn't seen yet, genre hidden - Result: it guessed correctly about 4 times out of 5, and correctly tagged a brand-new made-up astronaut plot as "Sci-Fi" **3. Lesson 2 — Unsupervised Learning (finds its own patterns)** - Show the computer the same movies, but *no* genre labels at all — just the plots - It looks for plots that use similar words and groups them together - It ends up with 6 groups, but has no idea what to call any of them - Result: Horror and Crime movies grouped together cleanly (shared words like "dark," "killer"); Comedy and Romance blurred together more (shared lighter, relationship-driven language) **4. The reveal** - Only now do we check the computer's groups against the real genre labels - Score: meaningful overlap, but far from a perfect match - Takeaway: it rediscovered real structure using nothing but word patterns — no labels needed **5. The core difference** | | Supervised | Unsupervised | |---|---|---| | Needs | Labeled examples | Just the raw data | | Does | Predicts a known category | Groups similar things together | | Real-world example | Spam filters, price prediction | Customer segments, topic discovery |
Here's the project as a plain-text workflow: **Project: Guess a movie's genre from its plot — two ways** **1. Setup** - 58 movies, 6 genres (Sci-Fi, Crime, Romance, Comedy, Horror, Animation) - Each movie has a title, a genre label, and a one-line plot - Both lessons below use the exact same movies — the only thing that changes is whether the genre label is used **2. Lesson 1 — Supervised Learning (learns from answers)** - Show the computer most of the movies *with* their genre labels attached - It studies the patterns: what kind of plot tends to go with each genre - Test it on the movies it hasn't seen yet, genre hidden - Result: it guessed correctly about 4 times out of 5, and correctly tagged a brand-new made-up astronaut plot as "Sci-Fi" **3. Lesson 2 — Unsupervised Learning (finds its own patterns)** - Show the computer the same movies, but *no* genre labels at all — just the plots - It looks for plots that use similar words and groups them together - It ends up with 6 groups, but has no idea what to call any of them - Result: Horror and Crime movies grouped together cleanly (shared words like "dark," "killer"); Comedy and Romance blurred together more (shared lighter, relationship-driven language) **4. The reveal** - Only now do we check the computer's groups against the real genre labels - Score: meaningful overlap, but far from a perfect match - Takeaway: it rediscovered real structure using nothing but word patterns — no labels needed **5. The core difference** | | Supervised | Unsupervised | |---|---|---| | Needs | Labeled examples | Just the raw data | | Does | Predicts a known category | Groups similar things together | | Real-world example | Spam filters, price prediction | Customer segments, topic discovery |
Created using ChatSlide
This study explores two approaches to genre learning through a dataset of 58 films across six genres. Supervised learning achieves a 4-in-5 test accuracy, effectively predicting labelled genres, while unsupervised learning uncovers plot patterns, revealing six initial unnamed groups. Notably, horror and crime genres cluster distinctly, whereas comedy and romance show significant overlap. The findings suggest a complementary use of both labelled predictions and raw data groupings for a...