Thumbnail frameworks — the concept layer
This is the "what to depict" layer; the SKILL.md house prompt structure is "how to render it". Load this in EXECUTE step 1, pick the framework(s), then assemble the house prompt.
Every thumbnail must open an information gap — the image raises a question the title / video answers. Before writing any prompt, brainstorm ≥5 concept options across the frameworks below and carry the strongest into the blocks; a concept often combines 2+ frameworks (e.g. Posed Portrait + Map + Landscape) when it deepens the gap without clutter.
Truthfulness law: the image may exaggerate but must honestly represent the video. A Social-UI / News-Clip / amplification that misrepresents the video breaks viewer trust — keep any baked text short and true.
Most frameworks map straight onto the house-structure blocks; the ones needing readable in-image text use the BAKED UI clause in the Text contract (block 3) and stay brand-generic (no real logos / networks).
| # | Framework | Realize it with |
|---|---|---|
| 1 | Before/After Transformation | Split frames before/after mode — max contrast between the two states of one subject |
| 2 | Social UI (tweet / DM / review) | KEY ELEMENTS = a generic chat bubble / DM row / star-review card beside the subject; short message via BAKED UI; NO real platform brand |
| 3 | Three-Step Progression | Split frames plain/custom, 3 vertical panels: start → mid story-beat → end; optional per-panel number/Day badge |
| 4 | Compelling Screenshot | NOT a generation — use a real frame from the source video, or a Posed Action Shot that matches the video's first frame. Flag when the user has a source clip |
| 5 | Posed Portrait | The default: SUBJECT rendered LARGE (fills much of the frame) + Identity Lock + YouTube rig + Emotion; very little or nothing in the background — the subject IS the focus |
| 6 | Posed Action Shot | Scene brief = one intriguing action mid-moment; COMPOSITION uncrowded, single hero action |
| 7 | Highlighting a Specific Day | Any framework + a big DAY N badge (BAKED UI); pick a day in the last ~20% of the arc |
| 8 | Graphical Representation | Non-photoreal: REPLACE the Frame contract with a clean diagram/graphic brief (bell curve, simple familiar chart); drop the photoreal + lighting-rig blocks |
| 9 | Landscape | Environment is the hero; keep ONE small subject on a power third; lighting rig optional |
| 10 | Map/Aerial | Frame = a map or aerial photo; KEY ELEMENTS = a highlighted route / point (circle, arrow, label via BAKED UI) |
| 11 | Product | Product is the SUBJECT (hero, sharp, its own label exempt per Angle Lock); open the gap by making the product the answer to the title's question |
| 12 | Adding Text | Text policy overlay (default) or baked TEXT; use it as a callout (arrow + word) or as the continuation / answer to the title's question |
| 13 | Repetition of Objects | KEY ELEMENTS = a large quantity of ONE object filling the frame, still legible; add one context element + a subject for scale |
| 14 | Size Difference | COMPOSITION = extreme scale contrast between two elements central to the story (giant vs tiny) |
| 15 | News Clip | Generic "breaking news" lower-third / chyron over the image (BAKED UI, short truthful line); generic broadcaster styling, NO real network |
| 16 | Amplified Reality | Posed Action Shot + exaggerate ONE real element from the story (KEY ELEMENTS oversized); keep it plausible — over-amplification kills trust |
Frameworks 1, 5, 12 are already the skill's defaults; the rest are selected here and realized through the existing house-structure blocks — do not invent a parallel renderer.
The 10 rules cheat-sheet (why these work)
The frameworks are vehicles for the underlying rules — apply them regardless of framework: one clear focal subject, a strong information gap, high contrast and readability at ~120px, emotion or intrigue on any face, no clutter, truthful to the video, and a concept that reads in under a second in a crowded feed.