[teamai] Push 1 resource(s) from root

This commit is contained in:
2026-09-10 16:51:10 +08:00
parent f3388c50e2
commit 93626b1d8d
3 changed files with 391 additions and 0 deletions
+1
View File
@@ -0,0 +1 @@
root
+247
View File
@@ -0,0 +1,247 @@
---
name: distill-to-skill
description: >-
Distill knowledge from any source — blog posts, articles, documentation, GitHub repos,
video transcripts, books, papers — into a well-structured agent skill. Use when the user
shares a URL, article, repo, or body of knowledge and wants it turned into a reusable skill.
Triggers on: "make a skill from this", "distill this into a skill", "create a skill from
this article", "turn this repo into a skill", "extract patterns from", "convert to a skill".
This skill complements the `skill-creator` skill — skill-creator handles the mechanics
(frontmatter, packaging, init scripts), this skill handles the distillation process.
---
# Distill to Skill
Turn any source of knowledge into a well-structured agent skill. This is the process
skill — it covers how to extract, filter, restructure, and encode knowledge. For the
mechanical aspects of skill creation (directory structure, frontmatter format, validation,
packaging), use the `skill-creator` skill.
## The Distillation Mindset
A skill is not a summary. It's a **decision-making tool** for an agent working on a task.
When distilling, constantly ask:
- "Would an agent mid-task benefit from knowing this?" → Keep it
- "Is this background context or motivation?" → Cut it
- "Is this specific to one language/framework but the idea is universal?" → Translate it
- "Could an agent figure this out on its own?" → Cut it
- "Does this change how the agent would write code or make decisions?" → Keep it
The goal: an agent loads this skill and immediately writes better code or makes better
decisions, without having read the original source.
## Workflow
### Step 1: Absorb the Source
Read the full source material thoroughly. For repos, explore the architecture, key files,
and patterns. For articles, read end to end. Don't skim — the best insights are often
buried in asides, footnotes, and "by the way" paragraphs.
**For articles/blog posts:**
- Fetch the URL and read the full content
- Note the core thesis (usually 1-2 sentences)
- Identify the actionable rules vs the explanatory prose
- Note any concrete code examples or patterns
**For repos:**
- Explore directory structure, entry points, key modules
- Read the core implementation files (not the tests or config)
- Identify the design patterns, not the specific implementation
- Note the TypeScript/type tricks, architectural decisions, and utility patterns
- Look at what's deliberately *absent* — that's often the most interesting insight
**For multiple sources on a theme:**
- Find the common thread across sources
- Note where sources agree (high-confidence patterns)
- Note where they diverge (context-dependent decisions)
### Step 2: Extract the Transferable Core
Separate the essence from the packaging:
| Keep | Cut |
|---|---|
| Universal principles | Author's personal journey |
| Concrete patterns with code | Motivational framing |
| Decision rules ("when X, do Y") | Background on why the field exists |
| Anti-patterns and pitfalls | Comparisons to other approaches |
| Copy-paste utilities | Historical context |
| Checklists | "Further reading" recommendations |
**The litmus test:** If you removed the original source from existence, would this
skill still be useful on its own? If yes, you've extracted the core correctly.
### Step 3: Decide the Skill Shape
**Single concept, self-contained → SKILL.md only (no references)**
Use when the idea can be fully expressed in ~100-250 lines. The concept is
cohesive enough that splitting it would lose the thread.
Examples from today's work:
- `parse-dont-validate` — One core idea (parse > validate) with practical rules
- `karpathy-guidelines` — A set of behavioral rules
- `agents-md` — How to write one specific file type
**Broad topic with depth → SKILL.md + references/**
Use when there are multiple distinct sub-topics that an agent might need
independently. SKILL.md carries the principles and quick-reference; references
carry the deep dives.
Examples:
- `lean-ts-patterns` — 7 principles in SKILL.md, 5 reference files by domain
- `agent-first-repo` — 3 pillars in SKILL.md, 3 reference files for each pillar
**Multiple independent ideas → Split into separate skills**
If the source contains 2+ concepts that would trigger in different contexts,
make separate skills. They can cross-reference each other.
Example: The OpenAI harness engineering article → split into `agents-md` (how to
write the file) + `agent-first-repo` (broader repo structure) because they trigger
in different contexts.
**Decision heuristic:**
```
Does this source contain one core idea?
YES → Single SKILL.md
NO → Are the ideas used together?
YES → SKILL.md + references/
NO → Separate skills that cross-reference
```
### Step 4: Translate to the User's Ecosystem
The source may be in Haskell, Rust, Go, or plain English. The skill should use
the user's preferred language and ecosystem.
- **Code examples:** Rewrite in the target language (typically TypeScript/Bun)
- **Library references:** Map to the target ecosystem's equivalents
- **Idioms:** Use the target language's patterns (e.g., branded types instead of newtypes)
- **Keep it runnable:** Code in the skill should be copy-pasteable and work
If the original insight is language-agnostic, use the target language for examples
but keep the prose universal.
### Step 5: Structure the Skill
Follow this template for SKILL.md:
```markdown
---
name: skill-name
description: >-
[What this enables]. [When to use it — specific scenarios].
Triggers on: [concrete trigger phrases].
---
# Title
[1-2 line summary. Source attribution if from a specific article/repo.]
## [Core Concept / Principles]
[The distilled rules. Concise. Imperative voice. Code examples inline.]
## [Practical Patterns / Copy-Paste Code]
[Things the agent can use immediately. Concrete, not abstract.]
## [Anti-Patterns / What to Avoid]
[Common mistakes. What NOT to do is often more valuable than what to do.]
## [Checklist / Code Review Guide]
[Verification points. Things to check when reviewing code.]
```
**For reference files**, each should:
- Start with a 1-2 line summary of what it covers
- Be self-contained — readable without SKILL.md for context
- Include the relevant companion skill name-drops (not full content)
- Stay under ~300 lines
### Step 6: Write the Description (Most Important Line)
The YAML `description` field is the **only thing** that determines whether the skill
triggers. It's loaded into context permanently. Write it carefully:
- Start with what the skill enables (not what it is)
- List specific scenarios and file types
- Include concrete trigger phrases the user might say
- Keep it to 3-5 lines of YAML
Bad: `"Patterns from a blog post about types."`
Good: `"Type-driven design: transform unstructured data into precise types at
system boundaries. Use when writing input validation, designing data types, or
reviewing code with redundant null checks. Triggers on: 'parse don't validate',
'make illegal states unrepresentable', 'input validation'."`
### Step 7: Validate
Before finishing, check:
- [ ] Could an agent use this skill without reading the original source?
- [ ] Is every section actionable (rules, patterns, code) not explanatory (history, motivation)?
- [ ] Are code examples in the user's preferred language and copy-pasteable?
- [ ] Is SKILL.md under ~300 lines? (Move depth to references/ if over)
- [ ] Does the description include concrete trigger phrases?
- [ ] Is there a checklist or code review guide for verification?
- [ ] Are companion skills referenced by name (not duplicated)?
- [ ] Would removing any section make the skill less useful? If not, cut it.
## Distillation Patterns
### The Inversion
Many articles explain bottom-up: problem → exploration → solution.
Skills should be top-down: **rule → example → anti-pattern**.
The agent doesn't need to be convinced. It needs to know what to do.
### The Translation
Academic/theoretical sources often use abstract examples. Translate to concrete,
real-world scenarios in the user's domain:
- "NonEmpty list" → `[T, ...T[]]` tuple type in TypeScript
- "Sum types" → discriminated unions with `kind` field
- "Smart constructor" → branded type with parse function
- "Monad" → async pipeline / Result type
### The Compression
A 5,000-word article typically distills to ~150-250 lines of skill. The compression
ratio is roughly 10:1 to 20:1. If your skill is approaching the same length as the
source, you're summarizing, not distilling.
### The Cross-Reference
When distilling a source that touches on ideas already captured in other skills,
don't re-explain — reference. Write 2-3 sentences of context for how the idea
applies here, then point to the companion skill for depth.
```markdown
Parse data at system boundaries into precise types — don't let raw/untyped data
flow deep into business logic. For the full treatment of branded types, smart
constructors, and the shotgun parsing anti-pattern, see the `parse-dont-validate` skill.
```
## Source-Specific Tips
**Blog posts:** Usually one core idea with 60% motivation, 30% examples, 10% actionable rules. Extract the 10%, expand it with your own examples.
**GitHub repos:** The code IS the content. Focus on architectural patterns, utility functions worth copying, TypeScript tricks, and what's deliberately absent. Ignore CI config, test infrastructure, and build tooling unless that's the point.
**Documentation:** Already structured, but optimized for lookup, not for decision-making. Restructure around "when to use X" rather than "what X does."
**Papers:** High insight density but buried in formalism. Extract the key theorem/insight, translate to practical code patterns, drop the proofs.
**Video transcripts:** Extremely low density. Scan for the 2-3 key moments where the speaker says something prescriptive, ignore the rest.
For concrete before/after examples of distillation, see
[references/examples.md](references/examples.md).
@@ -0,0 +1,143 @@
# Distillation Examples
Concrete before/after examples showing how source material becomes skill content.
## Example 1: Blog Post → Single SKILL.md
**Source:** "Parse, Don't Validate" by Alexis King (~5,000 words, Haskell)
**What the article contains:**
- Motivation: why `head :: [a] -> a` is partial (700 words)
- Two approaches: weaken output vs strengthen input (1,500 words)
- The `NonEmpty` list example with full Haskell code (800 words)
- "What is a parser?" philosophical discussion (600 words)
- Practical advice section (800 words)
- Recap and related reading (600 words)
**What the skill extracted:**
- The core idea in 6 lines (validate returns void, parse returns proof)
- The two strategies as a decision rule: "Try strategy 2 first. Fall back to 1."
- 7 practical rules, each with a TypeScript code example
- The shotgun parsing anti-pattern (3 sentences, not 3 paragraphs)
- A code review checklist (9 concrete smells)
**What was cut:**
- The entire Haskell-specific `NonEmpty` walkthrough → replaced with TS `[T, ...T[]]`
- "What is a parser?" philosophical section → collapsed to 1 sentence
- All motivation/persuasion → the agent doesn't need to be sold
- Recap and related reading → not actionable
- Footnotes about type theory → too academic
**Compression:** ~5,000 words → 210 lines (~750 words). Ratio: ~7:1
---
## Example 2: Article + Exemplary Doc → Workflow Skill
**Source:** matklad's "ARCHITECTURE.md" article (~800 words) + rust-analyzer's
architecture.md (~420 lines) as a concrete example
**What the article contains:**
- Why architecture docs matter (contributor 10x cost) (200 words)
- The rules: short, stable, codemap, name don't link, invariants, boundaries (400 words)
- Link to rust-analyzer as example (200 words)
**What the skill extracted:**
- 7 principles distilled from the article prose
- A 3-step workflow (explore → identify → write) that the article implies but doesn't state
- A concrete template with `### \`path/\`` headers, **Boundary:** and **Invariant:** callouts
- Style rules derived from studying the rust-analyzer example
- A quality checklist
**What was added (not in the source):**
- The workflow — the article says "what" but not "how an agent should do it"
- The template — extracted by studying rust-analyzer's structure
- The reference example — a generic TypeScript project demonstrating all patterns
**Key insight:** The article was 800 words of principles. The exemplary doc was 420
lines of practice. The skill bridged the two: principles + template + workflow.
---
## Example 3: Multiple Repos → Themed Skill with References
**Source:** 7 GitHub repos (citty, consola, ofetch, defu, scule, pathe, taze)
**What the repos contain:** ~15,000+ lines of source code across 7 repositories
**The distillation process:**
1. Explored each repo independently, documenting patterns
2. Identified the **common thread**: zero-dep, lightweight, TypeScript-first
3. Found 7 shared principles across all repos
4. Grouped patterns by domain: CLI, logging, fetch, data utils, TS tricks
5. Extracted copy-paste utilities (ANSI colors, isPlainObject, etc.)
**What the skill contains:**
- SKILL.md (193 lines): 7 principles, 4 copy-paste patterns, quick reference table
- 5 reference files (192-349 lines each): deep dives by domain
**What was kept:**
- Architectural patterns (factory over classes, one primitive compose everything)
- Copy-paste utilities under 25 lines
- TypeScript type tricks that are non-obvious
- Design decisions (why retries default to 0 for POST)
**What was cut:**
- Build configuration, CI setup, test infrastructure
- Implementation details specific to each repo's domain
- Anything that only makes sense in the context of that specific library
- Code that depends on those libraries' internal types
**Compression:** ~15,000 lines of code → 1,519 lines of skill. Ratio: ~10:1
---
## Example 4: Long-Form Article → Umbrella Skill + Companions
**Source:** OpenAI "Harness Engineering" article (~3,000 words)
**The decomposition decision:**
The article contained 5+ distinct ideas that trigger in different contexts:
1. How to write AGENTS.md → triggers when creating agent instruction files
2. Repo structure for agents → triggers when setting up new projects
3. Progressive disclosure → sub-topic of repo structure
4. Mechanical enforcement → sub-topic of repo structure
5. Entropy management → sub-topic of repo structure
Ideas 1 and 2 trigger independently (different user intents), so they became
separate skills. Ideas 3-5 are always needed in the context of idea 2, so they
became reference files within the `agent-first-repo` skill.
**The cross-reference pattern:**
The article also referenced two external concepts:
- matklad's ARCHITECTURE.md → already a skill (`architecture-md`)
- "Parse, don't validate" → already a skill (`parse-dont-validate`)
Rather than duplicating those skills' content, `agent-first-repo` includes
2-3 sentences of contextualized summary + a name-drop pointing to the companion
skill. This keeps each skill lean and avoids content drift between copies.
**Result:**
- `agents-md` — standalone, 184 lines
- `agent-first-repo` — 156 lines SKILL.md + 3 references (505 lines)
- Cross-references to `architecture-md` and `parse-dont-validate`
---
## The Pattern
Across all examples, the distillation process follows the same shape:
```
Source material (broad, explanatory, motivational)
↓ Extract transferable principles
↓ Cut motivation, history, persuasion
↓ Translate to user's language/ecosystem
↓ Add structure: rules → examples → anti-patterns → checklist
↓ Decide shape: single file / with references / split skills
↓ Cross-reference companions, don't duplicate
Skill (narrow, imperative, actionable)
```
The compression ratio is consistently 7:1 to 20:1. If your skill is longer than
1/5th the source material, you're likely summarizing rather than distilling.