skill-generalizer
by derekdylu · 0 users · release 2,
Turn a personal or local skill into a general, privacy-safe version that can be listed on a public skill marketplace. Scans for and removes personal information (names, emails, paths, account IDs, API keys, employer or project details), replaces hard-coded specifics with parameters and placeholders,
SKILL.md
Skill Generalizer
Convert a skill that was written for one person's setup into a skill anyone can use, with no personal information left in it. The output is a new, separate skill. Never modify or overwrite the original; the private version keeps working for its owner.
Two goals pull against each other, so keep both in mind: the result must be safe (nothing private survives, anywhere in the bundle) and still useful (the core workflow that made the skill valuable is preserved, not hollowed out into generic advice).
Inputs
- The source skill: its SKILL.md plus every bundled file (scripts, references, assets). If the user only names it, fetch it with the available skill-reading tool (for example DexDev
get_skill, thenget_skill_filefor each listed file) or read it from disk. Read all files, since private details hide in references and scripts more often than in the main body. - Optional: the marketplace's listing rules, license preference, and target audience. If not given, assume an open, permissive listing and ask only if the choice matters.
Workflow
1. Inventory
List every file in the bundle and the skill's purpose in one sentence. If the purpose itself only makes sense for one person (for example "file my own expense reports with my employer's template"), say so and ask whether to build a reusable version of the underlying task instead of proceeding.
2. Scan, then read manually
Run the bundled scanner on a local copy of the folder:
python scripts/scan_skill.py <skill-folder>
It flags secrets, emails, phone numbers, home-directory paths, IPs, internal hostnames, UUIDs, account-bound tool names (like mcp__Server-12345678__tool), ticket IDs, @handles/#channels, and URLs. Its output masks matches, so it is safe to show the user. If the script cannot be run, do the same pass by reading.
The scanner cannot see plain prose, so also read every file yourself looking for what patterns miss. See references/pii-checklist.md for the full list; the usual culprits are people's names, the user's employer/school/team, client or project codenames, real example data, locations, personal preferences baked in as rules ("always reply in X"), and screenshots or sample files containing real content.
3. Classify each finding
For each item decide one of:
| Action | When |
|---|---|
| Remove | It adds nothing for other users (a signature, a personal note, a stray path). |
| Parameterize | The skill needs some value there, but it varies per user. Replace with a placeholder like <your-project-slug> or a documented input, and tell the skill to ask the user or read it from their environment. |
| Replace with a neutral example | It is an illustrative example. Invent a clearly fictional one (Acme Corp, jane@example.com, /path/to/project). Never lightly alter a real value; fully replace it. |
| Keep | It is genuinely public and generic (official docs URLs, standard tool names, well-known open-source repos). |
Treat secrets differently: any real credential found must be reported to the user as compromised and in need of rotation, since it existed in plain text. Never print it in full, and never write it into the output.
4. Generalize the logic, not just the strings
Removing names is the easy half. Also:
- Replace hard-coded environment assumptions (specific OS, folder layout, tool versions, one company's stack) with detection or a short "adapt to" note.
- Convert personal rules into configurable defaults, for example "Reply in Traditional Chinese" becomes "Use the user's preferred language if known; otherwise match the language of the request".
- Replace account-bound tool references with the tool's generic purpose, for example "use the connected task tracker to ..." with a note on which tools are acceptable.
- Check the workflow for steps that only work with the owner's accounts, subscriptions, or private data. Either make them optional or document the requirement under a Requirements heading.
- Do not widen the skill beyond what it was good at. A generalized skill that claims to do everything triggers wrongly and does nothing well.
5. Rewrite the frontmatter for strangers
name: short, descriptive, lowercase-hyphenated, no personal or company words.description: say what it does and when to use it, in terms a stranger would phrase their request. Include several natural trigger phrases and a "use even if the user doesn't name it" clause, because skills tend to under-trigger. No personal context, no internal jargon.tags: 3 to 6 plain topic words.
6. Add marketplace-ready structure
Make sure the SKILL.md has, in order: a one-paragraph overview, Requirements (tools, accounts, languages, anything the user must have), the Workflow, Inputs and outputs, and Examples using only fictional data. Keep the body under about 500 lines and move depth into references/. Add a license note if the user chose one. Explain why behind rules rather than stacking capitalized commands, so the skill generalizes to cases the author did not foresee.
7. Verify
- Re-run the scanner on the generalized folder. Every remaining finding must be justified (for example public docs URLs) or fixed.
- Re-read every file once more as a stranger would. Ask: could anyone learn who wrote this, where they work, what they own, or what accounts they use?
- Mentally run the skill on a fresh, unrelated example to confirm it still works without the owner's setup. If feasible, test it on a realistic prompt.
- Check bundled assets and filenames, not just text: image content, spreadsheet cells, JSON samples, git metadata, document properties and comments.
8. Deliver
Produce:
- The generalized skill (SKILL.md plus files) as a new skill, using the user's skill tooling. Create it as a draft or private first when that option exists, and publish only after the user approves, because a public listing can be hard to retract.
- A Redaction report in the chat, never inside the skill itself: counts by category (removed, parameterized, replaced, kept), a short list of notable judgment calls, any credentials to rotate, and anything you were unsure about for the user to decide. Describe items by category and location, not by repeating the private values.
Finish by asking the user to review before anything goes on a public market. You can detect many leaks, but only the owner knows what is sensitive to them.
Example
Input skill (private): "Weekly report for Prof. Lin's lab: pulls tickets from LAB-xxx in our tracker, saves to /Users/sam/lab/reports, emails sam@university.edu, replies in Traditional Chinese."
Generalized skill: "Weekly project status report: gather the past week's completed and open tickets from the user's task tracker, summarize by theme, save to a user-specified folder, and optionally draft an email. Language defaults to the user's request language." Placeholders: <project-key>, <output-folder>, <recipient>. Report notes: 1 email, 1 home path, 1 ticket prefix, 1 person name removed; language rule converted to a default.
Edge cases
- The skill's value is private knowledge (a personal knowledge base, someone's writing style, a client playbook). Generalizing it would destroy it. Say so, and offer to publish only the reusable method as a template.
- Third-party content or code: preserve required licenses and attribution; do not strip them as "identifying information".
- Ambiguous names: a word could be a person, a product, or a common noun. When unsure, ask the user instead of guessing in either direction.
- Bundled binaries: if a file cannot be inspected (images, PDFs, archives), either inspect it visually or exclude it and say so.