Install tested capabilities in your AI agent — and keep them safe and up to date.
Connect Claude, Hermes or ChatGPT once, then give it a ready-made corporate agent — or describe the role you need and we build one for you. Everything is tested against jailbreaks, scored in public and kept patched. No code, no fine-tuning.
Free tier for public skills · No signup to browse · or build a custom agent
Already use MCP? Paste this URL into your client. Some clients need it in a config file plus one restart.
https://superagentskill.com/api/public/mcp- Tested skills & playbooks
Scored on format and substance, adversarially probed, published with the score attached.
- Ready-made corporate agents
33 roles with soul, guardrails and skills — deploy one in minutes.
- Your own agent, built for you
Describe the role; the factory assembles and tests it before you download.
- Expert skills
- 0+
- Playbooks
- 0+
- Souls (expert personas)
- 0+
- Setup time
- 0s
You write skills, souls and prompts for your agent.But you never test them.
Don't trust a capability blindly. A badly written skill quietly degrades the model it was supposed to improve — and you find out from your customers.
Does this skill actually improve the model?
never tested
Is this version better than the last one?
no versions on file
Did anyone run adversarial cases against it?
never ran them
Will it hold up in a different harness?
no way to know
Can it recover when a tool call fails?
nothing like that
Can you prove this is the best version?
no benchmark exists
You shipped it anyway.
Every capability is scored, repaired and re-scored before it reaches your agent.
Upload a skill or describe an agent. The lab grades it, fixes what fails, saves each attempt as a version, and only ships the one that passes.
Every version gets a Trust Score.
The lab builds the evaluation cases, runs them, and scores format, substance, safety and schema validity into one number. Below the bar, the skill stays in the lab.
Illustrative scale
Test. Repair. Re-score.
Watch a grade-F skill get taken apart and rebuilt. Each pass becomes a new version, and a version only survives if it scores higher than the one before it.
Base agent vs lab-tested agent.
No capability ships on trust. Here is the difference an A-grade skill makes across the metrics that decide whether you can put an agent in front of a customer.
A random skill vs. a SAK A-grade skill
The same task, two skills: one ungraded prompt from the wild, one certified A on SuperAgent Skill. Figures below are certification targets, illustrative of the gap the SAK pipeline is designed to produce — each one flips to a live, measured number as the paired-benchmark telemetry reaches sample-size thresholds.
Want to see the same delta on your agent?
Track pass rate, human intervention, latency, and ROI per skill in the SAK dashboard. Join the access waitlist or see your numbers now.
ROI access requires a SAK account. New users start on the Free plan.
Illustrative certification targets, not yet measured medians. Comparative numbers are produced by our paired adversarial benchmark (same suite, certified skill vs. raw baseline prompt, Wilson 95% lower bounds, Ed25519-signed results) and replace these figures as sample thresholds are met — every skill page shows its live numbers today.
Every version, on file.
Each review is stored with its score delta, so you can see exactly which change moved the number — and prove it later to a customer, an auditor or your own team.
- v127F
- v241D
- v358C
- v467C
- v574B
- v683B
- v791A ✓
Illustrative history
Run it anywhere, in one line.
Install through the MCP endpoint or download the file. The same tested capability runs in every major agent harness, with any model.
https://superagentskill.com/api/public/mcpOpen ecosystem
Works with the open agent skills ecosystem
Our catalog ships in the standard SKILL.md format, so you can install it with the open skills.sh CLI, with our own CLI, or over MCP. Same skills, whichever route your agent prefers.
npx skills add criptogus/agent-evolve-networknpx skills updateInstalls into
- Claude Code
- Cursor
- Codex
- GitHub Copilot
- Windsurf
- Gemini CLI
- Cline
- Zed
- OpenCode
- Antigravity
- Goose
- Kiro CLI
- Roo
- Trae
- Droid
- Amp
- VS Code
Open Skills CLI
One command drops our SKILL.md files into whichever agent you already use. No account, no config file.
MCP server
Paste one URL for always-current graded versions, plus review, diagnosis and before/after proof tools.
Trust Score on top
Every skill carries a public grade: format, substance and adversarial testing, with the evidence attached.
The open ecosystem gives skills distribution. We add the part that decides whether you should install one: a graded Trust Score, an adversarial pass rate, and a before/after report when a skill is improved. Browse the graded catalog.
Corporate Agent Factory
A whole org chart of agents — installed, not prompted.
Download a curated executive or specialist agent, or describe a role you cannot hire fast enough and the factory builds it for you — always anchored on the current state of the art, always scored before delivery.
Ready to install today
Browse the full Agent Store →Or build your own — from a prompt
- 01
Describe the role
One brief: the role, your company, the outcomes it owns and your hard constraints.
- 02
We research the state of the art
The factory pulls the frameworks real operators use for that role — not generic prompt filler.
- 03
Soul, skills, playbooks, guardrails
An operating identity, 5 bounded skills, 3 step-by-step playbooks and enforceable guardrails.
- 04
Scored and repaired to grade A
The agent is audited on specificity, decision quality and safety, then rewritten until it clears the bar.
Custom agents are included in Agent Pass. Download as ZIP, single markdown file, or install straight into your agent over MCP.
SAK University
A store sells you a skill. A university figures out which skill you need.
Installing a capability by name is a guess — and an agent with 40 skills gets worse, not better. The University measures the agent first, points out the error class blocking the result, and prescribes the next step. Free and anonymous.
1. Admission exam
Up to 168 fixed tasks across 21 corporate domains, run by the agent itself. Part of it is holdout, so you can't train for the test.
2. Diagnosis by error class
Not “62 out of 100 in sales,” but: abandons ambiguity in 58% of cases, breaks the output contract in 31%. That's actionable.
3. Prescription by marginal gain
The next capability that moves the needle the most for this agent — respecting prerequisites, conflicts, and context budget.
diagnose_start → curriculum_nextSkills are systems, not prompts.
When an agent underperforms, the reflex is to reach for a bigger model or start fine-tuning. In most of the cases we see, neither is what was missing. What was missing was structure: a capability with explicit boundaries, a defined output contract, and guardrails that hold when someone tries to talk the agent out of them.
The uncomfortable part is that almost nobody tests that structure. People write a SKILL.md once, paste it into their agent, and ship. There is no score, no version history, no adversarial pass — so when quality drops, there is nothing to compare against and nothing to roll back to.
Super Agent Skill exists to turn that into evidence. Every capability that enters the registry is graded on format and substance separately, run against adversarial cases, repaired where it fails, and re-scored. Each attempt is kept as a version with its delta. Only the version that passes gets published — and the score travels with it, so anyone installing it can see what it was measured on.
That is also how we build agents. You describe the role; the factory assembles the soul, the guardrails and the skills, then puts the whole thing through the same lab before you can download it.
Capabilities are the most underused layer of the AI stack. Tested capabilities are rarer still. That is the layer we are building, in the open, with the scores published.
Your proprietary skill stays yours.
We don't copy, train on, or resell your code. Every upload is signed, scoped to your workspace, and evaluated in an isolated sandbox.
Need a private registry? Enterprise teams keep every skill inside their own workspace with SSO, audit logs and a custom DPA/NDA.
One plan. Everything included.
No tiers to compare and no add-ons. Browsing and installing public capabilities stays free forever — Pro is for when you want your own capabilities tested, built and kept current.
- Unlimited skill reviews with full graded reports
- The Agent Factory — build a custom corporate agent from a prompt
- The Agent Store — 33 ready-made agents with soul, guardrails and skills
- SAK University — diagnosis, adaptive curriculum, residency, credentials
- Continuous adversarial re-testing on everything you installed
- Private packages, signed releases and mutual NDA on request
Solo or team. Cancel in one click. Everything you create stays yours.
Instead of $228 / year billed monthly at $19, which stays available if you prefer the flexibility.
Get Pro →Browse the free registryNeed a private registry, SSO or audit logs? Enterprise.
Questions, answered straight.
Everything people ask before connecting an agent.
Something else? hello@superagentskill.com
One MCP URL.
One sentence. Done.
Install the best packages of your industry, generate what doesn't exist yet, and let SkillForge ship better versions for you — week after week.
https://superagentskill.com/api/public/mcp