Corporate agent factory · 459+ signed skills

Install tested capabilities in your AI agent — and keep them safe and up to date.

Connect Claude, Hermes or ChatGPT once, then give it a ready-made corporate agent — or describe the role you need and we build one for you. Everything is tested against jailbreaks, scored in public and kept patched. No code, no fine-tuning.

Free tier for public skills · No signup to browse · or build a custom agent

GitHub starsEvery skill adversarially tested before publish

Already use MCP? Paste this URL into your client. Some clients need it in a config file plus one restart.

https://superagentskill.com/api/public/mcp
What you install
  • Tested skills & playbooks

    Scored on format and substance, adversarially probed, published with the score attached.

  • Ready-made corporate agents

    33 roles with soul, guardrails and skills — deploy one in minutes.

  • Your own agent, built for you

    Describe the role; the factory assembles and tests it before you download.

See what is in the registry →
Expert skills
0+
Playbooks
0+
Souls (expert personas)
0+
Setup time
0s
The problem

You write skills, souls and prompts for your agent.But you never test them.

Don't trust a capability blindly. A badly written skill quietly degrades the model it was supposed to improve — and you find out from your customers.

Does this skill actually improve the model?

never tested

Is this version better than the last one?

no versions on file

Did anyone run adversarial cases against it?

never ran them

Will it hold up in a different harness?

no way to know

Can it recover when a tool call fails?

nothing like that

Can you prove this is the best version?

no benchmark exists

You shipped it anyway.

The lab

Every capability is scored, repaired and re-scored before it reaches your agent.

Upload a skill or describe an agent. The lab grades it, fixes what fails, saves each attempt as a version, and only ships the one that passes.

01Measure

Every version gets a Trust Score.

The lab builds the evaluation cases, runs them, and scores format, substance, safety and schema validity into one number. Below the bar, the skill stays in the lab.

Trust Scoreship threshold 90
No skill52
First draft60
Lab-tested version91

Illustrative scale

02The loop

Test. Repair. Re-score.

Watch a grade-F skill get taken apart and rebuilt. Each pass becomes a new version, and a version only survives if it scores higher than the one before it.

skill.upgrade
F → A in one pass
1/6Connect the MCP
Your MCP endpoint · works in Claude
One URL — no SDK, no keys, no DevOps. Then just ask your agent to review a skill.
Works with Claude Hermes ChatGPT Cursor Lovable
03Benchmark

Base agent vs lab-tested agent.

No capability ships on trust. Here is the difference an A-grade skill makes across the metrics that decide whether you can put an agent in front of a customer.

The A-grade differenceIllustrative targets

A random skill vs. a SAK A-grade skill

The same task, two skills: one ungraded prompt from the wild, one certified A on SuperAgent Skill. Figures below are certification targets, illustrative of the gap the SAK pipeline is designed to produce — each one flips to a live, measured number as the paired-benchmark telemetry reaches sample-size thresholds.

Ungraded skillF · no proof

Copy-pasted prompt, no adversarial testing, unsigned release.

SAK A-grade skillA · certified

Adversarial harness passed, Ed25519-signed, re-tested daily by SkillForge.

Outcome per execution — last 90 days

Not pass rate: the outcome actually recorded on every production run on the SAK network. Same task, one random skill from the wild, one certified A-grade skill.

Random skillSAK · A-grade
Random skillF · no proof
Completion rate
of runs finish the task
42%
Human intervention
of runs need a rescue
49%
Latency saved
vs. workspace baseline
0s
Tokens saved
18,400 avg. per task
0
SAK · A-gradeA · certified
Completion rate
of runs finish the task
94%
Human intervention
of runs need a rescue
4%
Latency saved
p95 42s → 6.8s
−35.2s
Tokens saved
4,100 avg. per task
−14,300

How these numbers are measured

Every execution on the SAK network reports the same outcome fields: task completion, human intervention, end-user rating, retry count, latency vs. baseline, and tokens consumed. We pair the same task across two skills — one ungraded skill found in the wild, one SAK A-grade certified skill — then aggregate the last 90 days of production runs. The chart shows medians; the column cards show the headline deltas visitors feel first.

  • 01Same task. Identical prompt, toolset and success criteria for both skills.
  • 02Per-execution. Outcomes are recorded on every run, not sampled or inferred.
  • 0390-day median. Rolling window removes outliers and single-day spikes.
+52 p.p.
more tasks completed
−92%
human rescues needed
2.4×
more positive ratings
View chart data as a table
Outcome per execution, last 90 days (%)
OutcomeRandom skillSAK · A-grade
Task completed42%94%
No human intervention51%96%
Positive rating38%91%
First-try success29%87%
Under latency budget44%93%
Task success rate after install

Weekly median, first 12 weeks in production. Ungraded skills drift as models change; SAK skills climb — SkillForge re-tests and patches them automatically.

Ungraded skillSAK · A-grade
94% vs 38% at W12
Adversarial attempts blocked, by attack class

Share of red-team attempts the skill refused or neutralized. Every SAK skill must survive this gauntlet before it can publish — pass rates are public.

Ungraded skillSAK · A-grade
≥98.7% every class
View chart data as a table
Task success rate by week (%)
WeekUngradedSAK · A-grade
W146%88%
W245%89%
W344%89%
W444%91%
W543%91%
W642%92%
W742%92%
W841%93%
W940%93%
W1039%94%
W1139%94%
W1238%94%
Adversarial attempts blocked (%)
Attack classUngradedSAK · A-grade
Prompt injection18%99.2%
Jailbreak24%98.7%
Data exfiltration31%99.6%
Policy bypass22%98.9%
  • Task success rate
    Real-world executions that complete without human rescue
    Ungraded
    42%
    SAK · A
    94%
    +124%
  • Prompt-injection resistance
    % of adversarial attempts blocked by the skill
    Ungraded
    18%
    SAK · A
    99.2%
    +451%
  • Hallucination rate
    Fabricated facts, tools or citations per 1k runs
    Ungraded
    31%
    SAK · A
    1.4%
    −95%
  • Avg. tokens per task
    Wasted context = wasted money on every call
    Ungraded
    18,400
    SAK · A
    4,100
    −78%
  • p95 latency
    How long the slowest 5% of calls take
    Ungraded
    42s
    SAK · A
    6.8s
    −84%
  • Cost per 1,000 runs
    Blended LLM + retry + human-in-the-loop cost
    Ungraded
    $38.20
    SAK · A
    $4.90
    −87%
  • PII / secret leakage
    Sensitive strings surfaced in outputs or logs
    Ungraded
    6.1%
    SAK · A
    0.02%
    −99.7%
  • Mean time to patch a CVE
    From disclosure to a signed, verified release
    Ungraded
    27 days
    SAK · A
    under 24h
    −96%
2.2×
more tasks completed per agent per day (target)
$33/1k runs
saved on LLM spend at the same volume
300× fewer
leaks & jailbreaks reaching production (target)

Want to see the same delta on your agent?

Track pass rate, human intervention, latency, and ROI per skill in the SAK dashboard. Join the access waitlist or see your numbers now.

ROI access requires a SAK account. New users start on the Free plan.

Illustrative certification targets, not yet measured medians. Comparative numbers are produced by our paired adversarial benchmark (same suite, certified skill vs. raw baseline prompt, Wilson 95% lower bounds, Ed25519-signed results) and replace these figures as sample thresholds are met — every skill page shows its live numbers today.

04Audit trail

Every version, on file.

Each review is stored with its score delta, so you can see exactly which change moved the number — and prove it later to a customer, an auditor or your own team.

review history
  • v127F
  • v241D
  • v358C
  • v467C
  • v574B
  • v683B
  • v791A

Illustrative history

05Export

Run it anywhere, in one line.

Install through the MCP endpoint or download the file. The same tested capability runs in every major agent harness, with any model.

ClaudeCodexHermesChatGPTCursorClineAny MCP client
https://superagentskill.com/api/public/mcp
See the setup for my client →

Open ecosystem

Works with the open agent skills ecosystem

Our catalog ships in the standard SKILL.md format, so you can install it with the open skills.sh CLI, with our own CLI, or over MCP. Same skills, whichever route your agent prefers.

$npx skills add criptogus/agent-evolve-network
$npx skills update

Installs into

  • Claude Code
  • Cursor
  • Codex
  • GitHub Copilot
  • Windsurf
  • Gemini CLI
  • Cline
  • Zed
  • OpenCode
  • Antigravity
  • Goose
  • Kiro CLI
  • Roo
  • Trae
  • Droid
  • Amp
  • VS Code

Open Skills CLI

One command drops our SKILL.md files into whichever agent you already use. No account, no config file.

MCP server

Paste one URL for always-current graded versions, plus review, diagnosis and before/after proof tools.

Trust Score on top

Every skill carries a public grade: format, substance and adversarial testing, with the evidence attached.

The open ecosystem gives skills distribution. We add the part that decides whether you should install one: a graded Trust Score, an adversarial pass rate, and a before/after report when a skill is improved. Browse the graded catalog.

Corporate Agent Factory

A whole org chart of agents — installed, not prompted.

Download a curated executive or specialist agent, or describe a role you cannot hire fast enough and the factory builds it for you — always anchored on the current state of the art, always scored before delivery.

Ready to install today

Browse the full Agent Store →

Or build your own — from a prompt

  1. 01

    Describe the role

    One brief: the role, your company, the outcomes it owns and your hard constraints.

  2. 02

    We research the state of the art

    The factory pulls the frameworks real operators use for that role — not generic prompt filler.

  3. 03

    Soul, skills, playbooks, guardrails

    An operating identity, 5 bounded skills, 3 step-by-step playbooks and enforceable guardrails.

  4. 04

    Scored and repaired to grade A

    The agent is audited on specificity, decision quality and safety, then rewritten until it clears the bar.

Custom agents are included in Agent Pass. Download as ZIP, single markdown file, or install straight into your agent over MCP.

SAK University

A store sells you a skill. A university figures out which skill you need.

Installing a capability by name is a guess — and an agent with 40 skills gets worse, not better. The University measures the agent first, points out the error class blocking the result, and prescribes the next step. Free and anonymous.

1. Admission exam

Up to 168 fixed tasks across 21 corporate domains, run by the agent itself. Part of it is holdout, so you can't train for the test.

2. Diagnosis by error class

Not “62 out of 100 in sales,” but: abandons ambiguity in 58% of cases, breaks the output contract in 31%. That's actionable.

3. Prescription by marginal gain

The next capability that moves the needle the most for this agent — respecting prerequisites, conflicts, and context budget.

Take the admission exam View the adaptive trackAlso available via MCP: diagnose_start curriculum_next
An open letter from the founder

Skills are systems, not prompts.

When an agent underperforms, the reflex is to reach for a bigger model or start fine-tuning. In most of the cases we see, neither is what was missing. What was missing was structure: a capability with explicit boundaries, a defined output contract, and guardrails that hold when someone tries to talk the agent out of them.

The uncomfortable part is that almost nobody tests that structure. People write a SKILL.md once, paste it into their agent, and ship. There is no score, no version history, no adversarial pass — so when quality drops, there is nothing to compare against and nothing to roll back to.

Super Agent Skill exists to turn that into evidence. Every capability that enters the registry is graded on format and substance separately, run against adversarial cases, repaired where it fails, and re-scored. Each attempt is kept as a version with its delta. Only the version that passes gets published — and the score travels with it, so anyone installing it can see what it was measured on.

That is also how we build agents. You describe the role; the factory assembles the soul, the guardrails and the skills, then puts the whole thing through the same lab before you can download it.

Capabilities are the most underused layer of the AI stack. Tested capabilities are rarer still. That is the layer we are building, in the open, with the scores published.

Gustavo Caetano
Founder · superagentskill.com · @gustavocaetano
How the scoring works →
Security & IP Promise

Your proprietary skill stays yours.

We don't copy, train on, or resell your code. Every upload is signed, scoped to your workspace, and evaluated in an isolated sandbox.

You keep full ownership

Your source code, prompts and context remain your property. We claim no rights to proprietary skills you upload or generate inside your workspace.

No unauthorized training

We never use submitted skills, docs or transcripts to train shared models, create competing products, or build look-alike packages.

Isolated evaluation sandbox

Certification runs inside a sandboxed environment. Other users can't peek at your code, and test artifacts are purged after scoring.

Signed, verifiable releases

Every published artifact is Ed25519-signed. Buyers see metadata and Trust Score — not your source — unless you explicitly open-source it.

Row-level access control

Database access is scoped per user via RLS. Your private packages, billing data and execution logs are invisible to other accounts by default.

Encrypted in transit & at rest

TLS protects data in flight. Sensitive storage is encrypted at rest. Audit logs record every package access and administrative action.

Exportable audit trail

Review who accessed what, when. Enterprise plans include full audit-log exports and tamper-evident signing history for compliance.

Private registry & DPA/NDA

Enterprise teams can run a fully private registry with SSO, custom contracts and a Data Processing Agreement. Your IP never touches the public marketplace.

Need a private registry? Enterprise teams keep every skill inside their own workspace with SSO, audit logs and a custom DPA/NDA.

Pricing

One plan. Everything included.

No tiers to compare and no add-ons. Browsing and installing public capabilities stays free forever — Pro is for when you want your own capabilities tested, built and kept current.

Everything in Super Agent Skill
  • Unlimited skill reviews with full graded reports
  • The Agent Factory — build a custom corporate agent from a prompt
  • The Agent Store — 33 ready-made agents with soul, guardrails and skills
  • SAK University — diagnosis, adaptive curriculum, residency, credentials
  • Continuous adversarial re-testing on everything you installed
  • Private packages, signed releases and mutual NDA on request

Solo or team. Cancel in one click. Everything you create stays yours.

Pro · billed yearly
$228$140/ year
Save $88 · 39% off

Instead of $228 / year billed monthly at $19, which stays available if you prefer the flexibility.

Get Pro →Browse the free registry

Need a private registry, SSO or audit logs? Enterprise.

FAQ

Questions, answered straight.

Everything people ask before connecting an agent.

A file that teaches an agent to do one job well — instructions, examples, an output contract and guardrails. Not a prompt. Alongside skills we ship playbooks (multi-step workflows), souls (drop-in expert personas) and guardrails (what the agent must never do).

Something else? hello@superagentskill.com

One MCP URL.
One sentence. Done.

Install the best packages of your industry, generate what doesn't exist yet, and let SkillForge ship better versions for you — week after week.

$https://superagentskill.com/api/public/mcp