working local prototype · founded 2026

Compile intent into agent capability.

Skillaryx is building a control layer that turns scattered prompts, scripts, tools and runbooks into reusable, evaluated and governed capabilities for AI agents.

BUILD LOG #0002 · WORKING LOCAL MVP · BOOTSTRAPPED
SKILLARYX / CAPABILITY COMPILERREV 0.3
Describe a capability
01 PARSE
02 CONTRACT
03 EVAL
04 POLICY
05 READY
artifact / release_auditor.skillbld_4f91a2
skill: release_auditor
goal: verify before deploy
runtime: claude
inputs: repo, policy, test_plan
permissions: repo.read, tests.run
network: deny-by-default
evals: safety, regression, output_schema

✓ contract valid
✓ 18/18 evals passed
✓ policy satisfied
93
READINESS
behaviorPASS
regressionPASS
permissionsPASS
runtimeCLAUDE
SKILL PACKAGE · versioned
EVAL SUITE · repeatable
PERMISSIONS · explicit
RUNTIME ADAPTER · portable
BUILD LOG · observable
SKILL PACKAGE · versioned
EVAL SUITE · repeatable
PERMISSIONS · explicit
RUNTIME ADAPTER · portable
BUILD LOG · observable
01 / PRODUCT THESIS

Agent skills should feel more like software than prompt folders.

A reusable AI capability should have a contract, a version, tests, resources, permissions and a known execution target. Skillaryx is being built around that idea.

THESIS / 001

Prompts are instructions.
Skills are operational assets.

Once an agent can touch files, tools, shell commands or business systems, “good prompting” is no longer enough. The capability needs evidence, boundaries and lifecycle management.

PRINCIPLE / 002

Prove behavior before deployment.

Evals belong next to the skill definition, not in an after-the-fact checklist.

PRINCIPLE / 003

Permissions are part of the skill.

Access to files, tools, network and connectors should be explicit and reviewable.

02 / CONTROL LAYER

One artifact. Four systems of trust.

Skillaryx packages the pieces that usually live in separate files, prompts, notes and scripts into a single capability lifecycle.

01
◫

Skill Registry

Stable identities, versions, dependencies and runtime compatibility for reusable capabilities.

02
✓

Eval Runner

Behavior, regression, output-contract and safety checks that travel with the skill.

03
⌁

Policy Engine

Explicit boundaries for files, shell, network, APIs and connector access.

04
↗

Runtime Adapters

Execution adapters that translate a capability contract into a supported agent environment.

03 / WORKING PROTOTYPE

Not a slide deck. A running local capability lifecycle.

The current V0.1 prototype runs locally on Windows. It compiles a versioned skill package, validates its contract and permissions, runs blocking evaluations, enforces a policy gate, records execution evidence, and persists run history across restarts.

VERIFIED LOCALLY · V0.1

Execution is earned, not assumed.

A healthy fixture reaches the runtime only after validation, policy checks, and blocking evals pass. A failing blocking eval stops the runtime before invocation.

PASS pathValidation → Evals → Policy → Runtime
BLOCK pathBlocking eval fails → Runtime not invoked
EvidenceRun ID, timestamps, checks and decisions
PersistenceSQLite run history survives restart
Skillaryx V0.1 skill detail showing explicit permissions, blocking evaluations, and the governed execution panel.
Versioned skill contractPERMISSIONS · EVALS · RUNTIME GATE
Skillaryx execution result showing validation, evaluations, policy and mock runtime all passing.
Healthy execution → PASSRUNTIME INVOKED · EVIDENCE PRESERVED
Skillaryx execution result showing a failed blocking evaluation and runtime execution blocked.
Blocking eval → runtime stoppedFAIL · EXECUTION BLOCKED
Evidence shown above is from the local deterministic MOCK runtime used for safe regression testing. Claude is the planned first live provider runtime; these screenshots do not represent a live Claude API execution.
04 / CLAUDE-FIRST PATH

Claude is our first-class runtime path.

We plan to use Claude for skill authoring, evaluation and controlled agent execution, with MCP-compatible connectivity as part of the product direction.

PLANNED INTEGRATION / CLAUDE

Intent → contract → eval → execution.

Natural-language requirements become structured skill definitions. Repeatable evals verify behavior. Approved capabilities execute through explicit tools and permissions. The long-term goal is to keep the workflow contract stable while adapting execution to supported runtimes.

CLAUDE API
SKILL AUTHORING
EVALS
TOOL EXECUTION
05 / BUILD LOG

Start with one reliable lifecycle.

We are deliberately starting narrow: make one skill definition reproducible, testable and executable before expanding into a broader team platform.

NOW · Q4 2026

Working local prototype

Skill compiler, versioned registry, validator, policy gate, deterministic eval runner, runtime abstraction and persistent execution evidence.

NEXT · Q1 2027

Private alpha

Version history, execution logs, policy controls and additional runtime adapters.

LATER · 2027

Team control plane

Private catalogs, approvals, shared eval suites and organization-level governance.

BUILD LOG OPEN

Make agent capability reusable enough to trust.

Skillaryx has a working local prototype today. We are now moving from deterministic local execution toward a live Claude runtime and private alpha.

Contact the founder ↗