Runtime Governance · Published Evidence

YOUR AGENT
IS BLEEDING
MONEY.
RIGHT NOW.

$106,000 Reported Loss — AutoGen #7770 AutoGen #7770 — Agent destroyed AWS management account

SHACKLE is a Python runtime governance layer for supported agent execution paths: repeat-call controls, model-aware spend estimates, timeouts, and operator decisions. It also publishes SP/1.0.1, an open mediation contract with four versioned, hash-pinned certification profiles. The reference runtime passed those profiles on its recorded artifact; today's later hardening adds native CrewAI hooks, broader LiteLLM interception, and explicit coverage-gap reporting. Exact scope, evidence, and exclusions are public.

See How It Works ↓ View Source (AGPLv3) Pricing & Implementation ↓
$ python agent_kickoff.py
[Agent] Running research task...
[Tool] web_search("latest AI news") → 200 OK
[Tool] web_search("latest AI news") → 401 Unauthorized
[Tool] web_search("latest AI news") → 401 Unauthorized ⚠

⛓️ SHACKLE CIRCUIT BREAKER: REPETITIVE_TOOL_CALL

Agent: ResearchAgent
Tool: web_search
Input: {"query": "latest AI news", "error": "401 Unauthorized"}
Call Count: 3x → CIRCUIT OPEN

━━━ Session Stats ━━━
Tokens: In: 8,400 | Out: 1,200
Session Cost: $0.02850 (saved: ~$400.00)
Time Running: 47.2s

[R] Resume/Reset  [S] Skip  [A] Abort
▌
// The Crisis

Your Agent Is Bleeding Money. Right Now.

🔄 The Loop of Death

Agent hits an error. Retries same tool with same input. Burns tokens quadratically. You wake up to a $400 bill — or worse, a $106,000 AWS disaster — with no idea what happened.

💰 No Native Budget Enforcement

CrewAI, AutoGen, and LangGraph have zero built-in cost guardrails. Your agent runs until it runs out of your money. Framework developers admit this is unsolved.

⏱️ Hung Tools, Silent Failures

API hangs. Agent waits forever. No timeout. No alert. Just a frozen process burning compute budget. Across 20+ agents in production, this compounds exponentially.

📊 Cross-Framework Invisible

Production stacks run CrewAI + AutoGen + LangGraph simultaneously. Each framework has different failure modes. None has a unified circuit breaker. Until now.

// Case Study

This Actually Happened.

$106,000

An AutoGen agent managing 10,000+ AWS accounts destroyed the management account in a recursive cleanup loop. No circuit breaker. No budget guard. No kill switch.

Source: microsoft/autogen#7770 — @tzb1-ai, enterprise infrastructure operator

SHACKLE includes repeat-call and error-signal controls for supported, tested execution paths. This report is context for the problem—not evidence that SHACKLE would have prevented that specific incident.

// Solution

One Decorator. Zero Refactoring.

$ cat agent.py

from shackle import Guard

@Guard(budget=0.25, max_repeat_calls=3, timeout_seconds=180)
def safe_kickoff():
    return crew.kickoff()

safe_kickoff()

$ pip install pyshackle && python agent.py
✅ SHACKLE active. Guards: budget=$0.25, repeat=3x, timeout=180s

SHACKLE attaches to supported execution seams rather than claiming universal framework coverage. Current runtime hooks include native CrewAI before-tool / before-LLM / after-LLM hooks, LiteLLM completion and call-time gates (including early-bound imports), and LangChain BaseTool run/arun. AutoGen uses its supported wrapper or a LiteLLM-routed path; unsupported paths must be assessed explicitly.

🛑

Repeat-Call Controls

Detects repeated tool calls and tested error-cascade patterns on covered hook paths; it is not a universal execution sandbox.

💰

Model-Aware Cost Estimates

Matches known provider-prefixed and dated model IDs to pricing rows. Unknown rates fall back with a warning; provider billing can differ.

⏱️

Pre-Call Budget / Time Gate

Supported LLM gates check known exhausted budget and elapsed time before a subsequent request. Token settlement depends on integration data.

🖥️

Operator Decisions

Resume / Skip / Abort flows on supported interactive paths. Headless or event-loop boundaries fail closed instead of blocking for input.

🔒

Local Runtime Hooks

V1 hooks execute in the user's Python process. The v2 sidecar is a separate deployment option; neither establishes host isolation by itself.

🔀

Coverage Reporting

Native CrewAI hooks, supported LiteLLM paths, and LangChain BaseTool hooks are detected; strict mode can reject detected gaps. Check custom paths separately.

// Why SHACKLE Exists

A Focused Governance Layer.

SHACKLE focuses on runtime mediation for supported tool and LLM call paths: repeat-call controls, model-cost estimates, time limits, and explicit decisions. It complements framework controls and monitoring. Coverage depends on the installed integration and actual execution path; verify that boundary for each deployment.

SHACKLE capabilityWhat the current implementation does / boundary
Repeat-call controlsTracks repeated tool calls and tested error-signal patterns on hooked paths.
LLM cost estimatesUses a local pricing table with provider-prefix and dated-model resolution; unknown models warn and use the default row. This is not the final provider invoice.
Pre-call checksSupported LLM hooks can block a request when known budget or elapsed-time thresholds are already exhausted.
Operator decisionsProvides interactive control on supported paths; headless/event-loop constraints are handled fail-closed rather than blocking for terminal input.
Coverage visibilityReports detected framework hook gaps; optional strict mode refuses to run when detected gaps exist. Dynamic/custom paths still require deployment-specific review.
Certification evidencePublishes four hash-pinned profiles and an artifact-bound registry entry; evidence applies only to the identified artifact and declared scope.

Use it precisely: confirm the actual tool/LLM path is covered, use provider-side spend caps, and interpret certification only against the named profile hashes and tested artifact.

// Architecture

Hooks at Supported Boundaries.

SHACKLE attaches to configured framework/provider hooks in the Python process. The diagram is conceptual: every application must verify that its actual execution path uses a supported hook.

┌─────────────────────────────────────────────────────────┐ │ SHACKLE RUNTIME ENVELOPE │ │ │ │ ┌──────────┐ ┌──────────┐ ┌──────────┐ │ │ │ CrewAI │ │ AutoGen │ │LangGraph │ │ │ └────┬─────┘ └────┬─────┘ └────┬─────┘ │ │ │ │ │ │ │ └───────────────┬───────────────┘ │ │ │ │ │ ┌────────▼────────┐ │ │ │ SHACKLE │ ← litellm hook │ │ │ decide(state, │ ← BaseTool.run hook │ │ │ call) → Verdict│ ← Agent.execute_task │ │ └────────┬────────┘ │ │ │ │ │ ┌────────────┼────────────┐ │ │ │ │ │ │ │ ┌────▼───┐ ┌─────▼────┐ ┌────▼─────┐ │ │ │ ALLOW │ │ DENY │ │ HITL │ │ │ │ execute│ │ block + │ │ pause + │ │ │ │ tool │ │ log │ │ console │ │ │ └────────┘ └──────────┘ └──────────┘ │ │ │ │ V1 (Local): In-process, memory state, CLI HITL │ │ V2 (Sovereign): Sidecar daemon, Redis, Postgres, WSS │ └─────────────────────────────────────────────────────────┘
// Deployment

Two Deployment Modes. One Decision Contract.

🖥️ LOCAL HITL (V1)

  • In-process decorator — zero infrastructure
  • Terminal-based HITL console
  • Memory-only state (lost on crash)
  • Perfect for development & debugging
  • pip install pyshackle
  • Free — AGPLv3
View on GitHub →

🏰 ENTERPRISE SOVEREIGN (V2)

  • Sidecar daemon — persistent state
  • Web/mobile remote HITL console
  • Redis + Postgres — state survives crashes
  • Distributed budget across serverless/K8s
  • Cryptographically signed audit records (deployment controls remain separate)
  • Commercial license available
Contact for Pricing →

From Runtime Guard to Public Standard.

SHACKLE includes a runtime circuit breaker and publishes the SP/1.0.1 conformance standard for runtime mediation of agent tool calls — an authored, verifiable contract that any runtime can be tested against.

Closed Decision Surface

ALLOW / DENY / HITL are the named outcomes. Enforcement validates returned decision shapes rather than treating unknown values as implicit permission.

Conformance Model

Valid(τ) ⇔ Required(τ) ⊆ Supported(τ). A transition is valid only when required capabilities are supported.

Four Pinned Profiles

15 decision-surface, 14 SP/1.0.1 adversarial-input, 31 decision-result hardening, and 22 runtime-adversarial cases: 82 cases with published SHA-256 identities.

Verification Record

The registry binds the owner-verified reference runtime to commit a45eaa8, its archive digest, configuration, four profile hashes, and stated exclusions. Fresh runs: 323 tests in `tests/`, 394 at repo root, verifier PASS.

Core invariant: history-visible ≠ runtime-executable. A record that an action happened is not proof that the transition was supported. Conformance is evaluated against the named public fixtures and the exact result is reproducible.

Scope boundary: the registry now binds the owner-verified reference runtime to commit a45eaa863a7d2afb0847ce81d0952737abdf98fa and its recorded source archive digest, after fresh verification of all four unchanged, hash-pinned profiles. The independent reviewer’s CrewAI-over-LiteLLM regression was reproduced and corrected; that specific reproduction is not an independent rerun of the full profile battery. This is evidence for named profiles and artifacts, not a guarantee of absolute security, all deployments, legal outcome, or absence of unknown vulnerabilities.

Read Certification Policy → View Verification Report → View Registry →

A Public, Reproducible Certification Bar.

SHACKLE publishes a reviewable test-and-registry process for runtime mediation. Its four official profiles contain 82 hash-pinned cases across the decision surface, SP/1.0.1 adversarial inputs, decision-result hardening, and runtime adversarial behavior. Each result is tied to the tested source artifact, configuration, date, and declared scope.

What a listing means: the identified implementation passed the exact profiles and hashes named in its registry entry under the recorded procedure. The result can be independently reproduced from the public artifacts. It is evidence about those tests and that artifact; it is not a claim of absolute safety, absence of unknown vulnerabilities, production enforcement on every deployment, regulatory approval, or legal outcome.

Decision surface — 15

Published decision cases for the core ALLOW / DENY / HITL surface.

SP/1.0.1 adversarial input — 14

Adversarial input cases for the published mediation contract.

Decision-result hardening — 31

Malformed and out-of-contract decision-result enforcement cases.

Runtime adversarial — 22

Selected runtime boundary, approval, replay, and concurrency cases; not an exhaustive attack enumeration.

  1. Inspect the policy and pins. Read the certification policy, profile manifest, registry entry, and report before interpreting a result.
  2. Reproduce the named artifact. Use the exact source commit, artifact digest, configuration, profile IDs, and hashes recorded in the registry.
  3. Submit case-level evidence. Include expected and observed outcomes, unsupported cases, failures, environment, and raw reproduction details.
  4. Review before listing. Intake is not automatic certification; the evidence must be evaluated against the published requirements.

Same public bar for every implementation.

Any agent runtime—including competing frameworks and safety products—can test against the published profiles. The initial registry entry is the owner's self-verified reference-runtime result; it is not an independent certification or a third-party pass. Later code changes do not inherit that entry unless separately reverified and rebound.

// Evidence

Evidence Has Its Own Scope.

Read each result against the artifact and procedure it actually covers. Earlier independent reproductions are historical; the latest reference certification is owner-verified, and later runtime hardening has separate green CI.

Independent historical review

@nutstrut reproduced earlier fixture surfaces. His September SP/1.0.1 rerun identified two qualifications, which were addressed in a public follow-up. The four-profile result is owner-verified on the corrected artifact; the independent review reproduced the specific accounting regression but did not rerun the full profile battery. Review record · Voluntary invitation.

Reference certification record

The owner-verified profile result is bound to the recorded reference commit and four exact profile hashes; it is not an external certification. Read the verification report and registry.

Latest hardening CI

The corrected runtime commit has a green five-job CI run; Certification Verify and Pages deployment also succeeded on that source commit. The results are scoped to the named artifact and tests.

The failure is documented

The $106,000 runaway-agent incident is a public report (microsoft/autogen#7770), not a hypothetical. SHACKLE's error-cascade breaker trips a repeated error-bearing call (401/403/500/timeout) on its 2nd attempt instead of the 3rd — verified in the suite.

🔗 Seeking Integration Partners

Running multi-agent production systems? SHACKLE fills the circuit breaker gap in your delegation protocol. We're looking for integration partners with live production deployments. Cross-framework, sandbox-compatible, runtime-level.

Propose Integration →
// Integrations

Drop Into Supported Paths.

Current source includes native CrewAI hook governance, two-layer LiteLLM interception, LangChain BaseTool hooks, and a supported AutoGen wrapper. Coverage is path-specific: AutoGen is not natively hooked, provider routes differ, and a wrapper or LiteLLM gate only governs calls that actually pass through it. Use strict coverage reporting as one signal, then verify the application path.

LiteLLM guardrail

ShackleGuardrail enforces the pure SP/1.0 reference decide(); ShackleEngineGuardrail drives the full stateful circuit breaker (budget / repeat / timeout). Sync check()/record() for the SDK, plus async pre/post hooks for the LiteLLM proxy. litellm is optional.

AutoGen wrapper

@wrap_tool governs any AutoGen tool through the same engine, with canonical input dedup so dict key ordering can’t evade loop detection. create_shackle_agent() spins up a governed AssistantAgent.

Same contract everywhere

Every path enforces Valid(τ) ⇔ Required(τ) ⊆ Supported(τ) and the ALLOW / DENY / HITL decision surface. See INTEGRATIONS.md for proxy config and usage.

// Pricing

Stop the Bleeding. Today.

Open Source

$0
  • Full V1 decorator source
  • Terminal HITL console
  • Budget + repeat + timeout guards
  • AGPLv3 license
  • Community support
  • Distributed state
  • SOC2 audit logs
  • Commercial license
GitHub →

Enterprise Sovereign

Custom
  • V2 sidecar daemon + distributed state
  • Postgres audit logs (SOC2-ready)
  • Remote HITL console (web/mobile)
  • Multi-tenant isolation
  • Commercial license (no copyleft)
  • SLA-backed priority support
  • On-premise deployment
Contact →
// Provenance

Built by Someone Who Ships.

Dante Bullock — 52-year-old self-taught systems architect, Oakland, California. Founder of Sovereign Logic. No venture capital. No corporate incubator. Just code that works.

"I don't wait for VC validation. I scrape issue trackers, find the bleeding, and build the tourniquet."

GitHub: @Fame510  |  docspoc101@gmail.com