SHACKLE is a Python runtime governance layer for supported agent execution paths: repeat-call controls, model-aware spend estimates, timeouts, and operator decisions. It also publishes SP/1.0.1, an open mediation contract with four versioned, hash-pinned certification profiles. The reference runtime passed those profiles on its recorded artifact; today's later hardening adds native CrewAI hooks, broader LiteLLM interception, and explicit coverage-gap reporting. Exact scope, evidence, and exclusions are public.
Agent hits an error. Retries same tool with same input. Burns tokens quadratically. You wake up to a $400 bill — or worse, a $106,000 AWS disaster — with no idea what happened.
CrewAI, AutoGen, and LangGraph have zero built-in cost guardrails. Your agent runs until it runs out of your money. Framework developers admit this is unsolved.
API hangs. Agent waits forever. No timeout. No alert. Just a frozen process burning compute budget. Across 20+ agents in production, this compounds exponentially.
Production stacks run CrewAI + AutoGen + LangGraph simultaneously. Each framework has different failure modes. None has a unified circuit breaker. Until now.
An AutoGen agent managing 10,000+ AWS accounts destroyed the management account in a recursive cleanup loop. No circuit breaker. No budget guard. No kill switch.
Source: microsoft/autogen#7770 — @tzb1-ai, enterprise infrastructure operator
SHACKLE includes repeat-call and error-signal controls for supported, tested execution paths. This report is context for the problem—not evidence that SHACKLE would have prevented that specific incident.
SHACKLE attaches to supported execution seams rather than claiming universal framework coverage. Current runtime hooks include native CrewAI before-tool / before-LLM / after-LLM hooks, LiteLLM completion and call-time gates (including early-bound imports), and LangChain BaseTool run/arun. AutoGen uses its supported wrapper or a LiteLLM-routed path; unsupported paths must be assessed explicitly.
Detects repeated tool calls and tested error-cascade patterns on covered hook paths; it is not a universal execution sandbox.
Matches known provider-prefixed and dated model IDs to pricing rows. Unknown rates fall back with a warning; provider billing can differ.
Supported LLM gates check known exhausted budget and elapsed time before a subsequent request. Token settlement depends on integration data.
Resume / Skip / Abort flows on supported interactive paths. Headless or event-loop boundaries fail closed instead of blocking for input.
V1 hooks execute in the user's Python process. The v2 sidecar is a separate deployment option; neither establishes host isolation by itself.
Native CrewAI hooks, supported LiteLLM paths, and LangChain BaseTool hooks are detected; strict mode can reject detected gaps. Check custom paths separately.
SHACKLE focuses on runtime mediation for supported tool and LLM call paths: repeat-call controls, model-cost estimates, time limits, and explicit decisions. It complements framework controls and monitoring. Coverage depends on the installed integration and actual execution path; verify that boundary for each deployment.
| SHACKLE capability | What the current implementation does / boundary |
|---|---|
| Repeat-call controls | Tracks repeated tool calls and tested error-signal patterns on hooked paths. |
| LLM cost estimates | Uses a local pricing table with provider-prefix and dated-model resolution; unknown models warn and use the default row. This is not the final provider invoice. |
| Pre-call checks | Supported LLM hooks can block a request when known budget or elapsed-time thresholds are already exhausted. |
| Operator decisions | Provides interactive control on supported paths; headless/event-loop constraints are handled fail-closed rather than blocking for terminal input. |
| Coverage visibility | Reports detected framework hook gaps; optional strict mode refuses to run when detected gaps exist. Dynamic/custom paths still require deployment-specific review. |
| Certification evidence | Publishes four hash-pinned profiles and an artifact-bound registry entry; evidence applies only to the identified artifact and declared scope. |
Use it precisely: confirm the actual tool/LLM path is covered, use provider-side spend caps, and interpret certification only against the named profile hashes and tested artifact.
SHACKLE attaches to configured framework/provider hooks in the Python process. The diagram is conceptual: every application must verify that its actual execution path uses a supported hook.
SHACKLE includes a runtime circuit breaker and publishes the SP/1.0.1 conformance standard for runtime mediation of agent tool calls — an authored, verifiable contract that any runtime can be tested against.
ALLOW / DENY / HITL are the named outcomes. Enforcement validates returned decision shapes rather than treating unknown values as implicit permission.
Valid(τ) ⇔ Required(τ) ⊆ Supported(τ). A transition is valid only when required capabilities are supported.
15 decision-surface, 14 SP/1.0.1 adversarial-input, 31 decision-result hardening, and 22 runtime-adversarial cases: 82 cases with published SHA-256 identities.
The registry binds the owner-verified reference runtime to commit a45eaa8, its archive digest, configuration, four profile hashes, and stated exclusions. Fresh runs: 323 tests in `tests/`, 394 at repo root, verifier PASS.
Core invariant: history-visible ≠ runtime-executable. A record that an action happened is not proof that the transition was supported. Conformance is evaluated against the named public fixtures and the exact result is reproducible.
Scope boundary: the registry now binds the owner-verified reference runtime to commit a45eaa863a7d2afb0847ce81d0952737abdf98fa and its recorded source archive digest, after fresh verification of all four unchanged, hash-pinned profiles. The independent reviewer’s CrewAI-over-LiteLLM regression was reproduced and corrected; that specific reproduction is not an independent rerun of the full profile battery. This is evidence for named profiles and artifacts, not a guarantee of absolute security, all deployments, legal outcome, or absence of unknown vulnerabilities.
SHACKLE publishes a reviewable test-and-registry process for runtime mediation. Its four official profiles contain 82 hash-pinned cases across the decision surface, SP/1.0.1 adversarial inputs, decision-result hardening, and runtime adversarial behavior. Each result is tied to the tested source artifact, configuration, date, and declared scope.
What a listing means: the identified implementation passed the exact profiles and hashes named in its registry entry under the recorded procedure. The result can be independently reproduced from the public artifacts. It is evidence about those tests and that artifact; it is not a claim of absolute safety, absence of unknown vulnerabilities, production enforcement on every deployment, regulatory approval, or legal outcome.
Published decision cases for the core ALLOW / DENY / HITL surface.
Adversarial input cases for the published mediation contract.
Malformed and out-of-contract decision-result enforcement cases.
Selected runtime boundary, approval, replay, and concurrency cases; not an exhaustive attack enumeration.
Any agent runtime—including competing frameworks and safety products—can test against the published profiles. The initial registry entry is the owner's self-verified reference-runtime result; it is not an independent certification or a third-party pass. Later code changes do not inherit that entry unless separately reverified and rebound.
Read each result against the artifact and procedure it actually covers. Earlier independent reproductions are historical; the latest reference certification is owner-verified, and later runtime hardening has separate green CI.
@nutstrut reproduced earlier fixture surfaces. His September SP/1.0.1 rerun identified two qualifications, which were addressed in a public follow-up. The four-profile result is owner-verified on the corrected artifact; the independent review reproduced the specific accounting regression but did not rerun the full profile battery. Review record · Voluntary invitation.
The owner-verified profile result is bound to the recorded reference commit and four exact profile hashes; it is not an external certification. Read the verification report and registry.
The corrected runtime commit has a green five-job CI run; Certification Verify and Pages deployment also succeeded on that source commit. The results are scoped to the named artifact and tests.
The $106,000 runaway-agent incident is a public report (microsoft/autogen#7770), not a hypothetical. SHACKLE's error-cascade breaker trips a repeated error-bearing call (401/403/500/timeout) on its 2nd attempt instead of the 3rd — verified in the suite.
Current source includes native CrewAI hook governance, two-layer LiteLLM interception, LangChain BaseTool hooks, and a supported AutoGen wrapper. Coverage is path-specific: AutoGen is not natively hooked, provider routes differ, and a wrapper or LiteLLM gate only governs calls that actually pass through it. Use strict coverage reporting as one signal, then verify the application path.
ShackleGuardrail enforces the pure SP/1.0 reference decide();
ShackleEngineGuardrail drives the full stateful circuit breaker
(budget / repeat / timeout). Sync check()/record() for the SDK, plus
async pre/post hooks for the LiteLLM proxy. litellm is optional.
@wrap_tool governs any AutoGen tool through the same engine, with
canonical input dedup so dict key ordering can’t evade loop detection.
create_shackle_agent() spins up a governed AssistantAgent.
Every path enforces Valid(τ) ⇔ Required(τ) ⊆ Supported(τ)
and the ALLOW / DENY / HITL decision surface. See
INTEGRATIONS.md
for proxy config and usage.
Dante Bullock — 52-year-old self-taught systems architect, Oakland, California. Founder of Sovereign Logic. No venture capital. No corporate incubator. Just code that works.
"I don't wait for VC validation. I scrape issue trackers, find the bleeding, and build the tourniquet."