From model registration to production monitoring — Aegis Sovereign covers the full MLOps and AI governance lifecycle. Champion/challenger registry, LLM gateway with data-classification routing, cryptographic audit chain, policy-as-code enforcement, real-time kill switch, and a defence-in-depth AI Safety Platform with ML classifier ensemble, agentic tool guards, human review workflow, and automated red-team certification. All in one place.
A mandatory 8-phase safety pipeline wired into every request path, every agent handoff, and every model promotion decision. Not an optional add-on — a platform invariant enforced before any model response reaches your users.
JAILBREAK-001
Instruction override detected
PRIVACY-002
API key in output — auto-redacted
HARMFUL-001
CBRN keyword — blocked
Llama Guard, ShieldGemma, Azure AI Content Safety, and OpenAI Moderation run in parallel. Configurable voting strategy: highest severity, majority vote, or weighted score.
Tracks PII observed in early turns and blocks it from appearing in later outputs — catches multi-turn data exfiltration attacks that single-request scanners miss.
Approvals — Pending Review
jailbreak · 2m ago
harmful_content · 8m ago
privacy_leakage · 14m ago
A caller agent cannot grant a callee more authority than it possesses. Trust levels — untrusted → internal → trusted → admin — enforced before every agent handoff.
Before any tool executes, AgenticSafetyGuard checks risk level, argument patterns (block rm -rf /, DROP TABLE), and principal allow-lists. Blocked tools fail loudly — never silently.
Red-Team Certification — Latest Run
Block and flag rates, top violation categories, SLO burn-rate, and an overall posture score — computed from real safety audit events and LLM request volume, not synthetic numbers. Exposed at /safety/dashboard/overview and /safety/metrics/current.
A pre-deployment clearance gate aggregates bias, robustness, and safety-eval results into a single verdict. Promotion is blocked until the model earns clearance — queryable at /safety/clearance/{model}.
PII protection
0 leaks / 24 h
Block-rate budget
≤ 10% / 24 h
Injection budget
≤ 25 / 24 h
Safety posture score
> 90 / 100
Burn rate > 14.4× = 5% of weekly error budget consumed per hour → PagerDuty page fires automatically
Automated compliance evaluation against every major AI regulation — EU AI Act, DORA, MAS TRMG, NIST AI RMF, SR 11-7, HIPAA/FDA, and SOX/SEC. Generate audit-ready artifacts in minutes, not weeks.
Run full compliance evaluations against any framework in under 5 minutes. Scoring, findings, and recommendations generated automatically.
Generate signed compliance packages your auditor or Legal team can receive directly — no manual formatting.
Disparate Impact Ratio analysis across demographic groups, with intersectional breakdown and EEOC-aligned thresholds.
Every model event is cryptographically chained. Verify chain integrity at any point — tamper-evident by design.
Summarise Q3 earnings for APAC…
Patient DOB and diagnosis code…
Ignore previous instructions…
Draft board memo re: merger…
Every LLM prompt is classified by data sensitivity before it leaves your network. PII, PHI, and confidential content is routed to on-prem or regional models. Jailbreak attempts are blocked at the gateway. Full cost attribution and chargeback reporting by team, user, and model.
Automatic content classification routes sensitive prompts to compliant, in-jurisdiction endpoints.
Jailbreak detection, PII extraction attempts, and policy violation scoring on every request.
Per-team, per-model token spend with budget alerts and departmental cost attribution.
Test routing rules against sample prompts before pushing policy changes to production.
Register, version, and promote models with compliance gates built in. Champion/challenger A/B splits, MLflow experiment linking, auto-generated model cards, and mandatory bias + robustness evaluation before any model reaches production.
Full model lineage: training data hash, experiment ID, eval scores, and promotion history.
Live A/B traffic split with automatic rollback if challenger degrades on any metric.
Auto-generated cards with intended use, performance metrics, bias results, and compliance status.
Block promotion until bias DIR, robustness score, and all compliance controls pass.
Agent skills and the MCP servers they mount — instructions, bundled code, tools, and external endpoints — are the fastest-growing new attack surface. Aegis Sovereign governs them like models and prompts: registered, versioned, safety-scanned, and blocked from production until they pass. Your cloud, your data, your keys.
Static checks for prompt injection, dangerous code, undisclosed egress, tool-description poisoning, and hidden-text obfuscation — with a severity-scored verdict.
Nothing enables in production until its scan passes — plus OPA policy — returning 422 SKILL_NOT_SAFE or MCP_SERVER_NOT_SAFE with the blocking findings.
Track where a skill or MCP server came from and require review for unsigned or marketplace-sourced artifacts — unauthenticated network-exposed MCP endpoints are blocked outright.
Every invocation is written to the audit chain; production skills feed the EU AI Act Annex IV documentation as system components.
WebSocket-powered live dashboards stream drift, prediction volume, and latency in real time. When a model degrades, the kill switch pulls it from traffic in under two seconds — no manual intervention required.
Emergency Stop
Pull any model from production traffic in <2s. Automatic rollback to last stable champion.
Population Stability Index computed continuously — alert on PSI > 0.10, auto-kill on PSI > 0.25.
Slack, JIRA, PagerDuty, and email webhooks. Critical issues page on-call within 60 seconds.
Multi-persona KPIs: Legal sees compliance scores, Ops sees model health, Dev sees latency and throughput.
Per-team spend caps with real-time burn-rate display. Auto-throttle when budget exceeds 80%.
The Aegis Sovereign MCP Server exposes 31 governance tools directly to Claude Code and Claude Desktop via the Model Context Protocol. Run compliance checks, inspect audit chains, validate safety policies, and promote models — all from a natural-language prompt in your IDE.
Compliance Tools (3)
check_compliance_status, run_compliance_eval, run_compliance_eval_all (all 7 frameworks in parallel)
Registry Tools (5)
list_models, get_model, promote_model with pre-flight gate checks, get_model_card, revert_model
Audit Tools (4)
get_audit_log, verify_audit_chain, anchor_audit_chain, verify_audit_anchor (external ledger)
Monitoring Tools (4)
check_llm_usage budget bars, run_drift_check, get_workspace_metrics, get_drift_history
Safety Tools (15)
safety_validate_text, safety_validate_tool_call, safety_validate_delegation, safety_mask_pii, safety_check_model, safety_clearance_gate, safety_run_redteam, safety_slo_status, safety_generate_report, safety_classifier_status, safety_list_reviews, safety_submit_review
> Check EU AI Act compliance for model fraud-v2
Fetching compliance status...
✓ EU AI Act — Score: 87/100
⚠ 2 open findings: Art.13 transparency, Art.14 oversight
Suggested: run_remediation_suggestion for fix steps
> Promote fraud-v2 to production
Checking evaluation gates...
✓ Bias eval: DIR 0.91 (threshold 0.80) — pass
✓ Robustness: 0.74 (threshold 0.65) — pass
✓ Model promoted to production
> Verify audit chain integrity
Verifying chain from genesis...
✓ 247 events verified — chain intact
Four specialised LLM agents handle compliance pipelines, incident response, executive reporting, and remediation planning — triggered automatically by platform events, no manual scheduling required.
Triggered on every model registration
Selects the right regulatory frameworks for your model's jurisdiction and industry, runs all evaluations in parallel, auto-promotes on pass, or routes to Legal for HITL approval on failure.
Triggered on drift, robustness failure, bias failure
Classifies severity, generates an LLM incident summary, and notifies Slack, JIRA, and PagerDuty — critical incidents page on-call immediately. Dead-letters after 3 retries.
Runs every Monday 07:00 UTC
Aggregates compliance scores, identifies models with >10pp degradation over 30 days, and generates a board-ready executive summary distributed via webhook and stored in the audit log.
Triggered on any failed compliance evaluation
Analyses failing controls, maps them to specific regulatory text, and generates a prioritised remediation plan (critical/high/medium) with effort estimates and implementation steps.
Event-Driven Trigger Flow
Model deployments are version-controlled manifests. Each manifest is linked to a compliance evaluation — if the eval fails, the deployment is blocked. OPA/Rego policy bundles define governance rules that apply across all models, with mandatory Legal or CISO approval gates before production promotion.
YAML-based model manifests with compliance evaluation IDs baked in. Git is the source of truth.
Author and test governance rules as code. Policies evaluated on every model action, not just at deploy time.
Route high-risk promotions to Legal, CISO, or custom approver groups with audit trail for every decision.
Compliance degradation triggers automatic rollback to last known-good manifest state.
$ git push origin main
Trigger: model-manifest/fraud-v4.yaml changed
Fetching linked compliance eval: eval-0x8f4a…
✓ EU AI Act — 91/100 (threshold 80) — pass
✓ SR 11-7 — 88/100 (threshold 80) — pass
⚑ Routing to Legal approval: high-risk jurisdiction
$ Legal approved — promote fraud-v4
Approval recorded: legal@bank.com · 2026-04-24T09:14Z
✓ A/B split: fraud-v3 80% · fraud-v4 20%
✓ Champion/challenger active. Monitoring enabled.
Test model resilience against five attack types — FGSM, PGD, Carlini-Wagner, DeepFool, and Square Attack. Required for NIST AI RMF MEASURE-ME4, FDA SaMD submissions, and EU AI Act Article 9.
Drag AutoGen, LangGraph, and CrewAI agents onto a shared visual workspace. The Sovereign Canvas normalises message schemas, handles state serialisation, and routes tool calls across frameworks — so your teams can compose without compromise.
Unified message bus translates between agent protocols in real time.
Checkpoint and resume any agent graph across deployments.
Every node is wrapped in a zero-trust security mesh — PII masking, Azure PIM/AD integration, and real-time audit trails.
Automatically detect and redact sensitive data before it reaches any agent or log.
Just-in-time privileged access with conditional policies and MFA enforcement.
Every agent action is logged, hashed, and available for compliance review.
The Federation module connects multiple institutions — banks, insurers, hospital networks — in a shared governance network with differential privacy guarantees. Aggregate compliance scores, share threat intelligence, and co-author policies without sharing raw model data.
Up to 16 institutions share a governance overlay. Each retains full data sovereignty over its own models.
Aggregated compliance metrics are ε-DP guaranteed. No individual model's data can be reverse-engineered.
Adversarial prompt signatures and bias detection patterns propagate across participants without raw data exposure.
Each participant's jurisdiction is respected — EU members get GDPR/EU AI Act gates; APAC members get MAS TRMG.
The Sovereign Bridge will connect on-prem Kubernetes clusters to cloud-managed control planes via mTLS tunnels. Data never leaves your perimeter — only orchestration metadata crosses the bridge. On the roadmap alongside multi-region deployment.
End-to-end encrypted channels with automatic certificate rotation.
Run models on-prem with cloud-orchestrated batching and failover.
The Reaper continuously monitors token spend, context-window utilisation, and agent idle time — then prunes stale branches and compresses histories without dropping accuracy.
Lossless summarisation keeps context windows lean at scale.
Automatically suspends dormant agents. GPU-slot reclamation lands with the runtime-telemetry work on the roadmap.
Set per-agent spend caps with real-time alerting and auto-throttle.
Real-time P95 latency compared against standard SaaS orchestration platforms.
Provision a 14-day evaluation sandbox pre-seeded with 5 demo AI models, compliance evaluations already run, 200+ audit events, and a live LLM gateway. No credit card. Work email required.
5 Demo Models
Fraud, credit risk, clinical NLP, churn, credit scoring
Compliance Pre-run
Your chosen framework already evaluated with findings
200+ Audit Events
Full audit trail ready to verify and export
14 days · No credit card · Work email required · Sandbox expires automatically