Operate
From stable operations to the learning, self-driving organization
The Operate phase establishes the permanent, largely autonomous operation of the agentic target landscape: SRE-based reliability steering, agentic incident handling, continuous AI governance and an improvement loop feeding operational insights back into strategy and backlog — making the organization self-driving in the literal sense.
Position of the phase in the TRANSFORM delivery model
Process model
The phase delivers measurably reliable, compliant and economical operations: met SLOs, autonomous resolution of standard incidents, continuous model and cost management, and an established improvement cycle back into the TRANSFORM phases.
Service & SLO design
Define service level indicators and objectives per service and agent; agree error budgets as the steering instrument between business, operations and delivery.
Observability & AIOps
End-to-end telemetry (metrics, logs, traces, agent decisions); agentic anomaly detection and root cause analysis across the whole landscape.
Autonomous operations & incident management
Self-healing for standard incidents through operations agents; structured incident management with human escalation and blameless post-mortems per SRE practice.
AI governance & model operations
Continuous monitoring of model behaviour, drift and guardrail compliance; compliance evidence per EU AI Act and NIST AI RMF; controlled model and prompt releases via the eval suites.
Continuous improvement & value evidence
FinOps-based cost steering, measurement of realized value contribution against the OKRs; feeding improvement needs back into the Strategy and Plan phases.
Methodology mix
Self-driving operations with governance and continuous improvement
| Method | Purpose in the TRANSFORM context | Reference |
|---|---|---|
| Site reliability engineering (SRE) | SLIs, SLOs and error budgets as quantitative reliability steering; automation before staffing. | Beyer et al. (Google, 2016) |
| ITIL 4 — service value system | Framework for service management practices (incident, problem, change) in interplay with agile delivery. | PeopleCert/Axelos |
| AIOps | AI-supported anomaly detection, event correlation and root cause analysis in IT operations; scientifically grounded failure management methods. | Notaro et al. (2021) |
| EU AI Act & NIST AI RMF | Regulatory and risk-based frame for operating AI systems; basis of the compliance evidence. | EU 2024/1689; NIST |
| FinOps | Operating model for the economic steering of cloud and AI resources (incl. the agents' token economy). | FinOps Foundation |
| Continuous improvement (PDCA/Kaizen) | Systematic learning loop from operational data back into strategy, backlog and architecture; closes the TRANSFORM cycle. | Deming; ITIL CSI |
Platform support
ReqPOOL Suite: REAM (Enterprise architecture) · SLIP (Reverse engineering — whitebox)
REAM — lifecycle & architecture management
Lifecycle and architecture management in one inventory: every productive application with responsibilities, technology stack and lifecycle status.
REAM — end-of-support & technical debt
Automatic end-of-support detection and technical debt monitoring; the results feed the continuous transformation roadmap.
REAM — loop back to strategy
New demands from operations feed the next ideation cycle — the lifecycle closes into a continuous loop.
SLIP — living system documentation
Source-proven documentation, knowledge graph and AI assistant remain as working tools in operations — system knowledge no longer depends on individual heads.
SLIP — impact analyses in operations
Before every change, the assistant answers “What breaks if I change this?” — based on the actual code instead of outdated documents.
Platform usage best practices
- Prepare every change to the productive system with a SLIP impact analysis; unresolved dependencies block the change.
- Update the system documentation after every release — outdated documentation is an operational risk, not a formality.
- Assess end-of-support and debt signals from REAM quarterly and feed them as demands into the next ideation cycle.
- Enforce error budgets consistently — when exhausted, the release pipeline stops automatically; exceptions are decided by the steering committee.
- Run post-mortems blameless and feed actions as backlog items back into the Plan phase; insights without implementation are waste.
- Report value evidence quarterly against the Strategy phase's business case — operations prove the transformation.
Artifacts & outcomes
- SLO catalogue with error budget policy
- Established autonomous operations with runbook library
- AI governance evidence (EU AI Act, NIST AI RMF)
- Quarterly value and cost report (FinOps)
- Continuous improvement backlog into subsequent cycles
Quality gate — Operate
Transition to the next phase happens through a formal quality gate (go/no-go). Gate criteria are documented in the gate review and signed off by the engagement lead.
- SLOs met over a full reporting period
- Agent autonomy levels documented and approved
- Compliance evidence complete and current
- Improvement cycle anchored in Strategy/Plan
Scientific deep dives
Foundational work of the SRE approach; freely available.
Practice volume on SLO design and canary operations.
Scientific survey of AIOps research.
Binding legal framework for AI systems in the EU.
Risk-based framework for trustworthy AI.
Standard for the economic steering of cloud resources.
Best-practice guide: Operate phase
Compact checklist covering process model, methodology mix and platform best practices for use in your engagement.
Download the best-practice guide (PDF)