AWS DevOps Agent

An agentic DevOps team member

AWS Swiss User Meetup Bern · 16.09.2026 · Tobias Vonesch, copebit AG

A question first

Who of you has up-to-date runbooks?

Why

What you take home
1

How it's built

2

What it can and can't do

3

How to run it yourself

Why now
Meme: Buzz Lightyear to Woody — AI, AI everywhere
About me
Tobias Vonesch

Tobias Vonesch

  • From Baden, now in eu-central-2, soon back in Baden
  • Technical Sales Engineer & AWS Consultant @ copebit AG
  • BSc Computer Science
  • 6× AWS certified
  • AWS Community Builder
  • Avid runner & football fan

tobias.vonesch@copebit.ch · copebit.ch

1

How it's built

Resources · Interfaces · Pricing

Resources
What is an Agent Space

Sources

  • Primary account
  • Secondary accounts
  • Azure · on-prem

Agent

  • Topology graph
  • Skills & memories
  • Custom SRE agents

Capabilities

  • Telemetry
  • Code & pipelines
  • Tickets & chat
  • MCP servers · webhooks
Role 1DevOpsAgentRole-AgentSpaceread-only, assumed by the agent
Role 2DevOpsAgentRole-WebappAdminwhat operators may do
Interfaces
Admins

AWS Console

  • Agent Spaces & roles
  • Accounts & integrations
  • Who may sign in
Operators

DevOps Agent web app

  • Investigations & prevention
  • Topology & skills
  • Chat
Machines

API & remote endpoints

  • CLI · CDK · CFN · Terraform
  • Webhooks in · EventBridge out
  • MCP & A2A
Pricing model
$0.0083per agent-second
≈ $0.50per minute of agent work
≈ $30per hour of agent work

No tokens, only minutes

Sample calculation10 investigations × 8 min ≈ $40 / month  ·  2-month free trial  ·  Support credits 30 / 75 / 100 %

2

What it can and can't do

Three jobs · One workflow · Hard boundaries

Functionality
When an alarm fires

Incident response

Alarm → root cause → mitigation plan

When nothing is on fire

Incident prevention

Evaluations → recommendations

Whenever you ask

On-demand SRE tasks

Chat · charts · custom SRE agents

Workflow
1

Trigger

Ticket · webhook · chat

2

Triage

Severity · duplicates

3

Investigate

Logs · metrics · code · deploys

4

Root cause

Evidence · blast radius

5

Plan

Mitigation steps

6

You act

Approve · execute · hand off

Knowledge
Discovers

Topology graph

  • CloudFormation stacks
  • Resource Explorer tags
  • Observed dependencies
Reads

Your code

  • GitHub · GitLab
  • Package → infrastructure
  • Code-level fixes
Learns

Learned skills & memories

  • Agent Space Understanding · every 3 days
  • Tool Use Best Practices · every 30 investigations
You teach

Custom skills

  • SKILL.md — your runbook
  • Per phase: triage · RCA · mitigation
  • Upload or manage in Git
Boundaries

It does

  • Read what you allow
  • Correlate & explain
  • Write to tickets & chat
  • Log every step
  • Stay in its space

It does not

  • Change your infrastructure
  • Replace your monitoring
  • Scale without quotas
  • Pin inference to your Region
  • Vet your MCP servers
3

Run it in your environment

Prerequisites · 20-minute setup · Hardening

Prerequisites
  1. A supported Regioneu-central-1 · eu-west-1 · eu-west-2
  2. Rights to create IAM rolesAIDevOpsAgentFullAccess + iam:PassRole
  3. SCPs that allow itaidevops:* · bedrock:InvokeModel
  4. Telemetry to readCloudWatch · CloudTrail · alarms
  5. A way to sign inIdentity Center · IdP · console link
Setup
1

Create Agent Space

Name · auto-create both roles

2

Add sources

Secondary accounts

3

Connect capabilities

Telemetry · code · tickets · MCP

4

Open the web app

Grant users · first question

Production hardening
BoundaryOne space per team & environment
RoleYour own read-only role · CMK
ToolsAllowlist MCP tools · private connection
CostCost & Usage Report per space

Demo

An imagined Friday evening

Situation

17:40 payment API alarm
1 engineer on call
deploy at 17:31

Tension

3 suspects
4 dashboards
1 person

Action

17:49 root cause in the journal
pool size → RDS connections
plan: roll back

Result

17:55 fixed
15 min, not 2 h
≈ $4.50

Thank you. Questions?