How it works · the closed-loop architecture
Detect · Investigate · Decide · Approve · Act · Verify · Learn — verified recovery closes the loop

The vendor-neutral autonomous incident-response layer for streaming video.

Streamwake investigates streaming failures, governs the remediation, and verifies that viewers actually recovered before the incident closes — end to end, on the same bus, for every incident the platform runs. Vendor-neutral across encoder, CDN, DRM, and player — works across Mux, Bitmovin, AWS, Broadpeak, Wowza, any CDN, and PagerDuty, with no consolidation onto a Streamwake-owned monitoring suite.

The closed-loop, in one picture

Seven stages, an explicit approval step before any write.

Read top to bottom on mobile, left to right on desktop. The agent proposes between Decide and Approve, but nothing writes to vendor config until a permitted operator signs off. The loop closes only when Verify confirms recovery.

01Detectencoder · CDN · DRM · player02Investigatehistory + cross-domain03Decidetyped mode + confidence04Approveoperator signs off05Actsmallest safe action06Verifyre-probe · close clean07Learnpostmortem + replayAGENTgoverns+ writes
The seven stages

Detect. Investigate. Decide. Approve. Act. Verify. Learn.

The same loop, on the same bus, for every signal Streamwake observes. Read top to bottom — each stage hands off to the next, with an explicit human Approve between Decide and Act, and the verification step closes the loop before writing the postmortem.

01

Detect

  • Player-QoE, encoder health, CDN egress, DRM handshake latencies all POST to one bus.
  • Anything anomalous gets a priority and a stage attached at ingest, not at triage.
  • Composes with the metrics / logs / traces tier you already run — same picture, same bus.
02

Investigate

  • Joins encoder × CDN × DRM × player signals into a typed incident joined with recent history.
  • Cross-domain correlation surfaces the affected surface, geography, and timeframe before any human reads.
  • Returns a structured view rather than a flat alert list to triage.
03

Decide

  • LLM-classifies the failure mode — manifest drift, edge brown-out, DRM cold-start — with explicit confidence.
  • Picks the smallest safe remediation candidate and proposes it as a typed action.
  • Hands off a typed next-action to the Approve step, never to vendor config directly.
04

Approve

  • A permitted operator signs off within role-scoped blast-radius RBAC and a typed hand-off contract.
  • Only against whitelisted actions, only on channels you have explicitly approved.
  • Nothing writes to vendor config until a human approves — by design, never by automation.
Explicit approval before any write
05

Act

  • Picks the smallest safe remediation — reroute egress, roll a flag, re-package a title.
  • Executes through a role-scoped blast radius, only against actions you have whitelisted.
  • Lands as a typed, auditable event scoped to the channel — replayable end to end.
06

Verify

  • Reads the same bus post-action and confirms the failure class is gone before close.
  • Anything outside the configured window reopens the incident with new evidence attached.
  • Closes the loop with a single typed event ready for downstream consumers.
07

Learn

  • Pushes a structured postmortem to Slack / Linear / PagerDuty the moment the loop closes.
  • Reused playbooks index the next response — what happened, what was tried, what changed.
  • Replay artifacts ship alongside the postmortem so the next incident reuses the same shape.
The role-scoped blast radius

What the agent is allowed to touch.

By default Streamwake reads signals rather than writes to vendor config. Writes happen only against actions you have whitelisted, are scoped to the channel, are gated behind the explicit Approve stage between Decide and Act, and are recorded as typed, auditable events you can replay.

Default posture
Read signals. Land typed incidents. Wait for a verified close. The agent watches, classifies, and proposes — it doesn't push a remediation to the CDN until you tell it which ones are safe and which channels can take them.
When you turn writes on
One action, one channel, one runbook — the agent picks the smallest safe move off its whitelist and runs it. If the fix isn't safe to apply automatically, the agent hands off to a permitted operator with a typed next action — never forces one.
  • Pushing the typed fix back as a Datadog event / New Relic event / Grafana annotation.
  • Rerouting egress on a Metro / region you have explicitly whitelisted.
  • Re-packaging or re-keying a title through your packager credential, scoped to the channel.
Operator-confidential walkthrough

Bring us your last streaming incident. A walkthrough, on us.

You share the timeline — the affected geography, the symptom, the moment it cleared. We come back with an incident replay, the ranked root cause across encode / edge / DRM, and the typed remediation the Streamwake agent would have run on the same signal bus.

Free. Operator-confidential.Replies from a person on the team within two business days.or streamwake@polsia.app