Skip to main content
Aurora CognitiveAurora Cognitive

MANAGED SUPPORT

Someone owns it at three in the morning

Alerts tied to user impact, runbooks for the failures that actually happen, a rota with named escalation, and restores tested before you need one.

Pressure test my on call

30 minutes with an engineer, no sales call

What this is

Monitoring tells you, alerting wakes someone

Most teams have dashboards nobody watches and alerts everybody mutes. We separate the two: monitoring records what the system is doing, alerting fires only when a person needs to act because users are affected. Behind each alert sits a runbook, and behind the rota sits a named escalation path. Response windows are agreed per engagement and written into the terms, never implied.

This is for you if

  • An outage is discovered by a customer message rather than by a system you own.
  • Alerts fire so often that the team has stopped reading them, so the real one gets missed.
  • You have backups but nobody has ever restored one, so you do not know whether they work.

This is not the right fit if you want a shared inbox with no ownership defined. Support without a named owner, a severity scale and an agreed response window is a queue, and it fails at three in the morning exactly when it matters.

Scope

What we build

01

Alert policy tied to user impact

Each alert names the user facing symptom it represents. Anything that does not require action within the hour becomes a report, not a page.

02

Runbooks for the top failure modes

For each realistic failure: how to confirm it, how to reduce impact now, how to fix it properly, and who to tell.

03

On call rota with named escalation

A published rota, a primary and a secondary, and a clear point where the issue goes to your side or to a vendor.

04

Patch cadence

A regular window for operating system, runtime and dependency updates, with a separate faster path for security fixes.

05

Restore tests on a schedule

Backups restored into a scratch environment at a set interval, with the result and the elapsed time recorded each time.

How it goes

From reactive firefighting to defined ownership

  1. 01

    Failure mode review

    One week

    We walk your systems and recent incidents, list what realistically breaks, and rank by user impact.

  2. 02

    Alerting and runbooks

    Two to three weeks

    Noisy alerts retired, impact based alerts added, runbooks written for the ranked list, severity scale agreed.

  3. 03

    Rota and restore rehearsal

    One to two weeks

    Rota published with escalation names, patch windows set, and a restore performed end to end with your team watching.

  4. 04

    Steady state review

    Monthly, ongoing

    Incident review, alert noise check, patch and restore evidence, and any runbook updated by what actually happened.

Artefacts

What you receive

Written record

Failure mode list ranked by user impact

Runbook per failure mode

Severity scale with agreed response windows

Incident reviews with timelines

In your systems

Impact based alert rules

Monitoring dashboards for trend, not for paging

Patch schedule with a separate security path

Backup jobs with restore verification

Operations

Published on call rota with named escalation

Restore test log with elapsed times

Monthly operations report

Handover session recording

Response windows are agreed per engagement. Access stays shared with you, and we never hold sole access to a client account.

Questions

Questions we get about managed support

Next step

Tell us about the last incident nobody owned

Book a technical assessment

30 minutes with an engineer, no sales call