Skip to main content
Your engineering metrics have an incident side to them: which services had incidents, how bad they were, and how long it took to recover. That’s the data behind DORA metrics like change failure rate and time to restore, and it lives in incident.io. Engineering metrics platforms like DX, Cortex, and Port read it straight from our API and join it with the deployment and code data they already collect from your CI/CD and source control. The tools do their own calculations. Your job on the incident.io side is to make sure every incident records which service it affected, and has the timestamps your metrics are built from.

Set up your incident data

Record the affected service

Every tool joins incidents to its own services through a custom field on the incident. Make it a Catalog-backed multi-select field (e.g. Affected services) rather than free text, so the values line up with the services in your Catalog. See Custom fields to create one, and add it to your incident forms so responders actually fill it in. If responders already pick something the Catalog can map to a service, such as an affected product, you can set the field automatically from that instead, so it’s never left empty. For incidents created from alerts, set it from the alert’s attributes. See Linking incidents to services.
Keep service names in your Catalog the same as the names in your engineering metrics tool, so payments-api in one and Payments API in the other don’t end up as two services.

Define time to restore

MTTR is only as good as the timestamps it’s measured between. Duration metrics measure the time between two timestamps, and you get to choose which two. If you measure from when impact started rather than when someone got round to declaring the incident, you’ll get a more honest recovery time. Configure timestamps and duration metrics in Settings → Response → Lifecycle, on the Timestamps and metrics tab. Turn on Enable validation for the metrics you report on as well, so a mistyped timestamp can’t give you a negative or wildly inflated duration.

Decide which incidents count

By default, the API doesn’t return test or tutorial incidents, or incidents that were declined, canceled, or merged into another one, so none of those will reach your metrics tool. Incidents still in triage do come through. See Which incidents are counted for how this compares with Insights.

Connect your tool

Each tool connects with an API key. Give each one its own key, so you can see what it’s doing and revoke it without affecting the others.

DX

DX’s incident.io connector needs an API key with the View data and View catalog permissions, and the ID of your affected services field, which you can find with the list custom fields endpoint.

Cortex

Cortex’s incident.io integration lists the permissions its API key needs. To bring your Cortex services into incident.io’s Catalog, see Cortex.

Port

Port’s incident.io integration needs a read-only API key.

Other tools

Any tool that can call an HTTP API can read the same data. Start with the API reference, and the list incidents endpoint.