Skip to main content

Incidents

Incidents give Starfire operators a structured record for platform-impacting problems from first detection through recovery. An incident is not just a status label. It connects impact, components, timeline updates, mitigation, and verification so operators can understand what happened after the event is over.

When to create an incident

Create or update an incident when a problem has broader platform impact, such as:
  • a provider or model route is broadly unavailable
  • authentication fails for a significant population
  • billing provisioning or webhooks are degraded
  • FORGE or research jobs are failing at scale
  • Developer Platform requests are broadly failing
  • a database, queue, storage, or other core dependency is unhealthy
  • an administrative/configuration change causes service degradation
A single-user issue normally belongs in support/diagnostics unless evidence shows a wider problem.

Incident lifecycle

A useful lifecycle distinguishes investigation from confirmed resolution. Common stages can include:
The exact status vocabulary is controlled by the active Starfire build.

Components

Associate the incident with affected components rather than labeling the entire platform down by default. Component scope might include areas such as Chat, model routing, authentication, billing, research, FORGE, Developer API, storage, or other defined platform services.

Timeline

The incident timeline should record meaningful changes, including:
  • first known impact
  • detection source
  • affected components
  • diagnostic findings
  • status changes
  • mitigation attempts
  • rollback or configuration changes
  • recovery verification
  • final resolution
Write updates as facts. If a root cause is not yet confirmed, say that explicitly.

Mitigation vs resolution

A mitigation reduces impact. Resolution means the service has been restored and recovery has been verified. Do not mark an incident resolved simply because a change was deployed. Observe the relevant health and error signals first.

Post-incident learning

For significant incidents, the timeline should make later review possible: what failed, how it was detected, how long impact lasted, what restored service, and what follow-up work reduces recurrence.

User communication

Operational incident records and public status communication can use the same underlying component/health concepts without exposing private internal diagnostics or security-sensitive details.

System health & diagnostics

Use health signals and dependency state to investigate and verify an incident.