Guides

Alerts and metrics

Watch CPU, memory, and restarts in the dashboard, and page email or Slack with custom rules.

Deploy-failed and health-check-failed notifications are always on when you configure billing notification channels. Custom rules add restart loops, extra health sensitivity, and (when metrics are enabled) CPU/memory thresholds.

Service metrics

Open a service → Metrics. Charts cover the last 1h / 6h / 24h:

  • CPU cores and memory working set
  • Container restart count
  • Ingress request rate and p99 latency (web services with a public hostname)

Workers and cron jobs have empty request charts. If metrics are not configured on the API, the page explains that instead of showing zeros.

Custom alert rules

Workspace Settings → Alerts. Owner and admin can create, edit, disable, and delete rules. Developers can read them.

TypeFires when
Health failingThe service stays unhealthy for N consecutive evaluator ticks (default 3)
Restart loopRestarts in the window reach the count (default 5 in 10 minutes)
CPU high / Memory highUsage stays at or above the percent for the duration (default 90% for 5 minutes). Requires platform metrics.

Each rule can target one service or all non-preview services, and optionally a deployment environment (for example staging-only noise). Channels inherit the workspace email and Slack webhook; uncheck a box to skip that channel for the rule.

Cooldown is per service (default 60 minutes) so a workspace-wide rule can still notify about each failing app without repeating the same one. A monthly fire budget caps a noisy rule.

Built-in deploy and health emails still work if you never create a custom rule.

Customer traces

Settings → Tracing is optional. It injects OTEL_EXPORTER_OTLP_ENDPOINT into new releases so your app can export to a collector you operate. Do not point that URL at Bytstack’s platform Tempo.