Skip to main content

Reference

Metrics reference

Every metric Elevarq Signals exposes on its /metrics endpoint — type, labels, and meaning, plus the failure and skip reason vocabularies you can alert on.

Operational metrics exposed by Elevarq Signals under signals.metrics_path (default /metrics) when signals.metrics_enabled: true. This page is the operator-facing reference for what's published, what labels mean, and what to alert on. For the end-to-end Prometheus and Grafana setup, see Observability.

Scrape configuration

The endpoint inherits the daemon's bearer-auth and binds to the configured api.listen_addr. Prometheus scrape config:

scrape_configs:
  - job_name: signals
    metrics_path: /metrics
    scheme: http        # or https when the daemon terminates TLS
    static_configs:
      - targets: ['signals.host:8081']
    authorization:
      type: Bearer
      credentials_file: /etc/prometheus/signals.token

Operators that scrape from inside the same host typically rely on network-level controls (loopback binding) rather than the bearer token. Either is acceptable; the no-secrets-in-metric-labels invariant holds regardless.

Metric inventory

Every metric is published from a dedicated registry — none of the default Go runtime / process metrics are exposed. Names follow signals_<concern>_<unit>.

Collection cycle

MetricTypeLabelsNotes
signals_collection_cycles_totalCountertarget, statusOne increment per completed collection cycle. status ∈ {success, partial, failed}.
signals_collection_failures_totalCountertarget, reasonHard cycle failures by category. reason ∈ {connect_error, version_unsupported, timeout_setup, safety_check, persistence, internal}.
signals_collection_duration_secondsHistogramtarget, statusPer-cycle wall-clock duration. Buckets are exponential from 50ms.

Per-collector outcomes

MetricTypeLabelsNotes
signals_collectors_succeeded_totalCountertargetSum of per-cycle successful-collector counts.
signals_collectors_failed_totalCountertarget, reasonFailed collectors classified by reason: permission_denied, object_missing, timeout, execution_error.
signals_collectors_skipped_totalCountertarget, reasonSkipped collectors by reason: version_unsupported, extension_missing, config_disabled, budget_exhausted, privilege_owner_only, privilege_restricted, privilege_column_filtered.
signals_eligible_collectorsGaugetargetNumber of collectors that would run for this target after every gate is applied (version, extension, sensitivity, profile). Updated at the top of every cycle.

Snapshots

MetricTypeLabelsNotes
signals_last_successful_collection_timestampGaugetargetUnix-seconds of the most recent completed cycle. Use time() - <metric> for staleness.

Export

MetricTypeLabelsNotes
signals_export_requests_totalCounterstatusOne increment per /export request. status ∈ {ok, failed}.
signals_export_failures_totalCountererror_categoryFailures by category (invalid_time_format, invalid_target_id, invalid_time_range, snapshot_not_found, conflicting_selectors, internal).
signals_export_duration_secondsHistogramstatusPer-export wall-clock duration.

Daemon state

MetricTypeLabelsNotes
signals_sqlite_persistence_failures_totalCounter—Atomic-insert collection rollbacks.
signals_high_sensitivity_collectors_enabledGauge—1 if the daemon-wide high-sensitivity flag is on, 0 otherwise.
signals_circuit_stateGaugetarget, statePer-target circuit state. One row per (target, state); the active row has value 1. state ∈ {closed, open, paused}.

Label cardinality

LabelSourceCardinality bound
signals_instanceDaemon instance_id (DB meta), applied as a registry-wide const label on every metricOne value per running daemon.
targetsignals.targets[].name from configNumber of configured targets. Operator-bounded.
statusClosed daemon enum≤ 3 values per metric.
reasonClosed daemon enum≤ 5 values per metric.
stateCircuit-breaker enumExactly 3 values.
error_categoryClosed export error enum≤ 6 values.

Every series carries a signals_instance label whose value is the daemon's stable instance_id (the same identifier in export metadata and the /status payload). Scraping several daemons into one Prometheus, group or filter by it — e.g. label_values(signals_collection_cycles_total, signals_instance) drives the multi-instance dashboard's instance selector. It is a per-deployment constant, so it does not widen cardinality.

The no-secrets invariant is enforced at every label-write site: no SQL text, query payload, database name, hostname, username, secret, or customer data appears in any metric label.

The following are starting points — operators tune thresholds based on their poll interval and fleet shape.

Coverage / health

# Coverage drift: a target lost more than 10% of its collectors in
# the last 24h. Catches extension uninstalls, version downgrades,
# accidental profile changes.
(
  signals_eligible_collectors
  / signals_eligible_collectors offset 24h
) < 0.9
# Stale data: no completed cycle in 2x poll_interval (60s here).
time() - signals_last_successful_collection_timestamp > 120

# No metric at all = target never collected since daemon start.
absent(signals_last_successful_collection_timestamp{target="prod-db"})

Cycle health

# Failure ratio over the last 15m.
sum(rate(signals_collection_cycles_total{status="failed"}[15m])) by (target)
  / sum(rate(signals_collection_cycles_total[15m])) by (target)
  > 0.2

# Latency outlier.
histogram_quantile(0.95,
  rate(signals_collection_duration_seconds_bucket[15m])
) > 30

Operator-controlled state

# Any target in non-closed circuit state.
sum by (target, state) (signals_circuit_state{state!="closed"} == 1)

# Paused for more than an hour — operator may have forgotten.
signals_circuit_state{state="paused"} == 1
  unless (signals_circuit_state{state="paused"} offset 1h == 1)

Export

# Export error spike.
sum(rate(signals_export_failures_total[5m])) by (error_category) > 0

Persistence

# SQLite write failure — should never be > 0.
sum(rate(signals_sqlite_persistence_failures_total[15m])) > 0

What's deliberately NOT exposed

  • SQL text or query payload. Collector output never enters metrics. Operators that need that level of detail use the collected NDJSON.
  • Hostnames, dbnames, usernames. Targets are referenced by their operator-assigned name only.
  • Credentials / secret references. Per the no-secrets-in-labels invariant.
  • Per-collector duration histograms. Cardinality risk is high (≈60 collectors × N targets × bucket count). Per-collector outcomes are tracked as counters; durations are aggregated at the cycle level.

Workbench integration

The Elevarq Workbench Signals page consumes Signals state via the Elevarq Analyzer-mediated import contract, not directly off /metrics. This page is for operator consumption — Prometheus scrapes, on-call alerting, capacity planning. The analyzer-side ingest reads collector_status.json and snapshots.ndjson from exported ZIPs, not the metrics endpoint.