Metrics — go-discord-caller¶
All telemetry is emitted via OpenTelemetry (OTLP gRPC) — there is no direct
Prometheus client. Instruments are created from a single meter in
internal/telemetry/ and exported to the OTLP endpoint set by OTEL_ENDPOINT
(empty disables telemetry — see internal/config/config.go). The collector /
Alloy / Prometheus OTLP receiver translates each OTel instrument into one or
more Prometheus series.
This file lists every instrument by its OTel name (the dotted name in code) alongside the Prometheus series name produced by the default OTLP→Prometheus translation.
OTel → Prometheus naming rules¶
The Prometheus names below assume the default OTLP translation (Prometheus
native OTLP receiver / OpenTelemetry Collector prometheus exporter, with unit
and type suffixes enabled — the defaults):
| OTel | Prometheus |
|---|---|
. and other illegal chars |
_ |
unit ms |
_milliseconds suffix |
unit s |
_seconds suffix |
| monotonic counter | _total suffix (not doubled if the name already ends in total) |
Float64Histogram |
_bucket / _sum / _count series |
UpDownCounter, Gauge, ObservableGauge |
gauge, no suffix |
Names vary if the exporter is configured with
without_units/without_type_suffix, or if Prometheus is run with--enable-feature=otlp-deltatocumulativestyle flags. The collector also adds atarget_infoseries andotel_scope_*labels. Tune accordingly.
All instruments carry the resource attribute service.name="go-discord-caller".
🤖 Bot / Discord entities (internal/telemetry/bot_metrics.go)¶
| OTel instrument | Type | Prometheus series | Attributes | Description |
|---|---|---|---|---|
gdc.discord.guild |
ObservableGauge | gdc_discord_guild |
guild_id, guild_name |
Info gauge (always 1) for known guilds. Emitted from internal/manager/service.go. |
gdc.bot.online |
ObservableGauge | gdc_bot_online |
user_id, guild_id |
1 while a bot is a registered member of the guild; absent otherwise. From internal/manager/service.go. |
gdc.voice.callers |
UpDownCounter | gdc_voice_callers |
guild_id, channel_id |
Users with the caller role currently in a voice channel. From internal/bot/handlers.go. |
gdc.command.started.total |
Counter | gdc_command_started_total |
command, guild_id ⚠️ |
Slash command invocations entering the handler. Recorded before the handler runs, so a handler that never returns is still counted. From internal/bot/middleware.go. |
gdc.command.total |
Counter | gdc_command_total |
command, guild_id ⚠️ |
Slash command invocations that completed. A sustained gdc_command_started_total - gdc_command_total means handlers went in and never came out. From internal/bot/middleware.go. |
gdc.command.duration |
Histogram (s) |
gdc_command_duration_seconds (_bucket/_sum/_count) |
command, guild_id ⚠️ |
Slash command execution duration. From internal/bot/middleware.go. |
⚠️ The three
gdc.command.*instruments label the guild asguild.id(dotted) in code — seeRecordCommand. The OTLP→Prometheus translation rewrites this toguild_id, so dashboards still queryguild_id, but the attribute is inconsistent with theguild_idused by every other instrument. Worth normalising in code.
🛰️ Speaker pool (internal/telemetry/pool_metrics.go)¶
Emitted from internal/pool/service.go (observable gauges via a registered
callback; counters inline in the watchdog).
| OTel instrument | Type | Prometheus series | Attributes | Description |
|---|---|---|---|---|
gdc.discord.bot |
ObservableGauge | gdc_discord_bot |
bot_id, bot_name |
Info gauge (always 1) for known speaker bots. |
gdc.pool.bots.total |
ObservableGauge | gdc_pool_bots_total |
— | Total speaker bots registered in the pool. (Gauge — not a counter despite the total suffix.) |
gdc.pool.bots.connected |
ObservableGauge | gdc_pool_bots_connected |
— | Speaker bots with a healthy gateway connection. |
gdc.bot.gateway.latency |
ObservableGauge (ms) |
gdc_bot_gateway_latency_milliseconds |
bot_id |
Gateway WebSocket heartbeat RTT per bot. 0 until first ACK. See docs/LATENCY.md — this is the control channel, not audio latency. |
gdc.pool.reconnect.attempts.total |
Counter | gdc_pool_reconnect_attempts_total |
bot_id |
Watchdog gateway reconnect attempts. |
gdc.pool.reconnect.failures.total |
Counter | gdc_pool_reconnect_failures_total |
bot_id |
Watchdog gateway reconnect failures. |
🏰 Session lifecycle & fanout (internal/telemetry/session_metrics.go)¶
Emitted from internal/manager/voice_raid.go (start/stop) and the auto-router
pipelines in internal/manager/pipeline/ (route transitions); frame drops from
the opus pipeline via FrameDropper.
| OTel instrument | Type | Prometheus series | Attributes | Description |
|---|---|---|---|---|
gdc.voice.sessions.active |
UpDownCounter | gdc_voice_sessions_active |
guild_id |
Currently active voice raid sessions. |
gdc.voice.session.start.total |
Counter | gdc_voice_session_start_total |
guild_id |
Voice raid starts. |
gdc.voice.session.stop.total |
Counter | gdc_voice_session_stop_total |
guild_id |
Voice raid stops. |
gdc.session.speakers |
Gauge | gdc_session_speakers |
guild_id |
Speakers (pool speaker bots plus the owner bot, if it joined) in the active raid; reset to 0 on stop. |
gdc.fanout.frames.dropped.total |
Counter | gdc_fanout_frames_dropped_total |
guild_id, path |
Opus frames dropped on full channels. path: mixer / direct / channel_mixer / relay_bridge / provider / receiver. |
gdc.session.route_transitions.total |
Counter | gdc_session_route_transitions_total |
guild_id, from, to |
Auto-router source-mode transitions. from/to: off / copy / mix. |
🎙️ Opus / mixer timing (internal/telemetry/opus_metrics.go)¶
All histograms, unit ms, attribute guild_id. Recorded on the hot path via a
pre-baked OpusRecorder (Metrics.ForGuild) — zero-alloc per frame.
| OTel instrument | Prometheus series (_bucket/_sum/_count) |
Buckets (ms) | Emitter | Description |
|---|---|---|---|---|
gdc.opus.receive.duration |
gdc_opus_receive_duration_milliseconds |
0.5, 1, 2, 5, 12, 20 | internal/opus/voice_receiver.go |
ReceiveOpusFrame execution time (excludes channel-wait). |
gdc.opus.provide.duration |
gdc_opus_provide_duration_milliseconds |
0.5, 1, 2, 5, 12, 20 | internal/opus/voice_provider.go |
ProvideOpusFrame drain+return time (excludes frame-wait). |
gdc.opus.allow_user.duration |
gdc_opus_allow_user_duration_milliseconds |
1, 2, 5, 12, 20 | internal/manager/allow_user.go |
allowUser filter time per evaluated frame. |
gdc.mixer.tick.duration |
gdc_mixer_tick_duration_milliseconds |
0.5, 1, 2, 5, 10, 20 | internal/opus/mixer.go |
Mixer tick processing time. |
gdc.mixer.pipeline.latency |
gdc_mixer_pipeline_latency_milliseconds |
10, 15, 20, 25, 30, 35, 40, 50, 60, 80, 120 | internal/opus/mixer.go |
Latency from Discord receive to the mixer tick that consumes the frame. Floor ~20 ms (tick period), ceiling 60 ms (audioSourceCap). See docs/LATENCY.md. |
Implementation¶
Instruments are defined in internal/telemetry/, split by subsystem:
bot_metrics.go, pool_metrics.go, session_metrics.go, opus_metrics.go.
metrics.go wires them together (NewMetrics); setup.go configures the OTLP
exporters (traces, metrics, logs) and the periodic metric reader (15 s interval).
Per-guild recorders are obtained via Metrics.ForGuild(ctx, guildID)
(guild_metrics.go), which bakes the guild_id attribute once and returns a
reusable GuildMetrics value — keeping the hot path allocation-free.
Example PromQL¶
# p99 mixer pipeline latency per guild (ms)
histogram_quantile(0.99,
sum by (guild_id, le) (rate(gdc_mixer_pipeline_latency_milliseconds_bucket[5m])))
# frame drop rate by pipeline stage
sum by (path) (rate(gdc_fanout_frames_dropped_total[5m]))
# speaker bots connected vs registered
gdc_pool_bots_connected / gdc_pool_bots_total
# active voice raids
sum(gdc_voice_sessions_active)
Traces: session startup phases¶
/start takes seconds before any audio flows (measured in production: ~0.8–1.9 s
with 2 speakers, ~3 s with 5, ~3.2–4.1 s with 10), and the work is a chain of
Discord round trips. voice.session / voice.session.guest spans last as long
as the raid, so the startup cost has its own child span with one child per round
trip. startPhase (internal/manager/trace.go) opens them.
voice.session (whole raid)
└── voice.session.setup (/start → first frame flowing)
├── voice.speakers.setup speaker.candidates, speaker.joined, capture
│ └── voice.bot.attach bot.id, bot.kind=speaker (one per speaker, concurrent)
│ ├── voice.conn.open voice handshake through SessionDescription
│ ├── voice.members.prefetch user.count (capture modes only)
│ ├── voice.conn.apply provider/receiver wiring + SetSpeaking op
│ └── voice.deaf.reconcile deaf.want, deaf.cached
├── voice.bot.attach bot.kind=owner (runs after every speaker)
│ └── … same four children
└── voice.pipeline.build local graph construction, expected ~0
What to read off a trace:
voice.bot.attach{bot.kind=owner}starting only after the last speaker's attach ends is the serialized owner handshake.voice.deaf.reconcilespans withdeaf.cached=falsestacking one after another are the per-guild REST bucket serializing member PATCHes — disgo holds the bucket mutex across each request.voice.conn.openis the irreducible part: op4 → voice server update → WSS+TLS → identify/ready → UDP discovery → session description.
Everything is sampled (default ParentBased(AlwaysSample)). The parent
voice.session span only exports when the raid ends, so a trace looks
incomplete while a raid is live — the startup children are already there.