[teamai] Push 87 resource(s) from XingfenD

This commit is contained in:
2026-09-10 16:10:45 +08:00
parent 425c9c078a
commit 65c04def51
1314 changed files with 211681 additions and 0 deletions
@@ -0,0 +1,186 @@
# Alerting
> See [metrics.md](metrics.md) for multi-window burn-rate SLO alerting and PromQL patterns for application metrics.
## The Four Golden Signals
Alert on what matters to users. Google's SRE book defines four golden signals — every Go service SHOULD have alerts covering all four:
| Signal | What it measures | Example metric | Alert trigger |
| --- | --- | --- | --- |
| **Latency** | Time to serve a request | `http_request_duration_seconds` (Histogram) | P99 > 2s for 5 minutes |
| **Traffic** | Demand on the system | `http_requests_total` (Counter) | Zero requests for 10 minutes |
| **Errors** | Rate of failed requests | `http_requests_total{status=~"5.."}` (Counter) | Error ratio > 1% for 5 minutes |
| **Saturation** | How full the system is | `db_connections_active / db_connections_max` | Pool > 90% saturated for 5 minutes |
## Awesome Prometheus Alerts
[awesome-prometheus-alerts](https://samber.github.io/awesome-prometheus-alerts/) is a curated collection of ~500 ready-to-use Prometheus alerting rules. This collection serves as a starting point for infrastructure and dependency alerting.
### Categories
| Category | Rules | Covers |
| --- | --: | --- |
| **Basic Resource Monitoring** | ~107 | Host metrics, Docker containers, hardware |
| **Databases and Brokers** | ~233 | PostgreSQL, MySQL, Redis, MongoDB, Kafka, RabbitMQ, etc. |
| **Reverse Proxies and Load Balancers** | ~45 | Nginx, Apache, HAProxy, Traefik |
| **Runtimes** | ~4 | PHP-FPM, JVM, Sidekiq |
| **Orchestrators** | ~74 | Kubernetes, Nomad, Consul, ArgoCD |
| **Network, Security, and Storage** | ~40 | Ceph, MinIO, SSL/TLS, DNS |
### How to Use It
The rules are organized by technology. Each rule is a ready-to-use Prometheus alerting rule in YAML format with customizable threshold values and `for:` durations.
1. **Find by technology** — locate the relevant database, message broker, or infrastructure component
2. **Adapt the YAML rule** — adjust threshold values (`> 0.01`, `> 100`, etc.) and the `for:` duration to match your SLOs and traffic patterns
3. **Place in Prometheus config** — add to the `prometheus/rules/` directory
### Integration Example
Prometheus loads alerting rules from YAML files referenced in its config. After copying rules from awesome-prometheus-alerts, place them in your rules directory:
```yaml
# prometheus/rules/postgresql.yml
groups:
- name: postgresql
rules:
# From awesome-prometheus-alerts — PostgreSQL section
- alert: PostgresqlDown
expr: pg_up == 0
for: 0m
labels:
severity: critical
annotations:
summary: "PostgreSQL down (instance {{ $labels.instance }})"
description: "PostgreSQL instance is down.\n VALUE = {{ $value }}"
- alert: PostgresqlTooManyConnections
expr: sum by (instance, datname) (pg_stat_activity_count{datname!~"template.*|postgres"}) > pg_settings_max_connections * 0.8
for: 2m
labels:
severity: warning
annotations:
summary: "PostgreSQL too many connections (> 80%) (instance {{ $labels.instance }})"
description: "PostgreSQL has {{ $value }} connections on {{ $labels.datname }}."
```
```yaml
# prometheus.yml
rule_files:
- "rules/*.yml"
```
### Workflow for New Dependencies
When adding a new infrastructure dependency (database, cache, message broker, reverse proxy) to a Go service:
1. [awesome-prometheus-alerts](https://samber.github.io/awesome-prometheus-alerts/) has alert rules organized by technology
2. Adapt the relevant alert rules and thresholds to your environment
3. Verify the exporter is deployed (e.g., `postgres_exporter`, `redis_exporter`) — the alerts depend on metrics from these exporters
4. Add the rules to your `prometheus/rules/` directory
## Go Runtime Alerts
The Prometheus Go client automatically exposes runtime metrics. Alert on these to catch resource leaks and GC pressure before they impact users.
```yaml
# prometheus/rules/go-runtime.yml
groups:
- name: go-runtime
rules:
# Goroutine leak — count growing steadily indicates a leak
# Diagnose: GET /debug/pprof/goroutine?debug=1 to see goroutine stack traces
- alert: GoroutineLeak
expr: go_goroutines > 1000
for: 10m
labels:
severity: warning
annotations:
summary: "High goroutine count (instance {{ $labels.instance }})"
description: "Goroutine count is {{ $value }}, possible leak."
# GC taking too long — P99 GC pause > 100ms degrades tail latency
- alert: HighGCDuration
expr: go_gc_duration_seconds{quantile="1"} > 0.1
for: 5m
labels:
severity: warning
annotations:
summary: "High GC duration (instance {{ $labels.instance }})"
description: "Max GC pause is {{ $value }}s. Check heap allocations."
# Heap growing unbounded — likely a memory leak
- alert: HighMemoryUsage
expr: go_memstats_alloc_bytes / go_memstats_sys_bytes > 0.9
for: 5m
labels:
severity: critical
annotations:
summary: "High memory usage (instance {{ $labels.instance }})"
description: "Allocated heap is {{ $value | humanizePercentage }} of system memory."
# Too many threads — usually caused by blocking syscalls or cgo
- alert: HighThreadCount
expr: go_threads > 500
for: 5m
labels:
severity: warning
annotations:
summary: "High OS thread count (instance {{ $labels.instance }})"
description: "Thread count is {{ $value }}. Check for blocking syscalls."
```
## Alert Severity Levels
Use two severity levels to separate "wake someone up" from "look at it tomorrow":
| Severity | Action | `for:` duration | Example |
| --- | --- | --- | --- |
| **critical** | Page on-call | 2-5 minutes | Service down, error rate > 5%, data loss risk |
| **warning** | Create ticket | 10-30 minutes | P99 latency high, connection pool > 80%, goroutine leak |
The `for:` duration controls how long a condition must be true before the alert fires. Short durations catch fast incidents but risk false positives from transient spikes. Long durations reduce noise but delay response.
**Guidelines:**
- Critical alerts: `for: 2m` to `for: 5m` — fast detection, wake someone up
- Warning alerts: `for: 10m` to `for: 30m` — confirmed trend, create a ticket
- NEVER set `for: 0m` on non-binary alerts — one bad scrape triggers a false page
- Binary alerts (service up/down) can use `for: 0m` or `for: 1m`
## Common Mistakes
```yaml
# Bad -- irate() is too volatile for alerts, reacts to a single scrape interval
# A brief spike or a single slow request triggers the alert
- alert: HighErrorRate
expr: irate(http_requests_total{status=~"5.."}[5m]) > 0.01
# Good -- rate() smooths over the full window, reducing false positives
- alert: HighErrorRate
expr: rate(http_requests_total{status=~"5.."}[5m]) / rate(http_requests_total[5m]) > 0.01
for: 5m
```
```yaml
# Bad -- no "for:" duration, fires on a single bad scrape
- alert: HighLatency
expr: histogram_quantile(0.99, rate(http_request_duration_seconds_bucket[5m])) > 2
# Good -- must be true for 5 minutes to fire
- alert: HighLatency
expr: histogram_quantile(0.99, rate(http_request_duration_seconds_bucket[5m])) > 2
for: 5m
```
```yaml
# Bad -- alerting on raw gauge without trend analysis (flaps constantly)
- alert: HighQueueDepth
expr: myapp_queue_messages_pending > 1000
# Good -- alert on sustained growth trend
- alert: HighQueueDepth
expr: myapp_queue_messages_pending > 1000
for: 10m
```
@@ -0,0 +1,23 @@
# Grafana Dashboards for Go Services
These community Grafana dashboards visualize Go runtime performance out of the box. They display the metrics automatically exposed by `github.com/prometheus/client_golang` — no custom instrumentation needed.
## Recommended Dashboards
| Dashboard | ID | What it shows |
| --- | --: | --- |
| [Go Host & Runtime Metrics](https://grafana.com/grafana/dashboards/21221-go-host-runtime-metrics-dashboard/) | 21221 | Host metrics + Go runtime (goroutines, heap, GC, threads) in one view |
| [Go Processes](https://grafana.com/grafana/dashboards/6671-go-processes/) | 6671 | Multi-process comparison — CPU, memory, goroutines, GC across all Go services |
| [Go Metrics](https://grafana.com/grafana/dashboards/10826-go-metrics/) | 10826 | Focused Go runtime view — memory breakdown, GC pauses, allocations, goroutines |
## How to Install
Dashboards are imported via **Dashboards > New > Import** in Grafana using the dashboard ID (e.g., `21221`), then selecting the Prometheus data source.
These dashboards require the default Go collector metrics (`go_goroutines`, `go_memstats_*`, `go_gc_duration_seconds`, `process_*`). If you use the Prometheus client library with default collectors, everything works out of the box.
## When to Use Each
- **21221** (Host & Runtime) — day-to-day monitoring of a single Go service alongside its host. Best as the default Go dashboard.
- **6671** (Go Processes) — comparing multiple Go services or replicas side by side. Useful during deployments to spot regressions across instances.
- **10826** (Go Metrics) — deep-diving into memory and GC behavior of a single service. Best for investigating performance issues.
@@ -0,0 +1,189 @@
# Structured Logging with `slog`
→ See `samber/cc-skills-golang@golang-error-handling` skill for the single handling rule.
## Why Structured Logging
Structured logs emit key-value pairs instead of freeform strings. Log management systems (Datadog, Grafana Loki, CloudWatch) can index, filter, and aggregate structured fields — something impossible with `log.Printf` output.
```go
// ✗ Bad — freeform string, impossible to filter by user_id
log.Printf("ERROR: failed to create user %s: %v", userID, err)
// ✓ Good — structured key-value pairs, machine-parseable
slog.Error("user creation failed",
"user_id", userID,
"error", err,
)
// JSON output: {"time":"2025-01-15T10:30:00Z","level":"ERROR","msg":"user creation failed","user_id":"u-123","error":"connection refused"}
```
## Handler Setup
```go
// Production MUST use JSON — because plain-text multiline logs (eg. stack traces) would be split into separate records by log collectors
logger := slog.New(slog.NewJSONHandler(os.Stdout, &slog.HandlerOptions{
Level: slog.LevelInfo,
}))
// Development — human-readable text
logger := slog.New(slog.NewTextHandler(os.Stderr, &slog.HandlerOptions{
Level: slog.LevelDebug,
}))
slog.SetDefault(logger)
```
## Log Levels
```go
slog.Debug("cache lookup", "key", cacheKey, "hit", false)
slog.Info("order created", "order_id", orderID, "total", amount)
slog.Warn("rate limit approaching", "current_usage", 0.92, "limit", 1000)
slog.Error("payment failed", "order_id", orderID, "error", err)
```
**Rule of thumb**: if you're unsure between Warn and Error, ask "did the operation succeed?" If yes (even with degradation), use Warn. If no, use Error.
## Cost of Logging
Logging is not free. Each log line costs CPU (serialization), I/O (disk/network), and money (log ingestion/storage in your aggregation platform). The cost scales with volume, which is directly controlled by log level.
- **Debug level in production** can generate millions of log lines per minute in a busy service, overwhelming your log pipeline and inflating costs by 10-100x
- **Info level** is the typical production default — it provides enough visibility without excessive volume
- Debug level SHOULD be disabled in production — use `slog.LevelInfo` in production and `slog.LevelDebug` only in development or when actively debugging a specific issue
- For high-throughput services, consider [samber/slog-sampling](https://github.com/samber/slog-sampling) to sample verbose logs (e.g., emit 1 in 100 Debug logs) rather than dropping them entirely
## Logging with Context
MUST use the `*Context` variants to correlate logs with the current trace. When an OpenTelemetry bridge is configured, trace_id and span_id are automatically injected into log records.
```go
// ✗ Bad — no trace correlation
slog.Error("query failed", "error", err)
// ✓ Good — trace_id/span_id attached automatically when OTel bridge is active
slog.ErrorContext(ctx, "query failed", "error", err)
```
## Adding Request-Scoped Attributes
Use `slog.With()` to create a child logger that includes attributes on every log line. Middleware can inject request-scoped fields so all downstream logs carry the same context.
```go
func LoggingMiddleware(next http.Handler) http.Handler {
return http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
logger := slog.With(
"request_id", r.Header.Get("X-Request-ID"),
"method", r.Method,
"path", r.URL.Path,
)
// Store enriched logger in context for downstream use
ctx := WithLogger(r.Context(), logger)
next.ServeHTTP(w, r.WithContext(ctx))
})
}
```
## Log Sinks and the `slog` Ecosystem
`slog` supports pluggable handlers. The Go community provides handlers for most log backends:
**Standard library:**
- `slog.JSONHandler` — JSON to stdout/stderr
- `slog.TextHandler` — human-readable key=value
**Log record handling:**
- [samber/slog-multi](https://github.com/samber/slog-multi) — fan-out to multiple handlers, routing, failover
- [samber/slog-sampling](https://github.com/samber/slog-sampling) — sample high-volume logs to reduce cost
- [samber/slog-formatter](https://github.com/samber/slog-formatter) — format/transform log attributes
**HTTP middleware:**
- [samber/slog-http](https://github.com/samber/slog-http) — HTTP server middleware (net/http, chi, fiber, echo, gin)
- [samber/slog-gin](https://github.com/samber/slog-gin) — Gin framework middleware
- [samber/slog-echo](https://github.com/samber/slog-echo) — Echo framework middleware
- [samber/slog-fiber](https://github.com/samber/slog-fiber) — Fiber framework middleware
- [samber/slog-chi](https://github.com/samber/slog-chi) — Chi router middleware
**Third-party log sinks** (see [go.dev/wiki/Resources-for-slog](https://go.dev/wiki/Resources-for-slog)):
- [lmittmann/tint](https://github.com/lmittmann/tint) — colorized terminal output
- [samber/slog-datadog](https://github.com/samber/slog-datadog) — send logs to Datadog
- [samber/slog-sentry](https://github.com/samber/slog-sentry) — send errors to Sentry
- [samber/slog-loki](https://github.com/samber/slog-loki) — send logs to Grafana Loki
- [samber/slog-nats](https://github.com/samber/slog-nats) — send logs to NATS
- [samber/slog-syslog](https://github.com/samber/slog-syslog) — send logs to syslog
- [samber/slog-fluentd](https://github.com/samber/slog-fluentd) — send logs to Fluentd
- [samber/slog-logrus](https://github.com/samber/slog-logrus) — bridge to Logrus
- [samber/slog-zap](https://github.com/samber/slog-zap) — bridge to Zap
- [samber/slog-zerolog](https://github.com/samber/slog-zerolog) — bridge to Zerolog
- [samber/slog-slack](https://github.com/samber/slog-slack) — send critical logs to Slack
## Migrating from zap / logrus / zerolog
`log/slog` is the standard library logger since Go 1.21. If the project uses `zap`, `logrus`, or `zerolog`, migrate to `slog` — it has a stable API, broad ecosystem support, and eliminates an unnecessary dependency.
**Step 1: Bridge** — route `slog` output through the existing logger so you can migrate call sites incrementally without changing log output:
```go
// Example: bridge slog → zap (same pattern for logrus/zerolog)
import slogzap "github.com/samber/slog-zap/v2"
zapLogger, _ := zap.NewProduction()
slog.SetDefault(slog.New(
slogzap.Option{Level: slog.LevelInfo, Logger: zapLogger}.NewZapHandler(),
))
```
Available bridges: [samber/slog-zap](https://github.com/samber/slog-zap), [samber/slog-logrus](https://github.com/samber/slog-logrus), [samber/slog-zerolog](https://github.com/samber/slog-zerolog)
**Step 2: Replace call sites** — change all logger calls to `slog`:
```go
// zap → slog
// Before: zap.L().Info("order created", zap.String("order_id", id))
// After:
slog.Info("order created", "order_id", id)
// logrus → slog
// Before: logrus.WithField("order_id", id).Info("order created")
// After:
slog.Info("order created", "order_id", id)
// zerolog → slog
// Before: log.Info().Str("order_id", id).Msg("order created")
// After:
slog.Info("order created", "order_id", id)
```
**Step 3: Remove the bridge** — once all call sites are migrated, replace the bridge handler with a native `slog` handler and remove the old logger dependency:
```go
slog.SetDefault(slog.New(slog.NewJSONHandler(os.Stdout, &slog.HandlerOptions{
Level: slog.LevelInfo,
})))
```
## Common Logging Mistakes
```go
// ✗ Bad — errors MUST be either logged OR returned, NEVER both (single handling rule violation)
if err != nil {
slog.Error("query failed", "error", err)
return fmt.Errorf("query: %w", err) // error gets logged twice up the chain
}
// ✓ Good — return with context, log at the top level
if err != nil {
return fmt.Errorf("querying users: %w", err)
}
// ✗ Bad — NEVER log PII (emails, SSNs, passwords, tokens)
slog.Info("user logged in", "email", user.Email, "ssn", user.SSN)
// ✓ Good — log identifiers, not sensitive data
slog.Info("user logged in", "user_id", user.ID)
```
@@ -0,0 +1,536 @@
# Metrics with Prometheus
→ See `samber/cc-skills-golang@golang-troubleshooting` skill for using metrics to diagnose production issues. → See `samber/cc-skills@promql-cli` skill for executing and testing PromQL queries via CLI.
When using the Prometheus client library, refer to the library's official documentation for up-to-date API signatures and examples.
## Metric Types
| Type | What it measures | Example | When to use |
| --- | --- | --- | --- |
| **Counter** | Cumulative total (only goes up) | Total requests, total errors | Counting events |
| **Gauge** | Current value (goes up and down) | In-flight requests, queue size, temperature | Current state |
| **Histogram** | Distribution of values in configurable buckets | Request duration, response size | Latency, sizes — when you need percentiles |
| **Summary** | Client-computed quantiles | Request duration (pre-computed P50, P99) | Rarely — prefer Histogram |
## Histogram vs Summary
This is one of the most common sources of confusion. Both measure distributions, but they work very differently.
**Histogram** stores observations in configurable buckets (e.g., 5ms, 10ms, 25ms, 50ms, 100ms, ...). Percentiles are computed at query time by Prometheus using `histogram_quantile()`. Because the raw bucket counts are stored server-side, histograms can be **aggregated across multiple instances** — essential for services running multiple replicas.
**Summary** computes quantiles (P50, P99, etc.) on the client side before sending them to Prometheus. This means the quantile values are pre-baked and **cannot be aggregated** — if you have 10 instances, you cannot combine their P99 values into a meaningful overall P99.
**Recommendation**: Histogram SHOULD be preferred over Summary in almost all cases. Summary is only useful when you need exact quantiles for a single instance and don't care about cross-instance aggregation.
## Tracking Percentiles (P50, P90, P99, P99.9)
Define a Histogram with appropriate buckets, then query percentiles with `histogram_quantile()`:
```go
import "github.com/prometheus/client_golang/prometheus"
import "github.com/prometheus/client_golang/prometheus/promauto"
var httpRequestDuration = promauto.NewHistogramVec(
prometheus.HistogramOpts{
Namespace: "myapp",
Subsystem: "http",
Name: "request_duration_seconds",
Help: "HTTP request duration in seconds.",
Buckets: prometheus.DefBuckets, // .005, .01, .025, .05, .1, .25, .5, 1, 2.5, 5, 10
},
[]string{"method", "path", "status"},
)
// In your handler or middleware:
func instrumentHandler(next http.Handler) http.Handler {
return http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
start := time.Now()
sw := &statusWriter{ResponseWriter: w, status: 200}
next.ServeHTTP(sw, r)
httpRequestDuration.WithLabelValues(
r.Method,
r.URL.Path,
strconv.Itoa(sw.status),
).Observe(time.Since(start).Seconds())
})
}
```
**PromQL queries for percentiles:**
```promql
# P50 (median) over the last 5 minutes
histogram_quantile(0.50, rate(myapp_http_request_duration_seconds_bucket[5m]))
# P90
histogram_quantile(0.90, rate(myapp_http_request_duration_seconds_bucket[5m]))
# P99
histogram_quantile(0.99, rate(myapp_http_request_duration_seconds_bucket[5m]))
# P99.9
histogram_quantile(0.999, rate(myapp_http_request_duration_seconds_bucket[5m]))
# P99 broken down by path
histogram_quantile(0.99, sum(rate(myapp_http_request_duration_seconds_bucket[5m])) by (le, path))
```
## Naming Conventions
Metric names MUST follow the [Prometheus naming best practices](https://prometheus.io/docs/practices/naming/). The pattern is: `<namespace>_<subsystem>_<name>_<unit>`
**Rules:**
- Use a single-word application prefix (namespace) relevant to the domain
- A metric must refer to a single unit and single quantity
- Include the unit as a suffix, in **plural** form
- MUST use **base units** — not derived units
**Always use base units:**
| Measurement | Use | Not |
| ------------- | -------------- | --------------------------- |
| Time | `_seconds` | `_milliseconds`, `_minutes` |
| Data size | `_bytes` | `_kilobytes`, `_megabytes` |
| Temperature | `_celsius` | `_fahrenheit` |
| Ratio/percent | `_ratio` (0–1) | `_percent` (0–100) |
| Mass | `_grams` | `_kilograms` |
**Suffix conventions:**
| Suffix | When to use | Example |
| --- | --- | --- |
| `_total` | Counters MUST use this suffix | `myapp_http_requests_total` |
| `_seconds` | Duration measurements | `myapp_http_request_duration_seconds` |
| `_bytes` | Data sizes | `myapp_response_size_bytes` |
| `_info` | Pseudo-metrics exposing metadata | `myapp_build_info` |
| `_created` | Creation timestamp of a counter | `myapp_http_requests_created` |
```go
// ✓ Good — namespace, subsystem, descriptive name, base unit suffix
myapp_http_requests_total // Counter
myapp_http_request_duration_seconds // Histogram — seconds, not milliseconds
myapp_http_response_size_bytes // Histogram — bytes, not kilobytes
myapp_db_connections_active // Gauge
myapp_queue_messages_pending // Gauge
process_cpu_seconds_total // Counter — total CPU time in seconds
// ✗ Bad
request_count // no namespace, no unit suffix
httpDuration // camelCase, no unit
request_duration_ms // milliseconds instead of seconds
myapp_request_size_kb // kilobytes instead of bytes
```
**Label naming:** do not embed label names into the metric name. Use labels to differentiate characteristics:
```go
// ✗ Bad — operation embedded in metric name
myapp_http_get_requests_total
myapp_http_post_requests_total
// ✓ Good — use a label
myapp_http_requests_total{method="GET"}
myapp_http_requests_total{method="POST"}
```
**Semantic consistency:** `sum()` or `avg()` over all label dimensions of a metric should be meaningful. If not, split into separate metrics.
## Exposing Metrics
```go
import "github.com/prometheus/client_golang/prometheus/promhttp"
mux.Handle("/metrics", promhttp.Handler())
```
## Document Metrics with PromQL Comments
EVERY METRIC declaration SHOULD include the relevant PromQL queries and alert rules as comments directly above the variable. This makes metrics self-documenting — when a developer reads the code, they immediately see how the metric is used in dashboards and alerts, without hunting through Grafana or alert configurations.
```go
// ✗ Bad — metric exists but nobody knows how to query or alert on it
var httpRequestsTotal = promauto.NewCounterVec(...)
// ✓ Good — PromQL queries and alert rules are part of the code
//
// Dashboard: rate(myapp_http_requests_total[5m])
// Dashboard: sum by (status) (rate(myapp_http_requests_total[5m]))
// Alert: sum(rate(myapp_http_requests_total{status=~"5.."}[5m])) / sum(rate(myapp_http_requests_total[5m])) > 0.01
var httpRequestsTotal = promauto.NewCounterVec(...)
```
This convention has practical benefits: PromQL queries are reviewed in PRs alongside the metric, queries stay in sync with metric changes (label renames, bucket changes), and new team members can understand the metric's purpose at a glance.
## Metric Examples and PromQL Queries
Production-ready metrics covering all four types with comprehensive PromQL for dashboards and alerts.
For infrastructure and dependency alerting (databases, caches, message brokers, reverse proxies, Kubernetes), [awesome-prometheus-alerts](https://samber.github.io/awesome-prometheus-alerts/) provides a curated collection of ~500 ready-to-use Prometheus alerting rules organized by technology. See [alerting.md](alerting.md) for integration details and Go runtime alerts.
NEVER use `irate(...)` for alerts — use `rate(...)` instead.
### Counters — tracking events
```go
// Dashboard: rate(myapp_http_requests_total[5m])
// Dashboard: sum by (status) (rate(myapp_http_requests_total[5m]))
// Dashboard: sum by (path) (rate(myapp_http_requests_total[5m]))
// Dashboard: topk(5, sum by (path) (rate(myapp_http_requests_total[5m])))
// Dashboard: increase(myapp_http_requests_total[1h])
// SLI: 1 - (sum(rate(myapp_http_requests_total{status=~"5.."}[5m])) / sum(rate(myapp_http_requests_total[5m])))
// Alert: sum(rate(myapp_http_requests_total{status=~"5.."}[5m])) / sum(rate(myapp_http_requests_total[5m])) > 0.01
// Alert: sum(rate(myapp_http_requests_total{status=~"5.."}[1m])) / sum(rate(myapp_http_requests_total[1m])) > 0.05
var httpRequestsTotal = promauto.NewCounterVec(
prometheus.CounterOpts{
Namespace: "myapp",
Subsystem: "http",
Name: "requests_total",
Help: "Total number of HTTP requests.",
},
[]string{"method", "path", "status"},
)
// Dashboard: sum by (type) (rate(myapp_errors_total[5m]))
// Dashboard: topk(3, sum by (type) (rate(myapp_errors_total[5m])))
// Alert: rate(myapp_errors_total{type="database"}[5m]) > 0.5
var errorsTotal = promauto.NewCounterVec(
prometheus.CounterOpts{
Namespace: "myapp",
Name: "errors_total",
Help: "Total number of errors by type.",
},
[]string{"type"}, // "database", "external_api", "validation"
)
// Dashboard: sum by (payment_method) (rate(myapp_orders_created_total[5m]))
// Dashboard: increase(myapp_orders_created_total[24h])
// Alert: rate(myapp_orders_created_total[30m]) == 0
var ordersCreated = promauto.NewCounterVec(
prometheus.CounterOpts{
Namespace: "myapp",
Subsystem: "orders",
Name: "created_total",
Help: "Total number of orders created.",
},
[]string{"payment_method"},
)
```
**Key PromQL patterns for counters:**
```promql
# Requests per second (smoothed over 5 minutes)
rate(myapp_http_requests_total[5m])
# Traffic by status code — see distribution of 2xx/4xx/5xx
sum by (status) (rate(myapp_http_requests_total[5m]))
# Top 5 busiest endpoints
topk(5, sum by (path) (rate(myapp_http_requests_total[5m])))
# Absolute request count in the last hour (useful for reports)
increase(myapp_http_requests_total[1h])
# Error ratio — fraction of requests returning 5xx (SLI)
sum(rate(myapp_http_requests_total{status=~"5.."}[5m]))
/
sum(rate(myapp_http_requests_total[5m]))
# 4xx error ratio — client errors (useful for spotting bad deployments)
sum(rate(myapp_http_requests_total{status=~"4.."}[5m]))
/
sum(rate(myapp_http_requests_total[5m]))
# Alert: error rate > 1% for 5 minutes (for: 5m)
sum(rate(myapp_http_requests_total{status=~"5.."}[5m]))
/
sum(rate(myapp_http_requests_total[5m]))
> 0.01
# Alert: spike detection — error rate > 5% over 1 minute (for: 2m)
sum(rate(myapp_http_requests_total{status=~"5.."}[1m]))
/
sum(rate(myapp_http_requests_total[1m]))
> 0.05
# Alert: zero orders for 30 minutes — business is broken (for: 30m)
rate(myapp_orders_created_total[30m]) == 0
```
### Gauges — tracking current state
```go
// Dashboard: myapp_http_in_flight_requests
// Alert: myapp_http_in_flight_requests > 500
var httpInFlightRequests = promauto.NewGauge(
prometheus.GaugeOpts{
Namespace: "myapp",
Subsystem: "http",
Name: "in_flight_requests",
Help: "Number of HTTP requests currently being processed.",
},
)
// Dashboard: myapp_db_connections_active
// Dashboard: myapp_db_connections_active / myapp_db_connections_max
// Alert: myapp_db_connections_active{pool="write"} / myapp_db_connections_max{pool="write"} > 0.9
// Alert: predict_linear(myapp_db_connections_active[15m], 600) > myapp_db_connections_max
var dbConnectionsActive = promauto.NewGaugeVec(
prometheus.GaugeOpts{
Namespace: "myapp",
Subsystem: "db",
Name: "connections_active",
Help: "Number of active database connections.",
},
[]string{"pool"}, // "read", "write"
)
var dbConnectionsMax = promauto.NewGaugeVec(
prometheus.GaugeOpts{
Namespace: "myapp",
Subsystem: "db",
Name: "connections_max",
Help: "Maximum database connections in the pool.",
},
[]string{"pool"},
)
// Dashboard: myapp_queue_messages_pending
// Dashboard: deriv(myapp_queue_messages_pending[5m])
// Alert: myapp_queue_messages_pending{queue_name="orders"} > 1000
// Alert: deriv(myapp_queue_messages_pending[10m]) > 50
// Alert: predict_linear(myapp_queue_messages_pending[30m], 3600) > 10000
var queueSize = promauto.NewGaugeVec(
prometheus.GaugeOpts{
Namespace: "myapp",
Subsystem: "queue",
Name: "messages_pending",
Help: "Number of messages waiting to be processed.",
},
[]string{"queue_name"},
)
// Dashboard: myapp_workers_active / myapp_workers_max
// Alert: myapp_workers_active / myapp_workers_max > 0.8
var workersActive = promauto.NewGauge(
prometheus.GaugeOpts{
Namespace: "myapp",
Name: "workers_active",
Help: "Number of worker goroutines currently processing jobs.",
},
)
// Usage in middleware:
func instrumentMiddleware(next http.Handler) http.Handler {
return http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
httpInFlightRequests.Inc()
defer httpInFlightRequests.Dec()
next.ServeHTTP(w, r)
})
}
```
**Key PromQL patterns for gauges:**
```promql
# Current value — gauges are queried directly
myapp_http_in_flight_requests
# Saturation — what fraction of the pool is in use
myapp_db_connections_active{pool="write"} / myapp_db_connections_max{pool="write"}
# Rate of change — is the queue growing or shrinking? (items/second)
deriv(myapp_queue_messages_pending[5m])
# Prediction — will the connection pool be exhausted in 10 minutes?
# predict_linear extrapolates the trend from the last 15 minutes
predict_linear(myapp_db_connections_active[15m], 600) > myapp_db_connections_max
# Prediction — will the queue exceed 10k items in 1 hour?
predict_linear(myapp_queue_messages_pending[30m], 3600) > 10000
# Alert: connection pool > 90% saturated (for: 5m)
myapp_db_connections_active{pool="write"} / myapp_db_connections_max{pool="write"} > 0.9
# Alert: queue depth growing faster than 50 items/sec (for: 10m)
deriv(myapp_queue_messages_pending[10m]) > 50
# Alert: worker pool saturated (for: 5m)
myapp_workers_active / myapp_workers_max > 0.8
```
### Histograms — tracking distributions (recommended for latency)
```go
// Dashboard: histogram_quantile(0.50, sum(rate(myapp_http_request_duration_seconds_bucket[5m])) by (le))
// Dashboard: histogram_quantile(0.90, sum(rate(myapp_http_request_duration_seconds_bucket[5m])) by (le))
// Dashboard: histogram_quantile(0.99, sum(rate(myapp_http_request_duration_seconds_bucket[5m])) by (le))
// Dashboard: histogram_quantile(0.99, sum(rate(myapp_http_request_duration_seconds_bucket[5m])) by (le, path))
// SLI: sum(rate(myapp_http_request_duration_seconds_bucket{le="0.3"}[5m])) / sum(rate(myapp_http_request_duration_seconds_count[5m]))
// Alert: histogram_quantile(0.99, sum(rate(myapp_http_request_duration_seconds_bucket[5m])) by (le)) > 2
var httpRequestDuration = promauto.NewHistogramVec(
prometheus.HistogramOpts{
Namespace: "myapp",
Subsystem: "http",
Name: "request_duration_seconds",
Help: "HTTP request duration in seconds.",
Buckets: []float64{.005, .01, .025, .05, .1, .25, .5, 1, 2.5, 5, 10},
},
[]string{"method", "path", "status"},
)
// Dashboard: histogram_quantile(0.95, sum(rate(myapp_external_call_duration_seconds_bucket[5m])) by (le, service))
// Alert: histogram_quantile(0.99, sum(rate(myapp_external_call_duration_seconds_bucket[5m])) by (le, service)) > 5
var externalAPICallDuration = promauto.NewHistogramVec(
prometheus.HistogramOpts{
Namespace: "myapp",
Subsystem: "external",
Name: "call_duration_seconds",
Help: "Duration of external API calls in seconds.",
Buckets: []float64{.01, .05, .1, .25, .5, 1, 2.5, 5, 10, 30},
},
[]string{"service", "endpoint"},
)
// Dashboard: histogram_quantile(0.95, sum(rate(myapp_orders_amount_dollars_bucket[5m])) by (le))
var orderAmount = promauto.NewHistogramVec(
prometheus.HistogramOpts{
Namespace: "myapp",
Subsystem: "orders",
Name: "amount_dollars",
Help: "Order amount in dollars.",
Buckets: []float64{1, 5, 10, 25, 50, 100, 250, 500, 1000, 5000},
},
[]string{"payment_method"},
)
```
**Key PromQL patterns for histograms:**
```promql
# Percentile latencies — the core latency dashboard
histogram_quantile(0.50, sum(rate(myapp_http_request_duration_seconds_bucket[5m])) by (le)) # P50
histogram_quantile(0.90, sum(rate(myapp_http_request_duration_seconds_bucket[5m])) by (le)) # P90
histogram_quantile(0.95, sum(rate(myapp_http_request_duration_seconds_bucket[5m])) by (le)) # P95
histogram_quantile(0.99, sum(rate(myapp_http_request_duration_seconds_bucket[5m])) by (le)) # P99
histogram_quantile(0.999, sum(rate(myapp_http_request_duration_seconds_bucket[5m])) by (le)) # P99.9
# P99 latency broken down by endpoint — find the slowest paths
histogram_quantile(0.99, sum(rate(myapp_http_request_duration_seconds_bucket[5m])) by (le, path))
# Average latency (mean) — useful alongside percentiles
sum(rate(myapp_http_request_duration_seconds_sum[5m]))
/
sum(rate(myapp_http_request_duration_seconds_count[5m]))
# Apdex-like SLI — fraction of requests under 300ms (target threshold)
sum(rate(myapp_http_request_duration_seconds_bucket{le="0.3"}[5m]))
/
sum(rate(myapp_http_request_duration_seconds_count[5m]))
# Request throughput from histogram (requests/sec)
sum(rate(myapp_http_request_duration_seconds_count[5m]))
# External API P95 latency per service
histogram_quantile(0.95, sum(rate(myapp_external_call_duration_seconds_bucket[5m])) by (le, service))
# Alert: P99 latency > 2s (for: 5m)
histogram_quantile(0.99, sum(rate(myapp_http_request_duration_seconds_bucket[5m])) by (le)) > 2
# Alert: P95 latency > 500ms (for: 10m)
histogram_quantile(0.95, sum(rate(myapp_http_request_duration_seconds_bucket[5m])) by (le)) > 0.5
# Alert: external API P99 > 5s (for: 5m)
histogram_quantile(0.99, sum(rate(myapp_external_call_duration_seconds_bucket[5m])) by (le, service)) > 5
# Alert: less than 95% of requests under 300ms (SLO breach) (for: 10m)
(
sum(rate(myapp_http_request_duration_seconds_bucket{le="0.3"}[5m]))
/
sum(rate(myapp_http_request_duration_seconds_count[5m]))
) < 0.95
```
### Summary — client-side quantiles (use sparingly)
Summaries compute quantiles on the client and cannot be aggregated across instances. Use them only for single-process diagnostics where exact quantiles matter. Prefer Histogram in all other cases.
```go
// Dashboard: myapp_jobs_processing_seconds{quantile="0.5"}
// Dashboard: myapp_jobs_processing_seconds{quantile="0.99"}
// Note: these quantiles CANNOT be aggregated across instances
var jobProcessingDuration = promauto.NewSummary(
prometheus.SummaryOpts{
Namespace: "myapp",
Subsystem: "jobs",
Name: "processing_seconds",
Help: "Job processing duration in seconds.",
Objectives: map[float64]float64{0.5: 0.05, 0.9: 0.01, 0.99: 0.001},
MaxAge: 10 * time.Minute,
},
)
```
## Multi-Window Burn-Rate SLO Alerting
For critical services, simple threshold alerts ("error rate > 1%") fire too late for fast incidents and too early for slow ones. Multi-window burn-rate alerting scales alert urgency to how fast you're consuming your error budget.
For a **99.9% availability SLO** (0.1% error budget over 30 days):
| Window | Burn rate | Error rate | Severity | Meaning |
| --- | --- | --- | --- | --- |
| 5m + 1h | 14.4x | > 1.44% | Critical (page) | Budget exhausted in ~2 days |
| 30m + 6h | 6x | > 0.6% | Critical (page) | Budget exhausted in ~5 days |
| 2h + 24h | 1x | > 0.1% | Warning (ticket) | On track to exhaust budget |
```promql
# Fast burn — page immediately (for: 2m)
# Both short and long windows must fire to avoid noise from brief spikes
(
(1 - sum(rate(myapp_http_requests_total{status=~"2.."}[5m])) / sum(rate(myapp_http_requests_total[5m]))) > 0.0144
and
(1 - sum(rate(myapp_http_requests_total{status=~"2.."}[1h])) / sum(rate(myapp_http_requests_total[1h]))) > 0.0144
)
# Medium burn — page (for: 15m)
(
(1 - sum(rate(myapp_http_requests_total{status=~"2.."}[30m])) / sum(rate(myapp_http_requests_total[30m]))) > 0.006
and
(1 - sum(rate(myapp_http_requests_total{status=~"2.."}[6h])) / sum(rate(myapp_http_requests_total[6h]))) > 0.006
)
# Slow burn — ticket (for: 1h)
(
(1 - sum(rate(myapp_http_requests_total{status=~"2.."}[2h])) / sum(rate(myapp_http_requests_total[2h]))) > 0.001
and
(1 - sum(rate(myapp_http_requests_total{status=~"2.."}[24h])) / sum(rate(myapp_http_requests_total[24h]))) > 0.001
)
```
The short window catches the incident fast; the long window confirms it's sustained. Together they eliminate false positives from transient blips.
## High-Cardinality Labels
NEVER use high-cardinality labels (user IDs, full URLs, request IDs). Every unique combination of label values creates a separate time series in Prometheus. Unbounded labels cause memory explosion on the Prometheus server, slow queries, and can crash the monitoring stack.
```go
// ✗ Bad — unbounded cardinality (millions of unique values)
httpRequestsTotal.WithLabelValues(r.URL.Path) // /users/alice, /users/bob, /users/charlie...
httpRequestsTotal.WithLabelValues(userID) // one series per user
httpRequestsTotal.WithLabelValues(r.Header.Get("X-Request-ID")) // one series per request!
// ✓ Good — bounded, normalized labels
httpRequestsTotal.WithLabelValues(routePattern) // /users/:id (the route template, not the actual path)
httpRequestsTotal.WithLabelValues(r.Method) // GET, POST, PUT, DELETE (5 values)
httpRequestsTotal.WithLabelValues(statusBucket) // "2xx", "3xx", "4xx", "5xx" (4 values)
```
**How to limit cardinality:**
- Use route templates (`/users/:id`) instead of actual paths (`/users/alice`)
- Bucket status codes (`2xx`, `4xx`, `5xx`) instead of exact codes (200, 201, 204, 400, 401, ...)
- Never use user IDs, request IDs, session IDs, or email addresses as labels
- Use attributes/tags in traces instead — traces handle high cardinality naturally
- **Rule of thumb**: if a label can have more than ~100 unique values, it's too many
@@ -0,0 +1,72 @@
# Profiling and Continuous Profiling
→ See `samber/cc-skills-golang@golang-troubleshooting` skill (pprof.md) for on-demand debugging.
## What Profiling Is
Profiling analyzes the runtime behavior of your program — where CPU time is spent, how memory is allocated, which goroutines are blocked, and where lock contention occurs. While metrics tell you "the service is slow," profiling tells you "this specific function on line 42 is the bottleneck."
## On-Demand Profiling with `pprof`
pprof endpoints MUST be protected with basic auth — NEVER expose them publicly. They leak sensitive runtime information and can be abused for DoS.
→ See `samber/cc-skills-golang@golang-troubleshooting` pprof.md for the full pprof CLI reference (profile types, capturing, analyzing, commands).
## Continuous Profiling with Pyroscope
On-demand profiling requires you to be there when the problem happens. Continuous profiling runs always-on in the background with low overhead (~2-5% CPU), so you can look at profiles after the fact. Toggle it with an environment variable.
```go
import "github.com/grafana/pyroscope-go"
func setupContinuousProfiling() {
if os.Getenv("PROFILING_ENABLED") != "true" {
return
}
_, err := pyroscope.Start(pyroscope.Config{
ApplicationName: "my-service",
ServerAddress: os.Getenv("PYROSCOPE_URL"), // e.g., http://user:pass@pyroscope:4040
ProfileTypes: []pyroscope.ProfileType{
pyroscope.ProfileCPU,
pyroscope.ProfileAllocObjects,
pyroscope.ProfileAllocSpace,
pyroscope.ProfileInuseObjects,
pyroscope.ProfileInuseSpace,
pyroscope.ProfileGoroutines,
pyroscope.ProfileMutexCount,
pyroscope.ProfileMutexDuration,
pyroscope.ProfileBlockCount,
pyroscope.ProfileBlockDuration,
},
})
if err != nil {
slog.Error("failed to start pyroscope", "error", err)
} else {
slog.Info("continuous profiling enabled", "server", os.Getenv("PYROSCOPE_URL"))
}
}
```
## Cost of Continuous Profiling
Continuous profiling adds overhead to every running instance — CPU for collecting stack samples, memory for buffering, and network for transmitting profiles to the backend. While typically low (~2-5% CPU), this cost is **per-instance and always-on**.
**Cost factors:**
- **CPU overhead** — profiling itself consumes CPU cycles. In CPU-bound services, even 2-5% overhead matters.
- **Network/storage** — profile data is continuously shipped to Pyroscope/your backend. High-replica services multiply this.
- **All profile types enabled** — each additional profile type (mutex, block, goroutine) adds incremental overhead.
**Mitigation:**
- Toggle via environment variable (`PROFILING_ENABLED`) — enable only when needed or on a subset of instances
- Start with CPU + heap profiles only; add mutex/block/goroutine profiles when investigating specific issues
- In large deployments, enable continuous profiling on a fraction of replicas (e.g., 1 in 10) rather than all of them
## When to Profile
1. Metrics show high CPU/memory usage → look at CPU/heap profiles
2. P99 latency spikes → CPU profile + mutex profile to find contention
3. Goroutine count growing → goroutine profile to find leaks
4. Before and after an optimization → compare profiles to verify improvement
@@ -0,0 +1,258 @@
# Real User Monitoring (RUM) and Product Observability
## What RUM Is
Backend observability (logs, metrics, traces, profiles) tells you how your **system** behaves. RUM tells you how your **users** experience it. While frontend SDKs capture browser-side signals, the Go backend plays a critical role: tracking server-side business events, feeding Customer Data Platforms, and correlating user sessions with backend traces.
## RUM Capabilities
| Capability | What it reveals | Example tools |
| --- | --- | --- |
| **Product Analytics** | What users do — page views, clicks, feature adoption, retention | PostHog, Amplitude, Mixpanel |
| **Funnel Analysis** | Where users drop off in multi-step flows (signup, checkout, onboarding) | PostHog, Amplitude, Mixpanel |
| **CDP** | Unified user profile from all data sources — events, properties, segments | Segment, RudderStack |
## Identity Key: Use `user_id`, Never Email
The distinct_id (identity key) used across all RUM tracking MUST be your internal, immutable `user_id`. NEVER use email addresses.
```go
// ✗ Bad — email is mutable, PII, and breaks analytics when users change it
posthogClient.Enqueue(posthog.Capture{
DistinctId: user.Email, // "alice@example.com" → user changes email → events split into two users
Event: "order_completed",
})
// ✓ Good — user_id is immutable, stable, and not PII
posthogClient.Enqueue(posthog.Capture{
DistinctId: user.ID, // "usr_a1b2c3" — never changes, always the same user
Event: "order_completed",
})
```
**Why email is a bad identity key:**
- **Mutable** — users change their email. Events before and after the change appear as two different users, breaking funnels, retention analysis, and cohort tracking.
- **PII** — using email as the identity key means every event, session recording, and analytics query contains personally identifiable information. This complicates GDPR/CCPA compliance — you can't anonymize analytics without losing user identity.
- **Non-unique across systems** — the same email might belong to different accounts in different services or environments.
- **Leaks into third-party systems** — the distinct_id is sent to your analytics platform (PostHog, Segment, etc.). If it's an email, you've shared PII with every vendor in your analytics pipeline.
Use `user_id` as the identity key everywhere: PostHog `DistinctId`, Segment `UserId`, Amplitude `user_id`. Store email as a user property if needed for display, never as the primary key.
## Backend Role in RUM
The Go backend tracks server-side events, correlates sessions with traces, and feeds data into CDPs.
### 1. Server-Side Event Tracking
When critical business events happen server-side (payment completed, subscription upgraded, email sent), track them from Go so they appear in the same analytics pipeline as frontend events.
```go
import "github.com/posthog/posthog-go"
var posthogClient posthog.Client
func initPostHog() {
var err error
posthogClient, err = posthog.NewWithConfig(
os.Getenv("POSTHOG_API_KEY"),
posthog.Config{Endpoint: os.Getenv("POSTHOG_HOST")},
)
if err != nil {
slog.Error("failed to init PostHog", "error", err)
}
}
func (s *OrderService) Complete(ctx context.Context, order Order) error {
// ... business logic ...
// Track server-side event — appears alongside frontend events in PostHog
posthogClient.Enqueue(posthog.Capture{
DistinctId: order.UserID, // immutable user_id, not email
Event: "order_completed",
Properties: posthog.NewProperties().
Set("order_id", order.ID).
Set("amount", order.Total).
Set("payment_method", order.PaymentMethod).
Set("item_count", len(order.Items)),
})
return nil
}
```
### 2. Connecting Frontend Sessions to Backend Traces
Pass the frontend session ID or distinct ID through HTTP headers so backend traces can be correlated with RUM sessions. When a user reports "the page was slow," you can find their session recording AND the backend trace for the same request.
```go
func TracingMiddleware(next http.Handler) http.Handler {
return http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
ctx := r.Context()
span := trace.SpanFromContext(ctx)
// Attach RUM session ID to the backend span
if sessionID := r.Header.Get("X-Session-ID"); sessionID != "" {
span.SetAttributes(attribute.String("rum.session_id", sessionID))
}
// Attach analytics distinct ID for user correlation
if distinctID := r.Header.Get("X-Distinct-ID"); distinctID != "" {
span.SetAttributes(attribute.String("rum.distinct_id", distinctID))
}
next.ServeHTTP(w, r.WithContext(ctx))
})
}
```
### 3. CDP Event Ingestion
If you use a Customer Data Platform (Segment, RudderStack), the Go backend sends events through the CDP's server-side SDK. The CDP unifies these with frontend events into a single user profile.
```go
import "github.com/segmentio/analytics-go/v3"
var segmentClient analytics.Client
func initSegment() {
segmentClient = analytics.New(os.Getenv("SEGMENT_WRITE_KEY"))
}
func (s *UserService) Upgrade(ctx context.Context, userID string, plan string) error {
// ... business logic ...
// Track through CDP — unified with frontend events
segmentClient.Enqueue(analytics.Track{
UserId: userID, // immutable user_id, not email
Event: "plan_upgraded",
Properties: analytics.NewProperties().
Set("plan", plan).
Set("source", "api"),
})
// Update user profile in CDP
segmentClient.Enqueue(analytics.Identify{
UserId: userID,
Traits: analytics.NewTraits().
Set("plan", plan).
Set("upgraded_at", time.Now()),
})
return nil
}
```
## GDPR and CCPA Compliance
RUM collects user behavior data — clicks, page views, session recordings. This triggers privacy regulation requirements. Compliance is not optional; violations carry heavy fines (GDPR: up to 4% of global revenue, CCPA: $7,500 per intentional violation).
### Consent Management
GDPR/CCPA consent SHOULD be obtained before loading RUM SDKs or sending tracking events. This applies to both frontend scripts and server-side event tracking.
```go
// Server-side: check consent before tracking
func (s *OrderService) Complete(ctx context.Context, order Order) error {
// ... business logic ...
// Only track if user has consented to analytics
consent := auth.ConsentFromContext(ctx)
if consent.Analytics {
posthogClient.Enqueue(posthog.Capture{
DistinctId: order.UserID,
Event: "order_completed",
Properties: posthog.NewProperties().
Set("order_id", order.ID).
Set("amount", order.Total),
})
}
return nil
}
```
### Data Subject Rights Endpoints
GDPR and CCPA require you to let users access, export, and delete their data. Implement API endpoints that propagate these requests to all systems that hold user data — your database, your analytics platform, your CDP.
```go
// DELETE /api/users/:id/data — GDPR Article 17 "Right to Erasure"
func (h *PrivacyHandler) HandleDataDeletion(w http.ResponseWriter, r *http.Request) {
ctx := r.Context()
userID := chi.URLParam(r, "id")
// 1. Delete from your database
if err := h.userRepo.DeleteAllData(ctx, userID); err != nil {
slog.ErrorContext(ctx, "failed to delete user data", "user_id", userID, "error", err)
http.Error(w, "internal error", http.StatusInternalServerError)
return
}
// 2. Delete from analytics platform
if err := h.posthog.DeleteUser(ctx, userID); err != nil {
slog.ErrorContext(ctx, "failed to delete analytics data", "user_id", userID, "error", err)
}
// 3. Delete from CDP
if err := h.segment.DeleteUser(ctx, userID); err != nil {
slog.ErrorContext(ctx, "failed to delete CDP data", "user_id", userID, "error", err)
}
slog.InfoContext(ctx, "user data deletion completed", "user_id", userID)
w.WriteHeader(http.StatusNoContent)
}
// GET /api/users/:id/data — GDPR Article 15 "Right of Access"
func (h *PrivacyHandler) HandleDataExport(w http.ResponseWriter, r *http.Request) {
ctx := r.Context()
userID := chi.URLParam(r, "id")
export, err := h.userRepo.ExportAllData(ctx, userID)
if err != nil {
slog.ErrorContext(ctx, "failed to export user data", "user_id", userID, "error", err)
http.Error(w, "internal error", http.StatusInternalServerError)
return
}
w.Header().Set("Content-Type", "application/json")
json.NewEncoder(w).Encode(export)
}
```
### Privacy Checklist
- [ ] **Consent before tracking** — no analytics scripts load and no server-side events fire until the user consents
- [ ] **Consent + cookie banner** — clear opt-in (not pre-checked boxes), separate consent for analytics vs marketing vs functional (frontend responsibility, but backend must respect the consent flag)
- [ ] **Data minimization** — only collect what you need, never track PII in analytics events
- [ ] **Data retention policy** — auto-delete old analytics data (e.g., 2 years for aggregated analytics)
- [ ] **Data subject rights** — endpoints for data export (right of access) and deletion (right to erasure)
- [ ] **Data processing agreements** — signed DPAs with all third-party analytics/CDP vendors
- [ ] **Privacy policy** — lists all RUM tools, what data they collect, and how long it's retained
- [ ] **Identity key is not PII** — use `user_id`, not email, as the distinct_id across all platforms
- [ ] **Self-hosted option** — consider self-hosting (PostHog, Matomo) to keep data in your infrastructure and simplify compliance
## Self-Hosted vs SaaS
| Factor | Self-hosted (PostHog, Matomo) | SaaS (Amplitude, Mixpanel) |
| --- | --- | --- |
| **Data residency** | Full control — data stays in your infra | Data on vendor's servers |
| **GDPR compliance** | Simpler — no cross-border data transfer | Requires DPA, SCCs, or adequacy decision |
| **Cost** | Infrastructure cost, scales with volume | Per-event or per-seat pricing |
| **Maintenance** | You manage upgrades, scaling, backups | Vendor handles everything |
| **Features** | Catching up but improving fast | Often more polished and feature-rich |
For EU-focused products or strict data residency requirements, self-hosting PostHog is the pragmatic choice — it eliminates most GDPR concerns around cross-border data transfer.
## Cost of RUM
RUM costs scale with **event volume**:
- **Event-based pricing** — every page view, click, and custom event counts. A busy SaaS app can generate millions of events/month per user segment.
- **CDP costs** — CDPs charge per tracked user and per event. Segment at scale can cost more than your entire backend infrastructure.
**Mitigation:**
- Use server-side event filtering to drop low-value events before they reach the analytics platform
- Self-host where possible to convert per-event pricing into fixed infrastructure cost
- Set data retention limits on aggregated analytics
@@ -0,0 +1,198 @@
# Distributed Tracing with OpenTelemetry
→ See `samber/cc-skills-golang@golang-context` skill for propagating context across service boundaries. → See `samber/cc-skills-golang@golang-samber-oops` skill for structured errors with stack traces in spans.
When using the OpenTelemetry Go SDK, refer to the library's official documentation for up-to-date API signatures and examples.
## Why Tracing
When a request crosses multiple services, logs from each service are isolated. Tracing connects them: a single trace shows the full request path with timing for every operation. This is how you answer "why was this request slow?" in a microservices architecture.
## OTel SDK Setup
Set up the TracerProvider early in your application. On new projects, do this first — then add spans everywhere incrementally.
```go
import (
"go.opentelemetry.io/otel"
"go.opentelemetry.io/otel/exporters/otlp/otlptrace/otlptracegrpc"
"go.opentelemetry.io/otel/sdk/resource"
sdktrace "go.opentelemetry.io/otel/sdk/trace"
semconv "go.opentelemetry.io/otel/semconv/v1.26.0"
)
func initTracer(ctx context.Context) (func(), error) {
exporter, err := otlptracegrpc.New(ctx)
if err != nil {
return nil, fmt.Errorf("creating OTLP exporter: %w", err)
}
res, err := resource.New(ctx,
resource.WithAttributes(
semconv.ServiceNameKey.String("my-service"),
semconv.ServiceVersionKey.String("1.0.0"),
),
)
if err != nil {
return nil, fmt.Errorf("creating resource: %w", err)
}
tp := sdktrace.NewTracerProvider(
sdktrace.WithBatcher(exporter),
sdktrace.WithResource(res),
)
otel.SetTracerProvider(tp)
shutdown := func() {
_ = tp.Shutdown(context.Background())
}
return shutdown, nil
}
```
## Creating Spans
Every meaningful operation should have a span. Think of spans as the building blocks of a trace — they show where time was spent.
```go
import "go.opentelemetry.io/otel"
var tracer = otel.Tracer("myapp/order-service")
func (s *OrderService) Create(ctx context.Context, req CreateOrderRequest) (*Order, error) {
ctx, span := tracer.Start(ctx, "OrderService.Create")
defer span.End()
// Add attributes that help with debugging
span.SetAttributes(
attribute.String("order.payment_method", req.PaymentMethod),
attribute.Float64("order.amount", req.Amount),
)
order, err := s.repo.Insert(ctx, req.ToOrder())
if err != nil {
span.RecordError(err)
span.SetStatus(codes.Error, err.Error())
return nil, fmt.Errorf("inserting order: %w", err)
}
return order, nil
}
func (r *OrderRepo) Insert(ctx context.Context, order Order) (*Order, error) {
ctx, span := tracer.Start(ctx, "OrderRepo.Insert")
defer span.End()
_, err := r.db.ExecContext(ctx, "INSERT INTO orders ...", order.ID)
if err != nil {
span.RecordError(err)
span.SetStatus(codes.Error, err.Error())
return nil, fmt.Errorf("exec insert: %w", err)
}
return &order, nil
}
```
**Where to add spans** — spans MUST be created for:
- Every service method (business logic layer)
- Every database query
- Every external API call
- Every message queue publish/consume
- Any operation that takes measurable time or could fail
## HTTP Middleware with `otelhttp`
Automatically creates spans for incoming and outgoing HTTP requests:
```go
import "go.opentelemetry.io/contrib/instrumentation/net/http/otelhttp"
// Incoming requests — wrap your handler
mux.Handle("/orders", otelhttp.NewHandler(orderHandler, "CreateOrder"))
// Outgoing requests — HTTP clients MUST use otelhttp for automatic span propagation
client := &http.Client{
Transport: otelhttp.NewTransport(http.DefaultTransport),
}
```
## Span Status and Recording Errors
```go
import (
"go.opentelemetry.io/otel/codes"
)
// On success — no need to set status (Unset is fine)
// On error — MUST call both RecordError() and SetStatus(Error)
if err != nil {
span.RecordError(err)
span.SetStatus(codes.Error, "operation failed")
return err
}
```
## Structured Errors with `samber/oops`
Standard Go errors lose critical debugging information: there's no stack trace, no structured context, and no way to attach request-scoped metadata. When an error surfaces in a trace, you see `"connection refused"` but not where it originated or which user/tenant was affected.
[`samber/oops`](https://github.com/samber/oops) is a drop-in error library that fills these gaps. Every `oops` error carries a stack trace, structured attributes, and integrates naturally with both OpenTelemetry spans and `slog`:
```go
import "github.com/samber/oops"
func (s *OrderService) Create(ctx context.Context, req CreateOrderRequest) (*Order, error) {
ctx, span := tracer.Start(ctx, "OrderService.Create")
defer span.End()
order, err := s.repo.Insert(ctx, req.ToOrder())
if err != nil {
// oops wraps the error with stack trace, structured context, and error code
return nil, oops.
In("order-service").
Code("order_insert_failed").
With("order_id", req.OrderID).
With("user_id", req.UserID).
Wrapf(err, "inserting order")
}
return order, nil
}
```
When this error is logged or recorded on a span, you get the full stack trace, the domain (`order-service`), an error code (`order_insert_failed`), and structured attributes (`order_id`, `user_id`) — all machine-parseable and searchable in your observability platform.
`oops` errors work with `span.RecordError()`, `errors.Is`/`errors.As`, and `slog` — see the `samber/cc-skills-golang@golang-error-handling` and `samber/cc-skills-golang@golang-samber-oops` skills for full usage patterns.
## Trace Sampling
In high-throughput services, tracing every request is expensive. Use sampling to control the volume:
```go
tp := sdktrace.NewTracerProvider(
// Sample 10% of traces in production
sdktrace.WithSampler(sdktrace.TraceIDRatioBased(0.1)),
sdktrace.WithBatcher(exporter),
sdktrace.WithResource(res),
)
```
For more nuanced control, use `sdktrace.ParentBased()` to respect the parent's sampling decision — this keeps traces complete across service boundaries.
## Cost of Tracing
Tracing can be one of the most expensive observability signals. Every span generates data that must be serialized, transmitted, stored, and indexed. In a microservices architecture, a single user request can produce dozens or hundreds of spans across services.
**Cost factors:**
- **Span volume** — a service handling 10k req/s with 5 spans per request generates 50k spans/s. At 100% sampling, this is enormous.
- **Span attributes** — each attribute adds to the payload size. Large attributes (request/response bodies) multiply cost.
- **Storage and indexing** — tracing backends (Jaeger, Tempo, Datadog) charge by volume. Unsampled traces can easily become the largest line item in your observability bill.
**Mitigation:**
- Use sampling (see above) — start with 10% (`TraceIDRatioBased(0.1)`) and adjust based on traffic volume and budget
- For high-throughput services, consider head-based sampling (decide at trace start) or tail-based sampling (decide after the trace completes, keeping only interesting traces like errors or slow requests)
- Avoid attaching large payloads as span attributes — log them instead and correlate via trace_id