# Integrate this codebase with Ritele

You are connecting this codebase to Ritele, a self-hosted tool that reads standard OpenTelemetry (traces and metrics), draws the real architecture of a system, and judges the health of every component on it. There is no Ritele client library. You make each service send standard OpenTelemetry to one address, and nothing else.

Follow this brief exactly. Do not invent attribute names, metric names or options that are not written here. If something here conflicts with what you see in the code, say so instead of guessing.

## 1. Ask before you change anything

You need three facts. Ask for any you cannot read from the repository.

1. The Ritele OTLP/HTTP address as the services reach it, for example `http://ritele:4318`. Traces and metrics go to the same host; the SDK appends `/v1/traces` and `/v1/metrics`. Port 4318 is OTLP/HTTP and 4317 is OTLP/gRPC.
2. The environment name for each deployment (`production`, `staging`). It becomes `deployment.environment.name` and keeps environments apart.
3. Whether the services already send to an OpenTelemetry Collector. If they do, add Ritele to that Collector (section 6) instead of changing each service's endpoint.

Then list every deployable service, its language, and how it is deployed (Docker, Kubernetes, VM). One deployable is one component.

## 2. Rules for every service

- `service.name` is the component's identity everywhere. One per deployable, stable across deploys and replicas, never a version or a host name. Changing it later reads as one component disappearing and another arriving.
- `deployment.environment.name` is required in practice. Without it everything lands in an environment called `default`.
- Send traces and metrics. Health is judged from metrics, so a service that sends only traces gets a map and no status.
- Never sample the metrics path, and never put an identifier, email, token or URL with ids in a metric label or span attribute.
- Name routes, not URLs: server spans carry `http.route` (`/v1/customers/:id`), never the raw path.
- Do not add any `ritele.*` attribute that this brief does not list.

### Complete traces

A trace is complete when every call a request makes is a span under it. The trace view, the Paths, the critical path and How it works depend on it.

- Trace every outbound call as a child of the request: HTTP, gRPC, database, cache, queue and cloud SDK (AWS, GCP) clients. For each client the service uses, check that its language's auto-instrumentation covers it. If one is not covered, wrap the call in a client span (`SpanKind.CLIENT`) if that is a small change; otherwise leave it and report it. Do not guess that a client is covered.
- Propagate W3C `traceparent` on every hop. Keep the default propagators; never switch them off.
- Leave the SDK sampler at its default, `parentbased_always_on`: Ritele's Collector samples after the fact. If a service already samples, keep it parent-based (`parentbased_traceidratio`) with the same `OTEL_TRACES_SAMPLER_ARG` on every service, and do not add sampling where there is none.
- Carry context through every async hand-off, not only Kafka: SQS, RabbitMQ, Redis streams, background jobs. On produce, inject `traceparent` into the message headers (`propagation.inject`). On consume, extract it and parent the consumer span to it, or add a span link to it.
- Add named spans for the important internal steps of a request, such as `validate-transfer` or `score-risk`, with `tracer.startActiveSpan`. A handful per operation, not every function. End every span, in `finally`.
- Keep span names short and stable. No ids, amounts or user values in a name: put them in attributes, and keep attributes free of personal data.
- On database client spans, `db.operation.name` and `db.collection.name` (the table) should be set. Most instrumentations set them; add them only where one does not. Write queries with placeholders (`$1`, `?`), never values built into the SQL, so customer values stay out of query text.

## 3. Node.js services: use @omob/otel-kit

`@omob/otel-kit` is an MIT-licensed npm package that wraps the OpenTelemetry SDK. Use version 0.12 or later: before 0.12, a service started with `--require` ran two SDKs and reported its CPU and GC twice. Before 1.0 a caret range does not cross a minor, so write `"@omob/otel-kit": "^0.12.0"`, never `^0.10.0` or `^0.3.0`.

```bash
npm install @omob/otel-kit @opentelemetry/api
```

Create `src/instrumentation.ts`:

```ts
import { ExporterType, Telemetry } from "@omob/otel-kit";

const endpoint = process.env.OTEL_EXPORTER_OTLP_ENDPOINT;

Telemetry.start({
  serviceName: "wallet-service",
  serviceVersion: process.env.APP_VERSION,
  environment: "production",
  traces: {
    exporter: ExporterType.OTLP,
    otlp: { url: `${endpoint}/v1/traces` },
  },
  metrics: { exporter: ExporterType.OTLP, cpuUsage: true },
  instrumentation: { ignoreIncomingPaths: ["/health"] },
  handleShutdownSignals: true,
});
```

Replace `wallet-service` and `production` with this service's values. Then:

- Load it before everything else. CommonJS: `node --require ./dist/instrumentation.js dist/server.js`. ESM (`"type": "module"`): `node --import ./dist/instrumentation.js dist/server.js`, on the command line and not as an import inside the entry file, or Fastify, ioredis and kafkajs are patched too late.
- Do not register `@opentelemetry/auto-instrumentations-node` as well: no `--require` or `--import` of its `register` module, in the start command or `NODE_OPTIONS`, and no `registerInstrumentations(getNodeAutoInstrumentations())`. otel-kit registers them already, and a second set doubles every span and metric. Remove any you find. To configure one instrumentation, use `instrumentation.config`.
- Set `OTEL_EXPORTER_OTLP_ENDPOINT` on the service. With OTLP traces the kit sends metrics to the same place every 30 seconds with no further code.
- Do not set `OTEL_METRICS_EXPORTER=none`. It overrides the `metrics` block and silently switches Health off for that service. The Jaeger quick start suggests it; remove it.
- On Fastify, run `npm i @fastify/otel` and add `enable: [InstrumentationName.FASTIFY]` under `instrumentation`. Without it every request is one bare span with no `http.route`.
- Set `serviceVersion` when the service starts from a monorepo root, which would report the root package's version.

### Connection pools

A trace cannot show that a service holds seven of its ten connections. Register each pool after `Telemetry.start()`, never before, or the readings are discarded and nothing errors. Name it `host:port/database`; pools sharing a name add into one series.

```ts
import { observeConnectionPool } from "@omob/otel-kit";

observeConnectionPool({
  name: "10.0.1.5:5432/billing",
  system: "postgresql",
  read: () => ({
    max: pool.options.max,
    used: pool.totalCount - pool.idleCount,
    idle: pool.idleCount,
    pending: pool.waitingCount,
  }),
});
```

### Replicas and CPU

- `service.instance.id` is a random id per process by default. On Kubernetes set it to the pod name so a restart is not counted as a new instance.
- `CPU_LIMIT_MILLICORES` gives the kit the CPU limit of one replica. Set it only when the container has a CPU limit: with none, the Kubernetes downward API reports the node's capacity.

On Kubernetes, pass the pod's identity in through the downward API. Declare each variable before the one that uses it:

```yaml
env:
  - name: CPU_LIMIT_MILLICORES
    valueFrom: { resourceFieldRef: { resource: limits.cpu, divisor: 1m } }
  - name: POD_NAME
    valueFrom: { fieldRef: { fieldPath: metadata.name } }
  - name: POD_UID
    valueFrom: { fieldRef: { fieldPath: metadata.uid } }
  - name: POD_NAMESPACE
    valueFrom: { fieldRef: { fieldPath: metadata.namespace } }
  - name: NODE_NAME
    valueFrom: { fieldRef: { fieldPath: spec.nodeName } }
  - name: OTEL_RESOURCE_ATTRIBUTES
    value: service.instance.id=$(POD_NAME),k8s.pod.name=$(POD_NAME),k8s.pod.uid=$(POD_UID),k8s.namespace.name=$(POD_NAMESPACE),k8s.node.name=$(NODE_NAME),k8s.container.name=wallet-service
```

`k8s.container.name` is the container's `name` in the pod spec, which the downward API cannot supply; write it in.

### Messaging

kafkajs is traced by default. A consumer that reads with `eachBatch` starts a trace of its own and links to the producer's; its topic's volume comes from the `messaging.process.duration` it sends, so keep its metrics exported. For `@google-cloud/pubsub`, create the client with `enableOpenTelemetryTracing: true`, or it makes no spans.

## 4. Services in any other language

Use that language's standard OpenTelemetry SDK or agent (the Java agent, the .NET and Python auto-instrumentation, the Go SDK) and configure it with environment variables only:

```
OTEL_SERVICE_NAME=wallet-service
OTEL_RESOURCE_ATTRIBUTES=deployment.environment.name=production,service.version=2.4.1
OTEL_EXPORTER_OTLP_ENDPOINT=http://ritele:4318
OTEL_EXPORTER_OTLP_PROTOCOL=http/protobuf
OTEL_TRACES_EXPORTER=otlp
OTEL_METRICS_EXPORTER=otlp
OTEL_METRIC_EXPORT_INTERVAL=30000
OTEL_SEMCONV_STABILITY_OPT_IN=http,database,messaging
```

- Keep `OTEL_SEMCONV_STABILITY_OPT_IN` as written. Older HTTP instrumentation reports `http.server.request.duration` only with `http`, and the Java agent reports `db.client.operation.duration` and `messaging.process.duration` only with `database` and `messaging`. Without them the service's requests and its calls to databases are missing.
- Health reads these metrics: `http.server.request.duration`, `http.client.request.duration`, `rpc.server.call.duration`, `rpc.client.call.duration`, `db.client.operation.duration`, `messaging.process.duration`, `messaging.client.consumed.messages`, `db.client.connection.count`, `db.client.connection.max`, `db.client.connection.pending_requests`, `db.client.connection.wait_time`, `process.cpu.time`, and the `jvm.*` family on Java. If the library already emits them, nothing more is needed.
- For a pool the library does not instrument, emit the four `db.client.connection.*` metrics yourself. Send `db.client.connection.wait_time` under one name only, never both the old millisecond spelling and the current seconds one.
- To state the CPU limit of one replica, set the resource attribute `ritele.cpu.limit` to cores, `0.5` for a limit of `500m`.

## 5. Optional resource attributes

Add these only when you know the value. They shape the map and are never required.

- `ritele.domain`, `ritele.layer`, `ritele.zone`: the organisation's own grouping, layering and region.
- `ritele.component.type`: only where Ritele guesses the type wrong.
- `ritele.intended_dependencies`: the components this service means to call, for drift. Name them the way the map does: `postgresql:ledger`, `kafka:transfers`, `paystack`.

## 6. A Collector you already run

Add an exporter and one traces and one metrics pipeline beside the ones you have. Do not append a signal path to the endpoint: `otlphttp` adds `/v1/traces` and `/v1/metrics` itself, and an endpoint that already ends in one posts to `/v1/traces/v1/traces`.

```yaml
exporters:
  otlphttp/ritele:
    endpoint: ${env:RITELE_OTLP_ENDPOINT}
service:
  pipelines:
    traces/ritele:
      receivers: [otlp]
      processors: [batch]
      exporters: [otlphttp/ritele]
    metrics/ritele:
      receivers: [otlp]
      processors: [batch]
      exporters: [otlphttp/ritele]
```

Do not sample before this exporter.

## 7. Prove it, do not assume it

Nothing errors when a signal goes to the wrong place, so ask for it. After the services restart and have served traffic for a minute or two:

- In Ritele, open Settings › Data sources. Every service you changed should be listed with points per minute above zero.
- Or query the API on port 8080:

```graphql
{ dataSources(env: "production") { name kind pointsPerMin lastSeen silent } }
{ coverage(env: "production") { name coverage { traces { state } appMetrics { state } runtime { state } pool { state } } } }
```

`traces` and `appMetrics` should read `RECEIVED` for every service, `runtime` for every Node or Java service, and `pool` for every service you registered a pool in. A cell that reads `MISSING` means that signal never arrived; the section above for that signal is what to recheck. An empty `dataSources` list means nothing reached Ritele: check the address, the port, and that metrics are not switched off.

## 8. Report back

Finish by listing the files you changed, the environment variables each service now needs and their values, every service you could not instrument and why, every outbound client you could not trace and where, every internal span you added and where, and anything in this brief you were unsure of.
