Documentation
Getting the most out of Ritele
Ritele draws a live architecture map from the telemetry your services already emit, then answers questions about it — what talks to what, what is closest to giving out, what a change would do. How much it can answer depends entirely on what reaches it. This page is that ladder, the two ways to climb it, and what Health, the agent and notifications ask of you.
Five minutes in
One container holds the Collector, the ingest, the store and the UI. Nothing to configure before you start.
docker run -p 4317:4317 -p 4318:4318 -p 8080:8080 \ -v ./data:/data ritele/ritele
Point any OpenTelemetry SDK at http://localhost:4318, open http://localhost:8080, and the map draws itself as traffic arrives. There is no schema to define and no diagram to author — the first version of your architecture is whatever your traces say it is.
That is the floor, and it is genuinely useful on its own. Everything below is about the ceiling.
Connecting with an AI coding assistant? Integrating with AI, below, has a brief written for one.
Running it
There are two ways to run it, and both are one command. Start with one container; move to Postgres and ClickHouse when you outgrow it, and bring everything with you.
| One container | Postgres + ClickHouse | |
|---|---|---|
| Start it | docker run — the command above | docker compose up -d, with the compose file below |
| Where data lives | inside the container's /data volume: SQLite for the map, DuckDB for metrics and traces | your Postgres 16 (the map) and ClickHouse 24.8 (metrics and traces) — the compose file starts both, or point it at your own |
| How big | up to 500 services; give /data 8 GB | up to 5,000 services, as far as your databases hold |
| What you run | one container | Ritele plus two databases to back up and upgrade |
| Choose it when | you want the simplest thing that works | you outgrow 500 services, send heavy traffic, or want the databases managed on their own |
Both use the same image, pulled from Docker Hub — nothing is built on your machine. Your licence can set a lower service limit than either figure above.
Postgres and ClickHouse
Save this as docker-compose.yml and start it. Change the passwords before anything else reaches it.
services:
ritele:
image: ${RITELE_IMAGE:-ritele/ritele:latest}
ports: ["4317:4317", "4318:4318", "8080:8080"]
environment:
RITELE_STORAGE: external
DATABASE_URL: postgres://ritele:ritele@postgres:5432/ritele
CLICKHOUSE_URL: http://ritele:ritele@clickhouse:8123
depends_on:
postgres: { condition: service_healthy }
clickhouse: { condition: service_healthy }
restart: unless-stopped
postgres:
image: postgres:16
environment: { POSTGRES_USER: ritele, POSTGRES_PASSWORD: ritele, POSTGRES_DB: ritele }
volumes: [pg:/var/lib/postgresql/data]
healthcheck:
test: ["CMD-SHELL", "pg_isready -U ritele -d ritele"]
interval: 5s
timeout: 5s
retries: 12
clickhouse:
image: clickhouse/clickhouse-server:24.8
environment: { CLICKHOUSE_USER: ritele, CLICKHOUSE_PASSWORD: ritele, CLICKHOUSE_DEFAULT_ACCESS_MANAGEMENT: 1 }
volumes: [ch:/var/lib/clickhouse]
healthcheck:
test: ["CMD-SHELL", "wget -qO- http://localhost:8123/ping || exit 1"]
interval: 5s
timeout: 5s
retries: 12
volumes: { pg: {}, ch: {} }
docker compose up -d
How they perform
Both were measured before the Health release, at the one container's limits then: 500 services, a week of per-minute metrics, a month of hourly ones and a week of traces. The one container answers every screen well inside its budget — a 30-day map in about 0.2 s, a week's trace list in about 0.1 s — in 3.6 GB. It writes about 450 metric rows a second, roughly twice what an estate that size sends, and that single writer is its real ceiling. Postgres and ClickHouse hold the same estate in half the disk and take metric writes tens of times faster, from many processes at once; updates to the map itself are slower there, one database round trip each, and still keep up with an estate that size. Below the one container's limits it is as quick and has less to run; past them, or under heavy traffic, move to Postgres and ClickHouse. With Health, the one container budgets 8 GB of /data at its default retention — see what's kept, below, for the windows and the series cap, both adjustable from Settings.
What is kept, and for how long
| Data | Kept for | Set by |
|---|---|---|
| The map: components and the calls between them | 30 days unless you change it | you, per environment, in Settings |
| Per-minute metrics | 48 hours by default on one container, 7 days by default on Postgres and ClickHouse | Settings › Storage › Retention, 24 to 168 hours |
| Hourly metrics | 30 days by default | Settings › Storage › Retention, 7 to 30 days |
| Kept traces | 7 days by default | Settings › Storage › Retention, 1 to 7 days |
| Critical paths | 13 months | fixed |
Settings › Storage › Retention sets those first three windows and the series cap for the whole install. A value saved there wins over the container's own RITELE_HEALTH_MINUTE_HOURS and RITELE_MAX_SERIES_PER_ENV, which still apply wherever nothing has been saved. Shortening a window takes effect a day after you save it, the same delay as a shortened environment window below; the series cap, raised or lowered, applies at once. A licence with its own metrics window caps how far either metric window can be raised.
Settings → Storage shows every window in one place. You choose how often each environment is pruned to its window — hourly, daily or weekly — and Clear old data now runs the whole cleanup at once. It only removes what is already past its window. The fixed windows are kept to within the hour whatever you choose. A snapshot of the map is taken before a prune, so what it removes stays readable on the timeline.
Shortening an environment's window waits a day before it deletes anything, and Settings says when it takes effect. Lengthening it, or keeping everything, applies at once. Clearing an entire environment is switched off unless the instance sets RITELE_ALLOW_RESET=true.
Disk comes back
Deleting an old row frees space inside the metrics file, but the file itself does not shrink on its own — the freed blocks just sit ready to be reused. Once a day, at 03:00 UTC, Ritele checks whether enough is reclaimable to be worth it and whether the file is small enough to finish in time, and compacts the file down to size if so. Settings › Storage also has Reclaim now, for an admin who wants it done sooner or is willing to accept a longer pause than the daily check allows. Either way, every metric read and write pauses for the length of the compaction — typically a few minutes, based on how long the last one took — and resumes once it finishes; the original file is only replaced once the compacted copy has been checked against it.
On Postgres and ClickHouse, each hourly cleanup drops whole days once they fall entirely outside their window, and rewrites the one day straddling the edge so its expired rows are actually removed rather than just hidden. Both tiers give the disk back, not only the row count.
Settings › Storage's On disk row shows the metrics file's real size, how much of that is reclaimable, and the free space left on the volume — so Reclaim now is a decision made with numbers in front of you, not a guess.
Moving from one container to Postgres and ClickHouse
Stop Ritele first. In the folder holding docker-compose.yml and your ./data, start only the two databases, copy into them over the network compose created (named after the folder — ritele_default in a folder called ritele), then start Ritele on them. Starting Ritele before the copy would give it databases that already hold data, which the copy refuses. It copies everything the one container holds — the map and its history, the intended model, drift findings, scenarios, settings, metrics and traces — then compares row counts table by table and exits with an error on any difference.
docker compose up -d postgres clickhouse docker run --rm --network ritele_default -v ./data:/data \ -e DATABASE_URL=postgres://ritele:ritele@postgres:5432/ritele \ -e CLICKHOUSE_URL=http://ritele:ritele@clickhouse:8123 \ ritele/ritele node packages/storage/dist/copy-to-external.cli.js docker compose up -d
The capability ladder
Each rung is a distinct thing to send, and each one unlocks something specific. Nothing here is all-or-nothing: stop at any rung and everything below it keeps working.
01
The map
Components, the calls between them, and which are instrumented. Every datastore, cache, broker and third-party host your services reach appears as a node, whether or not it can report for itself.
Needs spans with service.name on the resource. That is all.
02
Environments kept apart
Staging traffic stops contaminating the production map, and each environment carries its own history, its own intended model and its own scenarios.
Needs deployment.environment.name on the resource.
03
Entry points and critical paths
Named routes rather than a blur of spans, and for each one the hop-by-hop path a request actually takes — with the self time at every hop, so you can see which one to improve rather than guess.
Needs http.route on server spans (not the raw URL), the messaging conventions on consumers, and a span for every outbound call, so each hop is on the path and none is charged to its caller.
04
Rate, errors and duration you can trust
Volume that counts every request rather than a sample of it. The map's traffic figures, the baseline comparisons and every capacity answer are built on this.
Needs a Collector carrying the spanmetrics and servicegraph connectors.
05
How full things are
Connection pools, queue depth, and the number nobody else can compute: how much of a shared database's connection budget four separate services are holding between them.
Needs the db.client.connection.* metrics from your services.
06
What a change would do
Ask what happens at three times the traffic, or if the database gets 50 ms slower, or if a component goes away — and get back what gives out first and at what multiple of today.
Needs rungs 3 to 5, plus a declared concurrency per component.
07
Your own vocabulary on the map
Domains, layers and zones, so the map is laid out the way your organisation actually thinks — and a layer violation becomes something the tool can find rather than something a reviewer notices.
Needs ritele.domain, ritele.layer, ritele.zone on the resource.
08
Drift from what you intended
Record the architecture you meant to have, and every unexpected edge, missing dependency, new component and layer violation is reported against it from then on.
Needs an intended model — accept the current map as intended, or import one from CSV, Mermaid or Structurizr.
09
Health: a status for every component, and why
Healthy, degraded, unhealthy or silent for each component, with the numbers behind it, and an issue that opens only after a problem has lasted. Where something is wrong, it names what is likely to have caused it.
Needs request and runtime metrics from your services, and pool metrics where they hold a pool. See What Health needs.
Which path
Both reach the top of the ladder. They differ in how much you assemble yourself.
Plain OpenTelemetry
any language, any SDK
- No new dependency, and nothing Ritele-specific in your services
- Works from Go, Python, Java, .NET, Rust — anything that speaks OTLP
- You wire the resource attributes, the pool metrics and the export paths yourself
- The right choice for a polyglot estate, or where adding a library needs a review
@omob/otel-kit
a Node package that wires OpenTelemetry for you
- One
Telemetry.start()call sets up traces, metrics and resource attributes observeConnectionPool()emits the four pool metrics from a single reader- Sensible instrumentation defaults, and the export paths are correct by construction
- The faster route for a Node estate, and the one that avoids the traps below
Plain OpenTelemetry
Three things on the resource, two conventions on spans, and the export pointed at the Collector.
Resource attributes
Set these once, wherever your SDK is initialised. The first two are the whole of rungs 1 and 2.
service.name billing-management # the component’s identity deployment.environment.name production # keeps environments apart service.version 2.4.1 # optional, shown on the component # Optional, and what makes the map yours rather than generic ritele.domain payments ritele.layer edge ritele.zone eu-west-2 ritele.component.type gateway # when the heuristic guesses wrong
Span conventions that matter
Ritele reads current OpenTelemetry semantic conventions and the previous spelling of each, so an SDK a version or two behind still works.
| Attribute | On | What it buys |
|---|---|---|
http.route | server spans | A named entry point. Without it every distinct URL looks like a different route and the list is unusable. |
http.request.method | server & client spans | Operation grouping on an edge. |
db.system.name, db.namespace | client spans | The datastore appears as its own node, correctly typed, named for the database rather than the host. |
messaging.system, messaging.destination.name | producer & consumer | The topic or queue becomes a node, and a producer is linked to its consumers through it. |
messaging.operation.type | consumer spans | Distinguishes consuming from producing, which decides where a worker's spans are charged. |
server.address, peer.service | client spans | An uninstrumented third party still appears, named for what it is. |
db.operation.name, db.collection.name | database client spans | The operation and table on the span in the trace view, so a slow query is named without its text. |
| Span name | every span | A critical-path hop is the service and span name together. A short, stable name (POST /transfers, score-risk) is one hop; one with an id in it is a new hop every time. |
Complete traces
The trace view, the Paths, the critical path and How it works are only as complete as the spans you send. An error’s cause is found by following the trace down to the span that failed.
| Do | Why |
|---|---|
| Trace every outbound call as a child of the request: HTTP, gRPC, database, cache, queue and cloud SDK (AWS, GCP) clients. | A call with no span is time charged to its caller, and the dependency is missing from the trace. Check your language's instrumentation list for each client you use; one with no instrumentation needs a manual client span. |
Propagate W3C traceparent on every hop, and keep the SDK sampler parent-based. | A service that starts a new trace instead of continuing one splits the request in two. Leave OTEL_TRACES_SAMPLER at its default, parentbased_always_on: the bundled Collector samples after the fact. If you do sample in the SDK, use parentbased_traceidratio with the same OTEL_TRACES_SAMPLER_ARG on every service, so none drops what another kept. |
Carry context through every async hand-off: SQS, RabbitMQ, Redis streams, background jobs, not only Kafka. Inject traceparent into the message headers on produce; extract it or add a span link on consume. | Without it the worker's spans form a separate trace, and the request that queued the work shows no cost for it. |
Name the important internal steps of a request (validate-transfer, score-risk) with tracer.startActiveSpan. A handful per operation, not every function. | Each becomes a hop with its own self time, so the path shows which step is slow rather than one long span. |
| Keep span names short and stable. Ids, amounts and user values go in attributes, which stay free of personal data. | Names are what hops are grouped by. |
On database spans, write queries with placeholders ($1, ?), never values written into the SQL. | Customer values stay out of query text. Ritele keeps the text with literals replaced by ? either way, unless RITELE_DB_QUERY_TEXT says otherwise. |
Exporting
Traces and metrics are separate pipelines and both are needed. Point them at port 4318 of this install, which is its bundled Collector, or at a Collector you already run that forwards to it. Never at ingest on 8081: the Collector is what turns spans into the unsampled volume of rung 4.
OTEL_EXPORTER_OTLP_ENDPOINT=http://ritele:4318 OTEL_EXPORTER_OTLP_PROTOCOL=http/protobuf
For Health
Health asks nothing Ritele-specific of a plain SDK. Export metrics as well as traces every 30 seconds, and keep the stable metric names. The Java agent, the .NET and Python auto-instrumentation and the Go SDK are configured the same way, through these variables.
OTEL_SERVICE_NAME=billing-management OTEL_RESOURCE_ATTRIBUTES=deployment.environment.name=production,service.version=2.4.1 OTEL_EXPORTER_OTLP_ENDPOINT=http://ritele:4318 OTEL_EXPORTER_OTLP_PROTOCOL=http/protobuf OTEL_TRACES_EXPORTER=otlp OTEL_METRICS_EXPORTER=otlp OTEL_LOGS_EXPORTER=otlp OTEL_METRIC_EXPORT_INTERVAL=30000 OTEL_SEMCONV_STABILITY_OPT_IN=http,database,messaging
OTEL_SEMCONV_STABILITY_OPT_IN is what makes the stable metrics appear. Older HTTP instrumentation reports http.server.request.duration only with http, and the Java agent reports db.client.operation.duration and messaging.process.duration only with database and messaging. Without them a service's requests, and its calls into a database, are missing. otel-kit needs none of this.
OTEL_LOGS_EXPORTER=otlp sends the service's logs the same way, for any plain OpenTelemetry SDK. See Logs, below, for what Ritele keeps from them.
With @omob/otel-kit
@omob/otel-kit is a small MIT-licensed npm package that wraps the OpenTelemetry SDK for Node. It is not required and it is not a client for us — everything below can be wired by hand from any language. What it saves you is the assembly: one Telemetry.start() call in place of exporters, resource attributes, instrumentation defaults and the two separate export paths, each of which is easy to get subtly wrong.
npm install @omob/otel-kit @opentelemetry/api
Use ^0.12.0 or later. Before 0.12, a service loaded with node --require ran two copies of the SDK, so its CPU read double and its GC time could read above 100%. Before 1.0 a caret range never crosses a minor, so ^0.10.0 stays on 0.10.x, and ^0.3.0 predates nearly everything Health reads. Bump the minor in your package.json on purpose.
| Version | Added |
|---|---|
| 0.2 | observeConnectionPool(), and instrumentation.only. |
| 0.4 | The architecture attributes under ritele.*. A service on 0.3 has them ignored. |
| 0.5 | metrics.cpuUsage and architecture.cpuLimit, and a service.instance.id on every resource. |
| 0.6 | Metrics sent by default whenever traces go over OTLP, every 30 seconds; runtime metrics kept under instrumentation.only; ritele.trace.sample_probability; the stable HTTP metrics, with no variable needed. |
| 0.7 | container.id, and OTEL_METRICS_EXPORTER=none overriding a metrics block. |
| 0.9 | ritele.v8js.memory.heap.limit, the ceiling heap use is judged against. |
| 0.10 | CPU_LIMIT_MILLICORES and the standard sampler variables read; ids in url.path masked; metrics never sent to a stray localhost. |
| 0.11 | Logs follow the traces URL, so a one-line logs block sends them. |
| 0.12 | One SDK per process. Before it, node --require started a second one, which doubled CPU and pushed GC above 100%. A warning when the auto-instrumentations are registered a second time. |
Put this in its own module and load it before anything else. CommonJS: node --require ./dist/instrumentation.js. ESM: node --import ./dist/instrumentation.js, on the command line rather than as an import in your entry file. Instrumentation has to be in place before the libraries it patches are loaded. Don't also register @opentelemetry/auto-instrumentations-node, through its register module or getNodeAutoInstrumentations(): the kit already does, and a second set records every span and metric twice. Change one instrumentation's behaviour through instrumentation.config instead.
import { ExporterType, InstrumentationName, Telemetry } from "@omob/otel-kit";
Telemetry.start({
serviceName: process.env.APP_NAME ?? "billing-management",
serviceVersion: "2.4.1",
environment: "production",
resourceAttributes: { "ritele.domain": "payments", "ritele.layer": "edge" },
traces: {
exporter: ExporterType.OTLP,
otlp: { url: `${process.env.OTEL_EXPORTER_OTLP_ENDPOINT}/v1/traces` },
sampleRatio: 1,
},
// Metrics follow the traces URL to /v1/metrics every 30 seconds with no block at all.
// This one only adds the CPU readings.
metrics: { exporter: ExporterType.OTLP, cpuUsage: true },
instrumentation: {
only: [InstrumentationName.HTTP, InstrumentationName.UNDICI, InstrumentationName.PG,
InstrumentationName.IOREDIS, InstrumentationName.KAFKAJS],
enable: [InstrumentationName.FASTIFY],
ignoreIncomingPaths: ["/health"],
},
handleShutdownSignals: true,
});ignoreIncomingPaths is worth setting. Health checks are usually the highest-volume route in an estate and they crowd out everything real on the entry-point list. On Fastify, npm i @fastify/otel as well, or every request is one bare span with no http.route.
otel-kit sends ids in url.path, url.query and url.full as * and keeps http.route, so requests still group by route. It never sends the command line, the path of the node binary or the user name, and on a personal machine it replaces host.name and host.id with stable pseudonyms. This install's Collector deletes url.query and url.full on arrival either way.
Logs
otel-kit 0.11 or later sends logs with one more block in Telemetry.start() — logs follow the traces URL (/v1/traces becomes /v1/logs). It forwards what the service writes through pino, winston or bunyan, with that write's trace id attached; a bare console.log is not captured. Setting OTEL_EXPORTER_OTLP_LOGS_ENDPOINT turns logs on with no code change, and OTEL_LOGS_EXPORTER=none turns them off.
logs: { exporter: ExporterType.OTLP },What Health needs
Request and runtime metrics arrive on their own. Pools, CPU and where an instance runs each need a line from you. Settings › Data sources and the Coverage table show which are still missing.
| Health reads | With otel-kit | If it is missing |
|---|---|---|
| Request rate, errors and latency | Automatic. Metrics follow the traces URL every 30 seconds, and http.server.request.duration is reported in seconds. | The App metrics column reads missing, and the service's span counts, scaled up by its sample probability, stand in. |
| Event loop, heap and garbage collection | Automatic, the heap ceiling included, and kept under instrumentation.only. runtimeMetrics: false turns it off. | The Runtime column reads missing, and heap use has no ceiling to be judged against. |
| Pool use and waiting callers | observeConnectionPool() for each pool, after Telemetry.start(), named host:port/database. | The Pool column reads missing. A saturated pool looks like a slow database. |
| CPU against its limit | metrics: { cpuUsage: true } and CPU_LIMIT_MILLICORES, set only where the container has a CPU limit. | CPU is not judged. |
| Which replica is which | A random service.instance.id per process. On Kubernetes set it to the pod name through OTEL_RESOURCE_ATTRIBUTES. | Every restart counts as a new instance. |
| Where an instance runs | k8s.pod.uid, k8s.pod.name, k8s.namespace.name, k8s.node.name and k8s.container.name through OTEL_RESOURCE_ATTRIBUTES. container.id is detected under Docker and host.id always. | An instance cannot be matched to its pod, container or host, so the service's Infra column stays empty. |
On Kubernetes
Pass the pod’s identity and its CPU limit in through the downward API rather than in code.
env:
- name: CPU_LIMIT_MILLICORES
valueFrom: { resourceFieldRef: { resource: limits.cpu, divisor: 1m } }
- name: POD_NAME
valueFrom: { fieldRef: { fieldPath: metadata.name } }
- name: POD_UID
valueFrom: { fieldRef: { fieldPath: metadata.uid } }
- name: POD_NAMESPACE
valueFrom: { fieldRef: { fieldPath: metadata.namespace } }
- name: NODE_NAME
valueFrom: { fieldRef: { fieldPath: spec.nodeName } }
- name: OTEL_RESOURCE_ATTRIBUTES
value: service.instance.id=$(POD_NAME),k8s.pod.name=$(POD_NAME),k8s.pod.uid=$(POD_UID),k8s.namespace.name=$(POD_NAMESPACE),k8s.node.name=$(NODE_NAME),k8s.container.name=billing-managementDeclare each variable before the one that uses it. divisor: 1m gives millicores, so a limit of 500m arrives as 500. k8s.container.name is the container's name in the pod spec, which the downward API cannot supply, so write it in. A container with no CPU limit should not set CPU_LIMIT_MILLICORES: the downward API then reports the node's capacity.
Messaging
kafkajs is traced by default. A consumer that reads with eachBatch starts a trace of its own, linked to the producer's, and its topic's volume is read from the messaging.process.duration it sends, so keep its metrics exported. For @google-cloud/pubsub, create the client with enableOpenTelemetryTracing: true, or it makes no spans.
The Collector
Rung 4 lives here. The quickstart image ships a Collector already configured; this is what it carries, for anyone running their own.
Two trace pipelines do different jobs from the same spans, and the split is the important part:
| Pipeline | Sampling | Produces |
|---|---|---|
traces/graph | none — every span | The servicegraph and spanmetrics connectors. This is where volume, rate and error counts come from, and it must not be sampled. |
traces/exemplar | tail sampled | The individual traces behind a critical path. Errors and doc-traces kept at 100%, everything else at 1%. |
The minimum configuration
connectors:
service_graph:
# Lets a datastore or third party become a node, named for what it is
virtual_node_peer_attributes: [peer.service, db.namespace, db.name,
server.address, rpc.service, db.system.name]
span_metrics:
metrics_flush_interval: 15s
service:
pipelines:
traces/graph: # unsampled, on purpose
receivers: [otlp]
exporters: [service_graph, span_metrics]
traces/exemplar:
receivers: [otlp]
processors: [tail_sampling, batch]
exporters: [otlphttp/ritele]
metrics/graph:
receivers: [service_graph, span_metrics]
exporters: [otlphttp/ritele]
metrics/scraped: # your services’ own metrics, incl. pools
receivers: [otlp]
exporters: [otlphttp/ritele]Connection pools
A trace says how long a query took. It cannot say that this service holds seven of its ten connections, or that requests started queueing before anything got slow. A pool keeps those numbers in memory and tells nobody.
Four metrics, all straight from the OpenTelemetry database-client conventions. Anything already emitting them is already done.
| Metric | Carries |
|---|---|
db.client.connection.count | db.client.connection.state = used or idle |
db.client.connection.max | The pool's ceiling |
db.client.connection.pending_requests | Callers waiting for a connection |
db.client.connection.wait_time | Seconds spent waiting |
pending_requests is the one to put in front of someone. It goes non-zero before an outage, while the pool is exhausted and work is queueing. Latency only shows the symptom afterwards.
With otel-kit
import { observeConnectionPool } from "@omob/otel-kit";
observeConnectionPool({
name: "10.0.1.5:5432/billing", // host:port/database
system: "postgresql",
read: () => ({
max: db.options.max,
used: db.totalCount - db.idleCount,
idle: db.idleCount,
pending: db.waitingCount,
}),
});The metrics pipeline
Health is read from metrics your services already know how to send. They travel beside traces, on their own path, and are kept only where something reads them.
| Step | What happens |
|---|---|
| Your SDK | Exports metrics to /v1/metrics on port 4318 (or OTLP/gRPC on 4317) — the same host as traces, a different path. |
metrics/scraped | The bundled Collector's pipeline for metrics your services send. It deletes enduser.*, user.*, headers, query text and process command lines before anything is stored. |
| Ingest | Keeps the metrics it knows, those below among them, as one-minute series per component and instance. Anything else is listed under Settings › Metrics as not stored, where an admin can choose to store it. |
What health reads
| Metric | Judges |
|---|---|
http.server.request.duration, rpc.server.call.duration | Rate, errors and latency of what a service serves. Older HTTP instrumentation only reports this with OTEL_SEMCONV_STABILITY_OPT_IN=http, and otel-kit 0.6 or later needs no variable. Without it, the service's own spans stand in. |
http.client.request.duration, db.client.operation.duration, rpc.client.call.duration | A database, cache or external API, judged from the calls its callers make to it. |
messaging.client.consumed.messages, messaging.process.duration | A consumer's throughput and processing time. The Java agent sends the second only with OTEL_SEMCONV_STABILITY_OPT_IN=messaging. |
db.client.connection.* | How full each pool is, and how many callers are waiting. |
nodejs.eventloop.*, v8js.*, ritele.v8js.memory.heap.limit, jvm.* | The runtime: event-loop delay and utilisation, heap and garbage collection. The heap ceiling is ritele.v8js.memory.heap.limit, which otel-kit 0.9 or later sends; the old v8js.memory.heap.limit meant each heap space's size and is not read. |
Series are capped per environment: 10,100 by default on one container, raisable to 100,000 from Settings › Storage › Retention; 250,000 on Postgres and ClickHouse. A metric is also capped at 5,000 series of its own, with RITELE_MAX_SERIES_PER_METRIC. Past either cap, new series fold into one overflow series per metric rather than being dropped silently, and Settings › Metrics shows how close each environment is.
The series cap, explained
A series is one metric, for one component or instance, with one set of label values — a request-duration metric for billing-management is a different series from the same metric for billing-worker, and an instance that restarts under a new identity starts series of its own. Settings › Metrics › Series by metric lists every metric an environment has stored, how many series it holds, and whether Health reads it — that badge marks the metrics a status actually depends on, as distinct from everything merely kept. From there, Stop storing turns a metric off without touching what is already kept, and Delete its data removes what is stored for it as well.
A series that stops reporting still holds its place in the cap until it is retired. An hourly job retires any series silent for RITELE_SERIES_SILENCE_DAYS — 30 days by default, never more — which frees the place but not the disk: the data ages out on its own window regardless. Settings › Storage also offers Retire them now, for series you would rather free up sooner than wait for the daily job.
From a Collector you already run
Already running a Collector? Add an exporter to this install and one traces and one metrics pipeline beside the ones you have. Settings › Data sources gives the same snippet, with each environment variable explained.
exporters:
otlphttp/ritele:
endpoint: ${env:RITELE_OTLP_ENDPOINT} # e.g. http://ritele:4318, no signal path
service:
pipelines:
traces/ritele:
receivers: [otlp]
processors: [batch]
exporters: [otlphttp/ritele]
metrics/ritele:
receivers: [otlp]
processors: [batch]
exporters: [otlphttp/ritele]Datastore metrics
A trace sees a query from the caller's side. How many connections the server allows, how many are in use and how much memory a cache has left live inside the datastore, and something has to connect to it and ask. That is a Collector: the one bundled in this install, or one you already run.
The bundled Collector connects to a datastore only when an operator adds a receiver file for it to /etc/ritele/collector.d/ and restarts the container; templates for PostgreSQL, MySQL, Redis, Kafka, MongoDB and HTTP checks ship beside it in /etc/ritele/datastores/. The credential is an environment variable on the Collector's container, named in the file as ${env:…}. Ritele never writes it to its storage and never returns it through its API.
receivers:
postgresql:
endpoint: db.internal:5432
username: ritele_monitor
password: ${env:PG_MONITOR_PASSWORD}
databases: [ledger]
collection_interval: 60s
redis:
endpoint: cache.internal:6379
password: ${env:REDIS_MONITOR_PASSWORD}
collection_interval: 60s
metrics:
redis.maxmemory: { enabled: true } # the ceiling used memory is a fraction of
processors:
resource/ledger:
attributes:
- { key: ritele.node.name, value: "postgresql:ledger", action: upsert }
- { key: deployment.environment.name, value: "production", action: upsert }
resource/cache:
attributes:
- { key: ritele.node.name, value: "redis:cache.internal", action: upsert }
- { key: deployment.environment.name, value: "production", action: upsert }
service:
pipelines:
metrics/ledger: { receivers: [postgresql], processors: [resource/ledger, batch], exporters: [otlphttp/ritele] }
metrics/cache: { receivers: [redis], processors: [resource/cache, batch], exporters: [otlphttp/ritele] }Stamp each scraped instance with ritele.node.name, named the way the map names it, so its readings land on that component. Use one pipeline per instance, since the stamp differs per instance and a processor cannot tell two receivers apart inside one pipeline. Readings then show on the component and fill its Server column in coverage.
The agent
Host, container and Kubernetes metrics come from the agent: this install's own image, run once per host or Kubernetes node with RITELE_COLLECTOR_PROFILE=agent. It runs the Collector alone and sends what it reads to RITELE_OTLP_ENDPOINT, this install's OTLP/HTTP address, under the environment named in RITELE_ENV, or the one named default when it is unset. It holds no credential, and its cloud detectors are off, so it calls no cloud metadata endpoint. What it sends is recorded only when the licence key grants infrastructure.
| RITELE_AGENT_TARGET | Reads |
|---|---|
host, the default | The machine, from host_metrics every 30 seconds: CPU, memory, load, filesystems, disks, network and paging. The process scraper is off, because it names every process and its owner. |
docker | The machine, and every container on it from docker_stats over the Docker socket, restarts included. |
kubernetes | The machine, and its node's kubelet from kubelet_stats: node, pod, container and volume usage, and each pod's and container's use of its limits. Run as a DaemonSet. Apps that send through their node's agent gain pod identity from k8s_attributes. |
cluster | Pod phases, workload replicas and node conditions from the Kubernetes API, from k8s_cluster. Run as one replica per cluster. |
On a Docker host
docker run -d --name ritele-agent --restart unless-stopped \ --user 1000 --group-add "$(stat -c %g /var/run/docker.sock)" --uts host --read-only \ -v /:/hostfs:ro -v /var/run/docker.sock:/var/run/docker.sock:ro \ -e RITELE_COLLECTOR_PROFILE=agent -e RITELE_AGENT_TARGET=docker \ -e RITELE_ENV=production -e RITELE_OTLP_ENDPOINT=http://ritele.internal:4318 \ ritele/ritele
| Given | Because |
|---|---|
/ mounted read-only at /hostfs | host_metrics reads the machine's filesystems and /proc through it. |
/var/run/docker.sock mounted read-only, and the socket's group | docker_stats lists the containers and reads their stats through the Docker API. |
--uts host | The host's hostname, which is what joins its containers and apps to it. |
--user 1000, --read-only | A non-root user on a read-only root filesystem. |
On Kubernetes
| Service account | Runs | May |
|---|---|---|
ritele-agent | The DaemonSet, one agent per node | get on nodes/stats and nodes/proxy; get, list, watch on pods, namespaces, nodes, ReplicaSets, Deployments, StatefulSets, DaemonSets, Jobs and CronJobs. |
ritele-cluster-agent | The one-replica Deployment | get, list, watch on events, namespaces, nodes, pods, ReplicationControllers, resource quotas, services and their status; the same workload kinds; HorizontalPodAutoscalers. |
Both roles are cluster-wide and have no verb that creates, changes or deletes anything. nodes/proxy reaches each kubelet's API, which is wider than its stats; grant it knowingly. The DaemonSet mounts the node's / read-only, runs as a non-root user with every capability dropped and a read-only root filesystem, and opens 4317 and 4318 on the node so apps can send through it. It connects to its own node's kubelet without verifying the certificate, since kubelet serving certificates are self-signed on most clusters.
What it creates
| Entity | Recorded from |
|---|---|
| Host | A resource that reports the machine itself, which is the agent's host_metrics. An app only links to the host it names. |
| Container | container.id, from docker_stats or an app outside Kubernetes. |
| Pod | k8s.pod.uid, from kubelet_stats, k8s_cluster or an app enriched by k8s_attributes. In Kubernetes a container is read as its pod. |
| Kubernetes node | k8s.node.name. |
| Workload | A pod's Deployment, StatefulSet, DaemonSet, CronJob, Job or ReplicaSet. |
| Service instance | service.instance.id on a service's own telemetry. |
| Datastore server, Kafka cluster | A datastore receiver's readings, and kafka.cluster.alias. |
The links between them — an instance runs in a container or pod, a container runs on a host, a pod belongs to a workload and is scheduled on a node — are what a service's Infra column follows. A host past the key's host limit is counted and not stored, and neither is anything on it; a host already stored stays. Each environment keeps at most 5,000 entities on embedded storage and 100,000 on Postgres and ClickHouse, or RITELE_MAX_ENTITIES_PER_ENV, and pods stop being added first.
Cutting series at the source
By default the cluster agent reports a pod phase, a ready state, a restart count and a CPU limit for every container in the cluster — kube-system included, each one its own series against the cap. Set RITELE_K8S_NAMESPACES to a comma-separated list and it saves only those namespaces' pods and containers. The trade-off: with a list set, the receiver no longer watches cluster-scoped objects, so node conditions stop being reported. The per-node DaemonSet's own reading of its node and kubelet is unaffected either way — it was never cluster-wide.
A Node service's heap is kept as one series per instance rather than one per V8 heap space, so an estate of Node services costs a fraction of the series it once did for the same heap readings.
Health
Every component on the map gets a status and the sentence behind it, in numbers. Open Health from the sidebar, or switch the map to Health mode to see each status where the component sits.
| Status | Means |
|---|---|
| Healthy | Every check it can be judged on is within its rule. |
| Degraded | A check is past its warning line. |
| Unhealthy | A check is past its critical line. |
| Silent | A component that reports on a regular rhythm has sent nothing for at least five minutes. The rhythm is learned from its last day of data. |
| Not monitored | Nothing it sends can be judged yet. It neither opens nor closes an issue. |
Health runs every minute, per environment. A status is entered after three of five bad minutes and left after five clean ones, so one slow minute does not open an issue; a fixed limit crossed in the latest minute, or an error ratio of half or more, enters at once. Checks that compare with usual wait for 24 hours of data to say what usual is, and the component says it is learning until then. A check with fewer than 30 requests in its window keeps its last status.
Rules are per component type, under Settings › Health rules. Any check on a component can be adjusted or muted from the component itself; a mute stops notifications, never the status.
Coverage
Each component says which signals reach it, so a status is never read as better informed than it is. The Components table shows the same thing as a grid.
| Column | Filled by |
|---|---|
| Traces | Spans from or to the component. |
| App metrics | The request metrics above, sent by the service. |
| Runtime | Event loop, heap and GC for Node; the jvm.* family for Java. |
| Pool | db.client.connection.* from any service that holds a pool. |
| Server | A receiver's readings stamped with ritele.node.name for that component. |
| Infra | For a service: the container, pod, Kubernetes node or host one of its instances runs on, from the agent, when the key grants infrastructure. The cell names the nearest one with readings. |
A cell reads received when something arrived in the last 15 minutes, silent up to a day, and missing after that.
What else Health does
| Feature | What it does | Needs |
|---|---|---|
| Infrastructure | Hosts, containers and pods get a status of their own from the agent's readings, listed under Health › Infrastructure. A container its engine stopped reads Stopped, not an outage. | The agent, and the key's infrastructure grant. |
| Likely source | One issue for the component that most likely broke, with the components it affects listed under it. A source is never milder than the issue it takes in, and the search stops at a message broker, so a starved consumer does not blame its producer. | Nothing to set. |
| Headroom | How far a service, database or queue is from its first limit. Try it in What-if opens on the hour that was modelled. | Headroom reads without a grant. The What-if step needs the key's what-if grant. |
| SLOs | A budget per service or entry point, built with a preview over the past, with burn-rate alerts and a monthly report as CSV or print. | The key's SLO grant. Without it the builder previews on the install's own data. |
| Change events | Deploys, restarts and status changes, marked on the charts and listed on the Events page. A deploy is read from a change in service.version. | service.version on the resource. |
| Trace links | An issue links to the traces from its bad minutes, and each opens in your own trace backend. | A URL template under Settings › Trace links. |
Setting a source up
Settings › Data sources generates the set-up for each source, with a row that turns to received when its first reading arrives: otel-kit, plain environment variables, an existing Collector, the agent on Docker, Kubernetes or a Linux VM, and receivers for PostgreSQL, MySQL, Redis, Kafka, MongoDB and HTTP checks.
Logs
However you send them — a plain SDK's OTEL_LOGS_EXPORTER=otlp, or otel-kit's own logs block, both above — Ritele turns your services' logs into one more Health signal: how many error and warning records a service produced, each minute. Nothing else survives the pipeline. No log line, attribute or body is stored here, and there is no log search on offer — only the counts, by severity.
The Error logs check works like the other checks built on a usual: degraded once a service's error-log rate climbs to three times its usual and at least once a minute, unhealthy at ten times usual and at least five a minute. Like every such check, it needs 24 hours of a service's own history before it knows what usual is, and reads as learning until then.
Connecting your Loki
Counts are built in — every service's error and warning rate is a Health signal with no setup at all. Lines are a step further: point Ritele at your own Loki and it reads them on request, masks what it's told to, and never stores one. Here is the setup end to end, in the order you'll actually hit each step.
1. Turn logs on
Open Settings › Logs. Under Install mode, choose Loki only — lines are read from your Loki and never stored — or set RITELE_LOGS=loki on the container; a mode chosen here always wins over the container's own setting. Loki and stored behaves the same way for now, with storage to follow later, so only pick it if you want to be ready ahead of time. Then, in Sources, choose Your Loki for each environment that should read from it. An environment reads from exactly one place, so its counts on the map and the lines it shows always agree with each other. Changing the install mode and choosing a source both need the admin role; every role can read these settings.
2. The Loki address
Give it Loki's own HTTP API base — not Grafana's address — for example http://loki.observability.svc:3100 inside a cluster, or your gateway's address in front of Loki. Set Auth to match how Loki is secured: None, Basic (a username and a password) or Bearer token. If your Loki serves more than one tenant, set Tenant · X-Scope-OrgID to the one your logs live under. If Loki sits behind a private certificate authority, paste its CA bundle so Ritele's server trusts it. A password or token is write-only — once saved it's never shown again — and saving one at all needs RITELE_SECRET_KEY set on the container first.
Would rather not expose Loki directly? Point Ritele at Grafana's own data source proxy instead: https://<grafana>/api/datasources/proxy/uid/<loki-uid>, with Auth set to Bearer token and a Grafana service-account token (Administration › Service accounts, in Grafana) as the value. Grafana forwards the request to Loki and applies its own permissions on the way through. Find <loki-uid> either in the address bar of any Grafana Explore page already querying your Loki — it's the datasource value there — or under Connections › Data sources › Loki, at the end of that page's own address.
3. Test the connection
Press Test connection. It calls Loki's label list and nothing else — no log line is read at this point. “Connected · Loki answered with 42 labels” means the address, the auth and the CA bundle, if one is set, all worked, and Loki handed back its list of label names. Anything else names what to check next: the address first, then the auth, then the CA bundle.
4. Tell it how your Loki names things
Ritele finds your lines by the labels your Loki already has — it never asks Loki to understand Ritele. Open one real log line for this environment, in Grafana or however you'd normally query Loki, and read its labels straight off it:
| Field | What to set | Example |
|---|---|---|
| Services are named by | The label whose value is a component's own name. | container, app or service_name |
| Environment label | Leave it blank if this Loki only ever holds a single environment's logs. Otherwise the label that carries the environment, and the value this one uses. | namespace holding prod |
| Level is in | A label, if your logs carry one. Leave it blank for a JSON logger that writes severity inside each line instead — Ritele reads it from there once a line comes back. | detected_level, or blank for pino |
| Trace ids are in | Structured metadata, the line text, or None. | line text for most JSON loggers; structured metadata for OpenTelemetry's own OTLP logs |
| Namespace label, Pod label | Optional. Kept with the setting and shown beside each line; reads by component don't narrow by either yet. | k8s_namespace_name, k8s_pod_name |
Set Services are named by and Ritele checks it straight away: it asks Loki for that label's values over the last hour and reports how many of this environment's own components match one exactly — something like “12 of 14 components in prod match a value of container exactly, from 340 values over the last hour.” A lower number than you expect is the first sign of a naming mismatch, and the values that didn't match are listed right beside it.
Read 20 lines runs the exact setup above for real, before anything is saved: the last hour's newest lines from the busiest matching component, masked the way this environment masks, and recorded in the audit like any other read. It's a check, not a save — nothing is kept from it. Save labels writes the setup only once you're happy with what came back.
5. Optional: open lines in Grafana too
Open in Grafana is a separate, optional setting, under Settings › Log links, that puts a button beside an issue, a trace or a component which opens the matching lines directly in your own Grafana — your browser follows the link, your Grafana login applies, and Ritele makes no call at all. It needs the address of a Grafana Explore page, not a dashboard: copy the address bar of Explore itself, which looks like …/explore?…&panes=…, never a dashboard's …/d/<uid>/… link. Paste a dashboard address and Ritele says so plainly and asks for an Explore one instead. From a dashboard, open a logs panel's menu and choose Explore to get there; or open Explore directly from Grafana's own navigation and run any query against your Loki data source.
If Grafana's own navigation has no Explore entry for you, your Grafana role is Viewer without Explore access, and the links Ritele builds won't open for you until that changes — ask whoever administers your Grafana.
6. Troubleshooting
| If you see | It usually means | Fix it by |
|---|---|---|
| No labels back, or the test is refused | The address, the auth, or the CA bundle is wrong, so Loki never actually answered. | Check the address is Loki's own API — not Grafana's, not a dashboard — match Auth to how Loki is really secured, and paste the full CA bundle for a private certificate authority. |
| Components don't match | The label's values don't spell a component the same way Ritele does: different case, a prefix, a short name. | Read the values that didn't match, beside the match count, and either change the label field or rename the handful that differ. |
| Every severity shows as — | Level is set to a label these logs don't actually carry, or a JSON logger only ever writes severity inside the line. | Leave Level blank and let Ritele read severity out of each line's body instead of a label that isn't there. |
| No trace link | Trace ids is set to None, or to structured metadata your Loki never attaches for this logger. | Switch Trace ids to the line text for a JSON logger that only ever writes the id inside the body; keep structured metadata for logs that arrived as OpenTelemetry's own OTLP. |
| “The read stopped at 200 lines” | Ritele reads lines a page of 200 at a time, newest first, so older lines in the window are still in Loki. | Press Load older lines under the list to read the next page back, or narrow the time range to the minutes you care about. |
Notifications
An issue can be announced where your team already looks. Nothing is sent anywhere until an admin adds a channel.
| Variable | Why |
|---|---|
RITELE_SECRET_KEY | Required, at least 16 characters. A channel's URL, routing key, API key or SMTP password is a credential, so it is sealed with AES-256-GCM under this key before it is stored. Back it up with the data: a changed key leaves every channel unreadable until it is saved again. |
RITELE_PUBLIC_URL | The address people open this install at. Without it a message carries no link rather than a broken one. |
RITELE_ROLE | Only an admin can add channels and rules. With sign-in on, the role is the signed-in person's, and this variable is ignored. It applies only with RITELE_AUTH=off, where unset means admin for every caller of port 8080; editor and viewer can read the settings but not change them. |
Open Settings › Notifications, add a channel and use Send test. The first channel in an environment comes with a default rule — unhealthy for five minutes, any domain — which you can change or switch off. How many channels an install may add is set by its licence key.
| Kind | Sent to | Receives |
|---|---|---|
| Webhook | An http or https URL | A JSON body, ritele.notification.v1. |
| Slack | An https incoming webhook | A text message. |
| Microsoft Teams | An https workflow webhook | An Adaptive Card with an Open in Ritele action. |
| PagerDuty | events.pagerduty.com or events.eu.pagerduty.com, with the routing key you save | An Events API v2 event, one alert per issue keyed ritele-issue-<id>. Send test is a change event, which pages nobody. |
| Opsgenie | api.opsgenie.com or api.eu.opsgenie.com, with the API key you save | One alert per issue, aliased ritele-issue-<id>. Send test is a P5 alert closed at once. |
| Alertmanager | An http or https base URL, posted at /api/v2/alerts | An alert, posted again every RITELE_HEALTH_NOTIFY_REFRESH_MINUTES (60 by default) while the issue is open, since Alertmanager resolves one nobody re-posts. |
Your SMTP server, over starttls (the default), tls or none | A plain-text message. A password is refused over none, and starttls stops if the server does not offer it. |
A rule announces an issue when it opens, when it gets worse and when it resolves, once each. A failed send is retried after 2, 4, 8 and 16 minutes. A channel is sent at most 30 messages a minute; the rest go as one digest. A mute holds back openings and escalations, but a resolution still goes wherever the opening went.
A rule can also match an SLO's burn rate and a component's owner, send to the channel set for that owner, remind while an issue stays open, and hold back during a weekly maintenance window. PagerDuty, Opsgenie and Alertmanager channels, owner routing, SLO burns and maintenance windows need the key's notifications grant. Incident tools are never folded into a digest: they de-duplicate per issue themselves.
How it works
For each service, how it handles each operation it serves: the calls it makes, in order, with the time each one takes. It is drawn from the traces Ritele keeps, so it shows what the service does, not what someone drew.
Select a component on the System map and open its How it works tab. The tab lists the operations the service serves, a summary of the chosen one, and its steps as a list. Open full view opens the same operation on its own screen, with the operations as a table and the steps drawn as a sequence diagram, one lane per component. The full view keeps the operation, path and scope chosen in the panel, and Open on map goes back with them. On a narrow screen the full view shows the steps as a list.
Only a component that sends its own traces has a How it works. A database, a queue or a service that sends no traces appears in its callers' steps instead.
What it shows
| On screen | What it is |
|---|---|
| Operations | Each route, RPC method or topic the service serves, or its own runs when it serves nothing. The panel shows each one's share of the service's traffic, p99 and error rate; the full view's table adds its kind and kept traces. |
| Summary | Two or three sentences from the traces: the calls in order, the status returned, where the time goes, an N+1 loop, and where a failure begins. It never says what a call is for. |
| Steps | Each call the operation makes, in the order it starts, with its typical (median) and p99 time. A database step is its statement as stored, or its operation and table when none is stored. An HTTP or RPC call answered by an instrumented service is named after that service and its route. A publish names its topic. A named internal span of the service is drawn as the service's own work. |
| A step's detail | Choose a step for its typical and p99 time, the share of requests that reach it and that fail at it, the reasons it fails with, and up to three example traces of the path, slowest first, beside how many kept traces took it. They open in the trace view, or in your own trace backend when Settings › Trace links has a template. |
| Loops and parallel calls | Calls made at the same time are drawn side by side, marked parallel. The same one to three calls repeated are one loop step, with how many times per request. In an HTTP or RPC operation, a loop of database or cache calls only, run one after another three or more times a request, is marked N+1. A loop in a message consumer or a job is not. |
| Time no span covers | Where no span of the service covers a noticeable stretch of the request, the diagram says so, as no span covers ≈31 ms here, before the step that follows it. |
| The caller | The first lane is whoever called the operation. A caller that sends no traces is drawn as caller not instrumented, never guessed. A consumer's caller is its topic. |
| What the numbers rest on | A line under every operation: how many kept traces, over which window, at which keep rate, and that only instrumented work shows. Every share is marked ≈ and is of all the operation's requests in the window. |
Common id shapes in a span name, digits, uuids, long hex strings and email addresses, are masked to ?, so process order 8812 and process order 8813 are one step. It is not a guarantee: a customer's name or another word-shaped id in a span name is shown as sent, so name spans by the operation, not the thing it acts on. Routes, peer names, topics and RPC methods taken from their attributes are not masked this way.
A database step shows its statement as ingest stored it, and RITELE_DB_QUERY_TEXT decides how: masked, the default, replaces literal values with ?; raw stores and shows them as sent; off stores none, so the step shows its operation and table. Masking is a filter for values, not a guarantee: a value inside an identifier, a short hex string or an unquoted word can pass. A change to the variable applies to spans ingested after it. Hours already built keep their statements until they age out, after 30 days at most, or the environment is cleared. See Security, R-12.
Paths, and how it fails
One operation can run more than one way. The usual path is drawn first, and up to two other successful paths and two ways it fails are offered beside it, each with its share of requests. A path is named only by what differs from the usual one: Cache miss, Also calls …, Without …, Repeats …, … once, Fails at … or Fails inside …. All paths draws them in one diagram: the steps they share, then a branch for each path from where they diverge.
| Drawn | Needs, in the window |
|---|---|
| The usual path | 5 kept successful traces of the operation, 3 of them on that path. |
| Another successful path | 3 kept traces of it. |
| How it fails | 20 kept failed traces of the operation, and 3 of that failure. |
Below that, the screen says how many traces it has of how many it needs, and offers a wider window when there is one, rather than drawing a path from a handful. Paths with fewer than 3 traces, or past the limit of three successful and two failure paths, are counted under the diagram as rarer paths.
End to end
This service draws only what the operation's own service does. End to end follows the trace into each instrumented service that answers its calls and draws that service's steps under the call, up to three services deep. A call to something that sends no traces stays one step. A path draws at most 40 steps, or 80 end to end; the rest are counted as steps not drawn.
How it is built
An hourly job builds it from the traces Ritele kept, in the same pass as the critical paths. An hour is built at the first run after it closes, so a new install can take up to two hours to show its first operations. On its first run with nothing built yet, the job builds the past day from the traces still held. An hour can be rebuilt only while its traces are held, 7 days at most.
The bundled Collector keeps every failed trace and every doc-trace, and a sample of the rest at the trace keep rate, 1% unless it is changed under Settings › Storage › Retention. So a failed or documented trace counts as one request, and any other as 100 ÷ the keep rate: at 1%, one kept success stands for 100 requests. That is why every share is an estimate. Each hour keeps the rate it was built with, so a rate changed later never reweighs an hour already built; when the rate changed after an hour began, the screen says that some hours may be weighted at the wrong rate.
| Not a step | Why |
|---|---|
A connection pool handing out a connection: a database connect span with no statement or operation | It is the request waiting for a connection, not a call it makes, and drawn it would double every database step. Its time counts as covered, so it never reads as a gap. Pool pressure is in Connection pools and Health. |
Framework spans: those carrying hook.name, fastify.type, express.type, koa.type, next.span_type or nestjs.type, and those whose names begin with request, handler, middleware or router | They wrap the request rather than doing a piece of its work. Their time counts as covered. |
| Work with no span | Nothing records it. It shows as time no span covers, never as a step that was guessed. |
When an admin clears an environment, which only an install with RITELE_ALLOW_RESET=true allows, its How it works and its notes go with it. Another replica can show its cached view for up to a minute after, and an hour already being rebuilt can still be written.
Notes
Editors and admins can label a path, with Label this path, and add a note to a step from its detail; viewers read them. A note is shown to everyone beside what the traces say, marked ✎ with who wrote it, when, and not from traces. A step's note stays on its step when the path's shape changes. When no path drawn for the window holds a note's step or path, the note is listed under the diagram, in Notes on what no drawn path shows.
A note holds at most 500 characters, and an operation at most 200 notes; clear one to add another. Control characters other than tab and newline, and the common zero-width and direction-override characters, are refused.
Exporting
From the full view, Export downloads what is drawn as Mermaid (.mmd), SVG, or PNG at twice screen size. Each carries the same steps, times and people's notes, and the line saying what the numbers rest on.
An export holds the step labels as shown, including any statement text this install keeps, and each note's text with the name of who wrote it. It is a plain file with no access control, so review it before it leaves your network.
Limits worth knowing
| Limit | What it means for you |
|---|---|
| Quiet operations | At the default 1% keep rate, an operation serving fewer than about 500 requests an hour has no drawn successful path. Failures are kept in full, but how it fails is drawn only once 20 failed traces are kept in the window. Widen the window or raise the keep rate. A rare successful path stays hidden unless the window spans enough hours. |
| Batch consumers | A consumer that takes messages in batches starts its own trace, linked to the publish it continues. The link is seen only when both traces were kept, so at a low keep rate the hand-off is named but not measured per publish. |
| Who can read it | Every signed-in role, viewers included, can read every environment's How it works, statement text included. With RITELE_AUTH=off there are no accounts: whoever reaches the port reads it, in the role RITELE_ROLE gives every caller. |
| Many paths over a long window | A window reads at most 16 MiB of stored paths. A path that does not fit is counted with the rarer paths, which under-counts it, and in the worst case a lighter path is drawn as the usual one. A shorter window holds fewer stored paths. |
| Busy services | A service's 50 busiest operations an hour are kept, set by RITELE_MAX_SERVED_OPERATIONS_PER_COMPONENT; the rest are counted together as overflow, with no paths. |
Not in this release
Planned, and not built yet. Nothing below is behind a switch or a licence key in this release.
| Planned | Until then |
|---|---|
| A Helm chart for the agent | Kubernetes manifests for kustomize, a Compose file and a systemd unit run the agent, and Settings › Data sources generates the command for Docker, Kubernetes and a Linux VM. |
| Kubernetes nodes, workloads and Kafka clusters judged for health | They appear under Health › Infrastructure with their readings but have no checks of their own yet. A node shows its host's status, a workload is read through its pods, and Kafka readings count towards the topic's checks. |
Capacity modelling
Rung 6 asks what a change would do. It is an open queueing model fitted to what you have actually been serving, and it needs one thing telemetry cannot supply.
Arrival rates come from the unsampled metrics path, service time from the critical paths, and fan-out from the observed call graph. What none of those contain is how much a component can do at once — a thread pool, a worker count, a connection limit. That appears in no span, so you declare it, per component, on the component itself.
Until you do — or until a component's own pool metrics state it — its limit is unknown, and a component with an unknown limit is not judged. It is left out of what gives out first, its utilisation claims nothing, and it is named in unknownLimits, so a prediction is never quietly built on a guess. What it would have to cover is still reported, as requests in flight, for you to hold against the limit you know.
CPU is the second limit. Where a service's processes report process.cpu.time and a CPU limit is known — ritele.cpu.limit on the resource, or one typed in the product, which wins — the model judges CPU as well as concurrency, and says which of the two gives out first.
The baseline is measured against per-minute detail, which Settings › Storage › Retention keeps for 48 hours to 7 days by default, depending on tier, and up to 7 days however it is set. Peak and Typical search whatever that window holds, and a custom window has to start within it.
What you get back
| Field | Means |
|---|---|
saturationOrder | What gives out first, and at what multiple of today's traffic. The headline answer. Only components with a known limit are listed, and anything that never gives out is absent rather than listed at infinity. |
saturationOrder[].resource | Which limit gives out: concurrency, the component's workers or connections, or cpu, the cores its replicas are allowed. |
loads[].limitKnown | Whether anyone declared, stated or measured the component's concurrency. When it is false, the two fields below say nothing about capacity. |
loads[].utilisation | Offered load over capacity under the scenario. At or above 1 the queue does not drain. A claim only where limitKnown is true. |
loads[].drains | Whether the queue empties at all. A queue that never drains waits for ever, which is not a number — this is the boolean that says so. Also a claim only where limitKnown is true. |
loads[].arrivalRate | Visits per second at the component, under the scenario and today. Only traffic with a traced path is counted; the rest is in untracedTraffic. |
loads[].inFlight | Requests in flight on average: visits per second times seconds per visit. The concurrency a limit would have to cover, reported whether or not a limit is known. |
loads[].cpu | Cores the component burns, summed over its replicas, under the scenario and today; the CPU limit per replica and where it came from; and utilisation, cores over replicas times limit. With no limit known there is no CPU claim, only cores. |
edges[] | Calls per second on each edge, under the scenario and today — how extra load at an entry point trickles down to the components behind it. |
latencies[] | Predicted p50 and p99 per entry point, against what it does today. Null when either end does not drain: there is no ratio, and a number there would invent one. |
unknownLimits | Components nobody declared, stated or measured a concurrency limit for. They are not judged: no utilisation claim and no place in saturationOrder. |
untracedTraffic | Calls the unsampled pipeline counted under an operation no traced path explains. Their cost is unknown, so they are in no figure above — named so you know the load is there. |
costedFromErrors | Entry points whose every sampled trace failed, so the cost per request is a failure's, not a success's. |
unmodelled | Components the traces name but nothing could be fitted to. Still in your system, absent from this answer. |
errorMargin | The model's own error, stated once. Not a confidence interval over your inputs. |
Traps worth knowing
Each of these has cost somebody real time, and each fails quietly rather than loudly.
Proving it works
Do not assume a signal is arriving because nothing errored. Ask for it.
Every screen is backed by a GraphQL API on the same port, so each rung can be checked directly.
# Rung 1–2 — is anything arriving, and in which environment? { environments } # Rung 3 — are entry points named, or are they raw URLs? { entryPoints(env: "production", from: "…", to: "…") { id rps p99 } } # Rung 4 — is the unsampled path feeding volume? { graph(env: "production", metrics: { from: "…", to: "…" }) { nodes { name red { callsPerMin errorRate p99 } } } } # Rung 5 — are pool readings landing, and who reported them? { saturation(env: "production", from: "…", to: "…") { component signal client { total limitTotal reporters } } } # Health — which sources are sending, and which signals reach which component? { dataSources(env: "production") { name kind pointsPerMin lastSeen silent } } { coverage(env: "production") { name coverage { traces { state } appMetrics { state } runtime { state } pool { state } } } }
On the saturation query, reporters is what makes a reading checkable — a total with no reporters behind it cannot be verified by the person reading it. An empty result means nothing is reporting yet, which is a different problem from a reading of zero.
On the Health queries, an empty dataSources means nothing has reached the install: check the address, port 4318, and that metrics are not switched off. In coverage, a cell that reads MISSING names the signal that never arrived, and its section above is what to recheck.
Integrating with AI
Give your coding assistant the brief below, and it has what it needs to connect a codebase without guessing: the three facts to ask you for, the rules, including complete traces, the otel-kit set-up for Node, the environment variables for everything else, and how to prove the signals arrived. It is plain Markdown, written for a model to follow.
- Copy the brief into your assistant's prompt, or give it the link beneath it to read.
- Add your own context at the end: which services are in scope, and the OTLP address and environment name if you know them. The brief tells the assistant to ask for what is missing.
- Review the diff. It should touch each service's start-up and its environment, plus the few internal spans and message-header hand-offs it adds.
- Restart the services, then check Settings › Data sources and the Coverage table. The brief gives the assistant the same checks, and asks it to report which outbound clients it could not instrument and where it added internal spans.
# Integrate this codebase with Ritele
You are connecting this codebase to Ritele, a self-hosted tool that reads standard OpenTelemetry (traces and metrics), draws the real architecture of a system, and judges the health of every component on it. There is no Ritele client library. You make each service send standard OpenTelemetry to one address, and nothing else.
Follow this brief exactly. Do not invent attribute names, metric names or options that are not written here. If something here conflicts with what you see in the code, say so instead of guessing.
## 1. Ask before you change anything
You need three facts. Ask for any you cannot read from the repository.
1. The Ritele OTLP/HTTP address as the services reach it, for example `http://ritele:4318`. Traces and metrics go to the same host; the SDK appends `/v1/traces` and `/v1/metrics`. Port 4318 is OTLP/HTTP and 4317 is OTLP/gRPC.
2. The environment name for each deployment (`production`, `staging`). It becomes `deployment.environment.name` and keeps environments apart.
3. Whether the services already send to an OpenTelemetry Collector. If they do, add Ritele to that Collector (section 6) instead of changing each service's endpoint.
Then list every deployable service, its language, and how it is deployed (Docker, Kubernetes, VM). One deployable is one component.
## 2. Rules for every service
- `service.name` is the component's identity everywhere. One per deployable, stable across deploys and replicas, never a version or a host name. Changing it later reads as one component disappearing and another arriving.
- `deployment.environment.name` is required in practice. Without it everything lands in an environment called `default`.
- Send traces and metrics. Health is judged from metrics, so a service that sends only traces gets a map and no status.
- Never sample the metrics path, and never put an identifier, email, token or URL with ids in a metric label or span attribute.
- Name routes, not URLs: server spans carry `http.route` (`/v1/customers/:id`), never the raw path.
- Do not add any `ritele.*` attribute that this brief does not list.
### Complete traces
A trace is complete when every call a request makes is a span under it. The trace view, the Paths, the critical path and How it works depend on it.
- Trace every outbound call as a child of the request: HTTP, gRPC, database, cache, queue and cloud SDK (AWS, GCP) clients. For each client the service uses, check that its language's auto-instrumentation covers it. If one is not covered, wrap the call in a client span (`SpanKind.CLIENT`) if that is a small change; otherwise leave it and report it. Do not guess that a client is covered.
- Propagate W3C `traceparent` on every hop. Keep the default propagators; never switch them off.
- Leave the SDK sampler at its default, `parentbased_always_on`: Ritele's Collector samples after the fact. If a service already samples, keep it parent-based (`parentbased_traceidratio`) with the same `OTEL_TRACES_SAMPLER_ARG` on every service, and do not add sampling where there is none.
- Carry context through every async hand-off, not only Kafka: SQS, RabbitMQ, Redis streams, background jobs. On produce, inject `traceparent` into the message headers (`propagation.inject`). On consume, extract it and parent the consumer span to it, or add a span link to it.
- Add named spans for the important internal steps of a request, such as `validate-transfer` or `score-risk`, with `tracer.startActiveSpan`. A handful per operation, not every function. End every span, in `finally`.
- Keep span names short and stable. No ids, amounts or user values in a name: put them in attributes, and keep attributes free of personal data.
- On database client spans, `db.operation.name` and `db.collection.name` (the table) should be set. Most instrumentations set them; add them only where one does not. Write queries with placeholders (`$1`, `?`), never values built into the SQL, so customer values stay out of query text.
## 3. Node.js services: use @omob/otel-kit
`@omob/otel-kit` is an MIT-licensed npm package that wraps the OpenTelemetry SDK. Use version 0.12 or later: before 0.12, a service started with `--require` ran two SDKs and reported its CPU and GC twice. Before 1.0 a caret range does not cross a minor, so write `"@omob/otel-kit": "^0.12.0"`, never `^0.10.0` or `^0.3.0`.
```bash
npm install @omob/otel-kit @opentelemetry/api
```
Create `src/instrumentation.ts`:
```ts
import { ExporterType, Telemetry } from "@omob/otel-kit";
const endpoint = process.env.OTEL_EXPORTER_OTLP_ENDPOINT;
Telemetry.start({
serviceName: "wallet-service",
serviceVersion: process.env.APP_VERSION,
environment: "production",
traces: {
exporter: ExporterType.OTLP,
otlp: { url: `${endpoint}/v1/traces` },
},
metrics: { exporter: ExporterType.OTLP, cpuUsage: true },
instrumentation: { ignoreIncomingPaths: ["/health"] },
handleShutdownSignals: true,
});
```
Replace `wallet-service` and `production` with this service's values. Then:
- Load it before everything else. CommonJS: `node --require ./dist/instrumentation.js dist/server.js`. ESM (`"type": "module"`): `node --import ./dist/instrumentation.js dist/server.js`, on the command line and not as an import inside the entry file, or Fastify, ioredis and kafkajs are patched too late.
- Do not register `@opentelemetry/auto-instrumentations-node` as well: no `--require` or `--import` of its `register` module, in the start command or `NODE_OPTIONS`, and no `registerInstrumentations(getNodeAutoInstrumentations())`. otel-kit registers them already, and a second set doubles every span and metric. Remove any you find. To configure one instrumentation, use `instrumentation.config`.
- Set `OTEL_EXPORTER_OTLP_ENDPOINT` on the service. With OTLP traces the kit sends metrics to the same place every 30 seconds with no further code.
- Do not set `OTEL_METRICS_EXPORTER=none`. It overrides the `metrics` block and silently switches Health off for that service. The Jaeger quick start suggests it; remove it.
- On Fastify, run `npm i @fastify/otel` and add `enable: [InstrumentationName.FASTIFY]` under `instrumentation`. Without it every request is one bare span with no `http.route`.
- Set `serviceVersion` when the service starts from a monorepo root, which would report the root package's version.
### Connection pools
A trace cannot show that a service holds seven of its ten connections. Register each pool after `Telemetry.start()`, never before, or the readings are discarded and nothing errors. Name it `host:port/database`; pools sharing a name add into one series.
```ts
import { observeConnectionPool } from "@omob/otel-kit";
observeConnectionPool({
name: "10.0.1.5:5432/billing",
system: "postgresql",
read: () => ({
max: pool.options.max,
used: pool.totalCount - pool.idleCount,
idle: pool.idleCount,
pending: pool.waitingCount,
}),
});
```
### Replicas and CPU
- `service.instance.id` is a random id per process by default. On Kubernetes set it to the pod name so a restart is not counted as a new instance.
- `CPU_LIMIT_MILLICORES` gives the kit the CPU limit of one replica. Set it only when the container has a CPU limit: with none, the Kubernetes downward API reports the node's capacity.
On Kubernetes, pass the pod's identity in through the downward API. Declare each variable before the one that uses it:
```yaml
env:
- name: CPU_LIMIT_MILLICORES
valueFrom: { resourceFieldRef: { resource: limits.cpu, divisor: 1m } }
- name: POD_NAME
valueFrom: { fieldRef: { fieldPath: metadata.name } }
- name: POD_UID
valueFrom: { fieldRef: { fieldPath: metadata.uid } }
- name: POD_NAMESPACE
valueFrom: { fieldRef: { fieldPath: metadata.namespace } }
- name: NODE_NAME
valueFrom: { fieldRef: { fieldPath: spec.nodeName } }
- name: OTEL_RESOURCE_ATTRIBUTES
value: service.instance.id=$(POD_NAME),k8s.pod.name=$(POD_NAME),k8s.pod.uid=$(POD_UID),k8s.namespace.name=$(POD_NAMESPACE),k8s.node.name=$(NODE_NAME),k8s.container.name=wallet-service
```
`k8s.container.name` is the container's `name` in the pod spec, which the downward API cannot supply; write it in.
### Messaging
kafkajs is traced by default. A consumer that reads with `eachBatch` starts a trace of its own and links to the producer's; its topic's volume comes from the `messaging.process.duration` it sends, so keep its metrics exported. For `@google-cloud/pubsub`, create the client with `enableOpenTelemetryTracing: true`, or it makes no spans.
## 4. Services in any other language
Use that language's standard OpenTelemetry SDK or agent (the Java agent, the .NET and Python auto-instrumentation, the Go SDK) and configure it with environment variables only:
```
OTEL_SERVICE_NAME=wallet-service
OTEL_RESOURCE_ATTRIBUTES=deployment.environment.name=production,service.version=2.4.1
OTEL_EXPORTER_OTLP_ENDPOINT=http://ritele:4318
OTEL_EXPORTER_OTLP_PROTOCOL=http/protobuf
OTEL_TRACES_EXPORTER=otlp
OTEL_METRICS_EXPORTER=otlp
OTEL_METRIC_EXPORT_INTERVAL=30000
OTEL_SEMCONV_STABILITY_OPT_IN=http,database,messaging
```
- Keep `OTEL_SEMCONV_STABILITY_OPT_IN` as written. Older HTTP instrumentation reports `http.server.request.duration` only with `http`, and the Java agent reports `db.client.operation.duration` and `messaging.process.duration` only with `database` and `messaging`. Without them the service's requests and its calls to databases are missing.
- Health reads these metrics: `http.server.request.duration`, `http.client.request.duration`, `rpc.server.call.duration`, `rpc.client.call.duration`, `db.client.operation.duration`, `messaging.process.duration`, `messaging.client.consumed.messages`, `db.client.connection.count`, `db.client.connection.max`, `db.client.connection.pending_requests`, `db.client.connection.wait_time`, `process.cpu.time`, and the `jvm.*` family on Java. If the library already emits them, nothing more is needed.
- For a pool the library does not instrument, emit the four `db.client.connection.*` metrics yourself. Send `db.client.connection.wait_time` under one name only, never both the old millisecond spelling and the current seconds one.
- To state the CPU limit of one replica, set the resource attribute `ritele.cpu.limit` to cores, `0.5` for a limit of `500m`.
## 5. Optional resource attributes
Add these only when you know the value. They shape the map and are never required.
- `ritele.domain`, `ritele.layer`, `ritele.zone`: the organisation's own grouping, layering and region.
- `ritele.component.type`: only where Ritele guesses the type wrong.
- `ritele.intended_dependencies`: the components this service means to call, for drift. Name them the way the map does: `postgresql:ledger`, `kafka:transfers`, `paystack`.
## 6. A Collector you already run
Add an exporter and one traces and one metrics pipeline beside the ones you have. Do not append a signal path to the endpoint: `otlphttp` adds `/v1/traces` and `/v1/metrics` itself, and an endpoint that already ends in one posts to `/v1/traces/v1/traces`.
```yaml
exporters:
otlphttp/ritele:
endpoint: ${env:RITELE_OTLP_ENDPOINT}
service:
pipelines:
traces/ritele:
receivers: [otlp]
processors: [batch]
exporters: [otlphttp/ritele]
metrics/ritele:
receivers: [otlp]
processors: [batch]
exporters: [otlphttp/ritele]
```
Do not sample before this exporter.
## 7. Prove it, do not assume it
Nothing errors when a signal goes to the wrong place, so ask for it. After the services restart and have served traffic for a minute or two:
- In Ritele, open Settings › Data sources. Every service you changed should be listed with points per minute above zero.
- Or query the API on port 8080:
```graphql
{ dataSources(env: "production") { name kind pointsPerMin lastSeen silent } }
{ coverage(env: "production") { name coverage { traces { state } appMetrics { state } runtime { state } pool { state } } } }
```
`traces` and `appMetrics` should read `RECEIVED` for every service, `runtime` for every Node or Java service, and `pool` for every service you registered a pool in. A cell that reads `MISSING` means that signal never arrived; the section above for that signal is what to recheck. An empty `dataSources` list means nothing reached Ritele: check the address, the port, and that metrics are not switched off.
## 8. Report back
Finish by listing the files you changed, the environment variables each service now needs and their values, every service you could not instrument and why, every outbound client you could not trace and where, every internal span you added and where, and anything in this brief you were unsure of.
Everything it reads
The complete list. Anything not here is not read, and nothing here is required except service.name.
Resource attributes
| Attribute | Effect |
|---|---|
service.name | The component's identity everywhere. The only thing genuinely required. |
deployment.environment.name | Which environment the component belongs to. Previous spelling deployment.environment also read. |
service.version | Shown on the component. |
ritele.node.name | On metrics scraped from a server by a Collector receiver, names the component on the map the reading belongs to. It does not rename a service: a service is named by service.name alone. |
ritele.component.type | Types the component explicitly, overriding the heuristic. Must be one of the kinds Ritele knows. |
ritele.domain | Groups components into a domain boundary on the map. |
ritele.layer | Places the component in a layer. Layer order is what makes a violation detectable. |
ritele.zone | Region or availability zone. |
ritele.intended_dependencies | The components this service means to call, comma-separated or as a string array — otel-kit's architecture.intendedDependencies emits it. Name each one the way the map does (the naming table below). Feeds drift. |
service.instance.id | Tells replicas of one service apart. CPU is counted per process, and the replica count is the median number reporting per minute. Without it, host.name with process.pid stands in. |
container.id, k8s.pod.uid, k8s.pod.name, k8s.namespace.name, k8s.node.name, k8s.container.name, host.id, host.name | Where an instance runs. An instance is matched to its container, else its pod, else its host, and those to the agent's own readings of them. |
ritele.trace.sample_probability | The share of traces the SDK keeps, from 0 to 1. Where a service sends no request metrics, health scales its span counts back up by it. |
ritele.cpu.limit | Cores one replica may use — 0.1 for a Kubernetes limit of 100m. With process.cpu.time, it lets the what-if model judge CPU. A Node process (process.runtime.name = nodejs) is counted as one core at most, since its JavaScript runs on one thread. |
Span attributes
| Attribute | Effect |
|---|---|
http.route, http.request.method | Named entry points and operation grouping. Previous spelling http.method also read. |
db.system.name, db.namespace | Datastore identity and type. Previous spellings db.system and db.name also read. |
messaging.system, messaging.destination.name | Topic or queue as a component, and producer-to-consumer linkage. |
messaging.operation.type | Producing versus consuming. Previous spelling messaging.operation also read. |
server.address, peer.service, rpc.service | Names an uninstrumented peer so it still appears. |
Metrics
| Metric | Effect |
|---|---|
db.client.connection.count | Connections held, by state. Plural db.client.connections.usage also read. |
db.client.connection.max | The pool ceiling. |
db.client.connection.pending_requests | Callers queued for a connection — the early warning. |
db.client.connection.wait_time | Time spent waiting. The deprecated spelling is milliseconds and the current one is seconds; Ritele converts on the name it was sent, so send one or the other, never both. |
process.cpu.time | CPU per process, per minute, in seconds, summed over cpu.mode. With ritele.cpu.limit it makes CPU a limit the what-if model can judge. |
http.*, rpc.*, messaging.*, db.client.operation.duration, nodejs.*, v8js.*, jvm.*, process.* | Health: request, consumer and runtime checks, per component and instance. The families that feed a check are listed under the metrics pipeline. |
How component names are formed
| Component | Named from | Example |
|---|---|---|
| Service | service.name, exactly as sent. | wallet-service |
| Database | <db.system.name>:<db.namespace>. Where the call carries peer.service, that names the part after the colon instead. | postgresql:ledger |
| Cache | The same rule, for redis, memcached and valkey. A cache rarely has a namespace, so the part after the colon is usually server.address. | redis:10.0.0.4 |
| Queue or topic | <messaging.system>:<messaging.destination.name>, keeping only the last / segment of the destination. | kafka:transfers |
| External API | peer.service, else server.address, else rpc.service. | paystack, api.paystack.co |
An entry in ritele.intended_dependencies is matched to a component in this order: the exact name; then system:name with the common spellings of a system folded together, so postgres:ledger finds postgresql:ledger; then a bare name that matches the part after the colon of exactly one component, so ledger finds postgresql:ledger when nothing else is called ledger. An entry that matches nothing is kept as written and reported as a missing dependency until a call to it is seen. Entries that look like a URL or an identifier (containing ://, @, ?, = or a space) are dropped. Declarations are read from every span a service sends, not only the traces that are kept, and into the intended model on every drift run, every 15 minutes, when the licence grants the intended model — the first 100 distinct entries per service, each bounded at 200 characters, and at most 2,000 declared dependencies per environment. A component whose dependencies a person has already described, by import, by accepting the map or by marking it, keeps that description and its declaration is ignored. A declared component cannot be unmarked by hand, because the next run would put it back; change the service's list instead. Sending a shorter list, an empty one, or no attribute at all withdraws what was dropped. An unexpected call is only reported from a component that is described — one with intended dependencies of its own, or one a person put in the model — so being named in another service's list does not flag every call a component makes. When a service runs several replicas, the list from the replica that reported last is the declaration, so it can change back and forth during a rolling deploy.