Measures
Metrics need signal-specific SQL. A counter reports its increase. A gauge reports its latest value or a chosen aggregate across series. A rate divides each series’ increase by elapsed time. A measure declares the reading; the engine derives the SQL for the query’s time grain.
{ "with": { "NetRxBytes": { "counter": { "value": "Value", "temporality": "cumulative", "where": "MetricName = 'system.network.io' AND Direction = 'receive'", "title": "Received", "format": "bytes" } }, "NetRxRate": { "rate": { "of": "NetRxBytes", "title": "Received / s", "format": "bytes" } } }}NetRxBytes is now an ordinary view scalar: select it, group by anything, put it in a stat and a plot on the same page. A counter’s per-bucket readings sum to the reading of the whole window, so a chart and a KPI over one agree.
Measures live on a timeseries table
A measure reads samples by series. Declare it on a timeseries table, which defines each series’ identity and timestamp:
{ "tables": { "metrics_sum": { "timeseries": { "from": "otel_metrics_sum", "timestamp": "TimeUnix", "identity": ["ServiceName", "ResourceAttributes", "MetricName", "Attributes"], "accumulation_start": "StartTimeUnix" } }, "parts": { "table": { "from": "system.parts", "with": { "Bytes": "sum(bytes_on_disk)" } } } }}| field | means |
|---|---|
from | the physical table (or a sibling alias to derive from). Defaults to the alias |
timestamp | the sample’s own time — what orders the samples of one series |
identity | what makes two rows samples of the same series |
accumulation_start | when the series began accumulating. Optional; see resets below |
A plain table accepts only SQL scalars; a measure on one is not expressible.
A timeseries either declares a series or extends one. Which it is follows from from, and so does what it has to state:
from | means | identity / timestamp |
|---|---|---|
| a table name, or omitted | declare a series over those rows | required |
{ "extend": "<alias>" } | take that alias’s shape and change some of it | optional, inherited |
{ "host_metrics": { "timeseries": { "from": { "extend": "metrics_sum" }, "identity": ["ServiceName", "HostName", "MetricName", "Attributes"] } }}extend inherits timestamp, identity and accumulation_start field-wise, so narrowing a stream you imported is one line rather than a copy that drifts from the pack that owns it. Anything you do state overrides the inherited one — including timestamp, for a stream ordered on a different column. It names a new alias rather than an override: metrics_sum keeps its own shape and a query picks which it reads. Because it says the target is an alias, a name the view doesn’t define is an error rather than a physical table nobody has.
A bare from naming a sibling alias still reads that alias’s rows, but declares a series from scratch — so it states its own identity and timestamp, and inherits nothing. A plain table derived from a timeseries inherits the shape unchanged.
Keep the emitter in identity. Two hosts’ en0 are different series. Keying on the attributes alone merges them, and every reading taken across more than one emitter under-reports — measured at up to 4.1x on real data. It also catches identifier reuse: a PID that gets recycled is a new resource, so its counter starts over rather than looking like a reset (70 of 5,159 PIDs were reused inside one six-hour window).
Narrowing identity is a claim about the data — “one resource per service” — and it is checkable: if it is false, series merge and every reading falls silently.
One table is one series shape. Two shapes over the same rows are two tables. That lets a query partition once, however many measures it selects: measured at 2 sorts / 1 window against 6 / 3 for three measures on separate keys, and 8.6x at six.
A measure names a metric rather than a shape. NetRxBytes is a measure; “the increase of whatever counter you are looking at” is not. The generic version has to hold for every metric the table can reach, which means trusting a flag rather than naming a known metric — and one emitter that sets IsMonotonic on a value that resets (process.runtime.cpython.gc_count does) makes it read 150x high with no predicate available to exclude it that keeps it generic.
So the layering runs: a base view declares the sample shape — identity, timestamp, accumulation_start — and the metric’s own pack declares the measure over it. That is why @opentelemetry/views/metrics_sum carries no measures at all, while @opentelemetry/views/hosts names NetRxBytesDelta, CpuBusySecsDelta and MemUsedBytes on the shape it inherits.
The kinds
counter — how much it went up
The reading is the increase over the window, taken per sample against the previous sample of the same series and clamped at zero.
{ "counter": { "value": "Value", "temporality": "cumulative", "where": "IsMonotonic AND MetricName = 'system.disk.io'" }}temporalityis required.cumulativesamples are running totals, so the reading is their increase;deltasamples are already increments, which needs no series key and no window at all and lowers to a singlesumIf. The two readings differ by orders of magnitude, so there is no safe default — and for OTel data the row says which it is, inAggregationTemporality.whereis the only place to narrow a counter. Filtering the reading afterwards cannot un-count the rows it excluded.- Keep
whereto columnsidentityalready covers — a key part, or a subscript of a map key part. Such a predicate is constant across a series, so it selects whole series rather than punching holes in one, and every measure on the table keeps sharing a single partition. - Only select rows that can rise. A value that legitimately falls — memory in use, thread counts, queue depth — has no increase to report, and summing its positive steps reports a figure many times the level it is measuring. Read those as a
level. (In OTel these are exactly the rows withIsMonotonic = false; the spec models them as gauges.)
A series whose accumulation began inside the window is credited in full. Observed once, it has no increase, so without that term an ephemeral series vanishes from the reading entirely.
That same term handles resets, because a reset is a new accumulation. The window partitions on the series and its accumulation_start, so a counter that restarts opens a new partition instead of producing a negative step that clamps to zero and loses everything since. Both need the table’s accumulation_start; without one, both are dropped.
level — what it was
Each series collapses to one number (reduce), and those combine across series (combine).
{ "level": { "value": "Value", "where": "MetricName = 'system.memory.usage'", "combine": "sum" }}reduce | one series reads as |
|---|---|
latest (default) | its newest sample — the snapshot a gauge is |
max / min | its peak / trough over the window |
sum | its samples added up |
avg | its samples averaged |
combine is required — sum, max, min, or avg. Combining one series’ reading with another’s is a modelling decision rather than a default: summing is right for a partitioned total (heap used across pods) and wrong for a shared one (a queue depth every replica reports), and the two are indistinguishable from the schema.
“Total memory” is latest × sum: the newest reading of each state, added. “Peak RSS across processes” is max × sum. “The busiest core” is max × max — and when reduce and combine are the same distributing fold (sum, max, min), the series key drops out entirely and the measure is a single aggregate over the raw rows.
avg × avg does not distribute, and that is the point. A plain average over the rows weights each instance by how often it reports, so an instance scraped twice as often counts twice. Reducing per series first and averaging those readings gives every series one vote. Measured on nodejs.eventloop.utilization across 6 instances, the unbiased reading was 2.6× higher than the average over samples. Use avg × avg for a sampled gauge — utilization, queue depth, load — where latest would swing with wherever the last scrape landed.
A level keys on the series and never on accumulation_start: it is what the thing was, so splitting a series at a reset would count each epoch’s last reading as a separate series.
rate — per second
{ "rate": { "of": "NetRxBytes" } }Each series’ own rise divided by its own elapsed time, and those per-series rates added up. of must name a counter measure the same table defines.
Computing per series is essential. A group almost always holds many series — every device, every process — and a single group-wide division gets it wrong twice: the rise a bucket counts includes the step into the bucket, so it spans one more sample interval than the samples inside the bucket do; and a series present for part of the group’s span has its rise divided by the whole span. On real data those cost 2x and 40x respectively.
A series with one sample in the group contributes 0 — there is no elapsed time to divide by. So a bucket narrower than the emitter’s scrape interval reads low: floor the bucket width at twice the scrape interval ("resolution": { "high": { "min_interval": { "value": 1, "unit": "minutes" } } }) so every series has two samples to work with.
Unlike a counter, a rate is not additive across buckets — it is a per-second reading, so summing buckets is meaningless. Averaging them is not the same as the window’s own rate either, since the buckets are not equally populated.
distribution — the shape of the observations
A histogram point carries a fixed ladder of buckets rather than a value, so it reads differently from the three above: distribution declares the shape and is not itself selectable — a distribution is not a number. Point readings at it by name and select those.
{ "with": { "Latency": { "distribution": { "buckets": "BucketCounts", "bounds": "ExplicitBounds", "sum": "Sum", "min": "Min", "max": "Max", "temporality": "cumulative", "where": "MetricName = 'http.server.duration'" } }, "LatencyP95": { "quantile": { "of": "Latency", "at": 0.95, "format": { "duration": {} } } }, "LatencyMean": { "mean": { "of": "Latency" } }, "Requests": { "observations": { "of": "Latency" } } }}| reading | is | needs |
|---|---|---|
quantile (at) | interpolated inside the bucket it lands in | — |
observations | how many were recorded | — |
mean | their total over their count | sum |
minimum / maximum | the smallest / largest the points reported | min / max |
Declaring it once makes the readings cheap: a p50, a p95 and a p99 over one distribution cost one fold of the bucket arrays rather than three. observations needs no column — the data model defines the buckets as summing to the count, and they do (verified equal to an independent max − min reading on live data, 2,274 either way).
temporality is required for the same reason a counter needs it: cumulative buckets are running totals, so the reading is their per-series increase — the counter’s arithmetic applied elementwise to an array.
A distribution must select exactly one metric. The fold adds bucket arrays elementwise, and two metrics with different ladders have nothing meaningful to add. ClickHouse pads rather than refusing — summing a 16-bucket ladder into a 23-bucket one — so a loose where is a silently wrong answer rather than an error. A quantile refuses (reads NULL) when it sees more than one ladder, but keep the where to one metric and don’t rely on that.
Two things a quantile can only approximate, both inherent to bucketed histograms: it interpolates linearly inside a bucket, so it is never finer than the ladder; and a quantile landing in the overflow bucket above the highest bound reads that bound, because nothing above it was recorded.
Check the claims against your data
A measure is a set of assertions about a stream, and every one of them can be false while the frame renders perfectly — the reading is a number either way, just the wrong one. Nothing catches these for you, so check them once when you declare a measure, and again when an emitter changes.
Each check below is one query against the table the measure is declared on. <identity> is that table’s identity list, <ts> its timestamp, <start> its accumulation_start, and <where> the measure’s own predicate.
Is identity really what makes a series? A series cannot hold two samples at one instant, so a group with more than one row is two series merged — every reading taken across the pair under-reports.
SELECT countIf(n > 1) AS merged, count() AS groupsFROM (SELECT count() AS n FROM <table> WHERE <where> GROUP BY <identity>, <ts>)Narrowing is a claim, and the answer depends on scope: keying the OTel sums on ServiceName, HostName, MetricName, Attributes merged 1,440 of 843,202 groups across a whole deployment, and 0 of 678,917 once scoped to a single emitter’s service.
Does the counter’s value only rise? Take rise against fall within one accumulation epoch, so a genuine reset — which opens a new epoch — is not counted as a fall.
SELECT sum(if(Value > Prev, Value - Prev, 0)) AS rise, sum(if(Value < Prev, Prev - Value, 0)) AS fallFROM (SELECT Value, lagInFrame(toNullable(Value)) OVER (PARTITION BY cityHash64(<identity>, <start>) ORDER BY <ts>) AS Prev FROM <table> WHERE <where>)The two populations are nine orders of magnitude apart, so any bound in between works: measured over 1h of live OTel sums, non-monotonic metrics gave fall/rise of 0.911 and monotonic ones 1e-9.
Don’t trust a monotonicity flag the emitter sets. process.runtime.cpython.gc_count declares itself monotonic while reporting the current per-generation counts, which drop to zero on every collection — measured falling 27,981 against a rise of 27,769 over 3h, a counter reading ~150x the real increase. It was the only such metric of 170, and it is invisible in aggregate: against the whole table’s rise those same falls read 1e-9. Probe the measure’s own rows.
Is temporality the right way round? A cumulative series accumulates against one start; a delta point carries its own.
SELECT countIf(starts = samples) AS delta_shaped, countIf(starts = 1) AS cumulative_shaped, count() AS seriesFROM (SELECT count() AS samples, uniqExact(<start>) AS starts FROM <table> WHERE <where> GROUP BY <identity> HAVING count() > 1)Measured over the same window, all 10,586 cumulative series had exactly one start and all 59 delta series had one per sample — nothing sat between the two, so a mixed answer means the source mixes both, which is two measures rather than one.
What you get back
A measure expands into ordinary view scalars, so the reading and its supporting scalars both show up in completion and in a SELECT *:
| name | is |
|---|---|
X | the reading |
_series_key | the series a row belongs to, as one hash |
_series_epoch | the series and its accumulation epoch, for cumulative |
_epoch_prev_time | the previous sample’s time within the epoch |
_series_prev_time | the previous sample’s time within the series |
XPrev | the previous sample’s value — the only per-measure helper |
The _series_* scalars are shared by every measure on the table, so however many a query selects it sorts the partition once. Read X; the rest are supporting scalars, and a query that reaches none of them costs nothing — the compiler prunes what a query doesn’t reference.
The compiler also narrows the scan to the rows the selected measures need, by AND-ing the union of their wheres into the table’s predicate. That is worth an order of magnitude (8.4x fewer rows, 21x less CPU on live data), and it is why where is the only place to narrow a measure — a filter applied after the reading cannot prune anything. A scalar that merely composes measures inherits their scopes, so a CpuSecs adding a cumulative and a delta counter narrows exactly as the two do. A measure that reads every row stands the narrowing down, since it reads what it scans, and projecting * stands it down too. A predicate over a column with a text index (on otel_logs, the attribute maps, Body and TraceId) stays inside the aggregate: ClickHouse reads a WHERE over such a column from the index before it chooses a projection, and a scan narrowed by it forfeits every projection of the table, so the compiler leaves that term out of the scan.
Narrowing the scan drops the groups it empties. A GROUP BY returns only the keys its scan still has rows for, so a query selecting MemUsedBytes leaves out a host that reports no memory metric. That is the honest reading: the -If aggregates behind a measure return 0 for an empty selection, so the alternative is a host listed at 0 bytes used, which it never reported.
Only a measure’s where earns that trade. An -If the query writes itself does not: countIf(Status = 'ERROR') grouped by service reads a true 0 for a service with no errors, and dropping the row would lose that answer. The same goes for an aggregate scalar a view declares as plain SQL ("Errors": "countIf(IsError)") rather than as a measure — a predicate written there says nothing about which rows the reading is about, so every group stays. A SELECT DISTINCT folds rows the same way a GROUP BY does and is treated the same. A query that folds no rows narrows by every source of predicates, because narrowing takes no row away that it would otherwise return.
A folding query pushes only the conjuncts of a measure’s where that can skip part of the scan. MetricName = 'codex.tool.call' AND Attributes['success'] != 'true' narrows to the metric alone. The second half reads a value out of a map without matching it, and an index over a map’s values answers which parts hold a given value and nothing about which parts hold some other value — so it skips nothing, while it would empty every group whose calls all passed. A tool with no failures now reads 0, over the same data.
Two shapes are pushed, and everything else is left inside the aggregate:
- A conjunct that reads its columns whole — named outright (
Device != 'lo') or through a call over one (toStartOfFiveMinutes(Timestamp) >= …, the shape a sort key takes). These can match the table’s key or an index expression however they compare. - A positive match on a value read out of a column —
Attributes['decision'] IN ('denied', 'deny'), answered by the map’s value index.
So !=, NOT IN, NOT LIKE, LIKE, NOT BETWEEN and IS NOT NULL over a map lookup all stay in the aggregate and keep their groups. The list is what gets pushed rather than what gets held back, so a spelling nothing here covers widens the scan instead of dropping rows.
To prune a folding query by a predicate that is not a measure’s, write it in the query’s own where (or in a derived table, for a scope a whole family of panels shares). That prunes the scan exactly as narrowing would, and it states that the groups it empties should be gone.