> ## Documentation Index
> Fetch the complete documentation index at: https://docs.conduktor.io/llms.txt
> Use this file to discover all available pages before exploring further.

# Gateway request lifecycle metrics

> Per-stage latency metrics that break a request/response inside Gateway into 11 timed segments, from socket read to socket write, to help investigate latency.

Gateway's round-trip latency metric (`gateway_latency_request_response`) tells you a request was slow, but not *where* it was slow. When latency spikes, you are left choosing between thread dumps and profilers to work out whether the bottleneck is in Gateway authorization, an Interceptor, the upstream queue, or Kafka itself.

Request lifecycle metrics help identify where requests and responses spend most of their time. They split each request into 11 timed segments — from the moment Gateway reads a request off the client socket to the moment it writes the response back — each exposed as a Prometheus histogram tagged by Kafka API key and Interceptor name (when applicable).

<Note>
  Lifecycle metrics are disabled by default because they add high-cardinality series to your Prometheus. Enable them when you need a per-stage breakdown, and [scope them](#scope-the-metrics) to the API keys and Interceptors you care about.
</Note>

## Enable lifecycle metrics

Set the master switch to `true`:

```bash theme={null}
GATEWAY_FEATURE_FLAGS_LIFECYCLE_METRICS=true
```

If you deploy Gateway with the [Conduktor Helm chart](https://github.com/conduktor/conduktor-public-charts/tree/main/charts/gateway), enable and scope the metrics through Helm values instead — see [Configure with the Helm chart](#configure-with-the-helm-chart).

## Scope the metrics

Once enabled, all 11 metrics record for every Kafka API key, and the two per-Interceptor metrics record for every Interceptor. That's the highest-cardinality configuration. You can narrow it with the settings in [Lifecycle metrics scoping](/guide/conduktor-in-production/deploy-artifacts/deploy-gateway/environment-variables#lifecycle-metrics-scoping).

A few Kafka API keys, mainly `PRODUCE` and `FETCH`, carry most of the traffic. Most of the others are administrative and send few requests, so their charts are nearly empty. Scope the API keys to the ones you want to investigate.

Valid API keys are the [`ApiKeys` enum constant names](https://github.com/apache/kafka/blob/trunk/clients/src/main/java/org/apache/kafka/common/protocol/ApiKeys.java) — uppercase with underscores, such as `PRODUCE`, `FETCH`, and `OFFSET_COMMIT` (not the CamelCase protocol names like `OffsetCommit`). They're case-insensitive. If you set a name Gateway doesn't recognize, it fails to start and logs the full list of valid values. The following table suggests API keys for each monitoring goal:

| Goal | Example API keys | Why | What drives the rate |
| - | - | - | - |
| Data path | `PRODUCE`, `FETCH` | The data path — one request per batch, continuously | Throughput |
| Consumer group | `HEARTBEAT`, `OFFSET_COMMIT`, `OFFSET_FETCH`, `FIND_COORDINATOR`, `JOIN_GROUP`, `SYNC_GROUP` | Per-consumer-group housekeeping on timers | Consumer count × interval (heartbeat \~3s, auto-commit \~5s by default), group membership changes |
| Connection | `METADATA`, `API_VERSIONS` | Connection setup and rebalances | Topology changes, reconnects |
| Administration | `CREATE_TOPICS`, `DELETE_TOPICS`, `DESCRIBE_CONFIGS`, `ALTER_CONFIGS`, `LIST_GROUPS` | Control plane, driven by tooling not data flow | Admin activity — near zero in steady state |

For example, to scope the metrics only for the data path and two Interceptors, `encrypt` and `decrypt`, set the following configuration:

```bash theme={null}
GATEWAY_LIFECYCLE_METRICS_API_KEYS=PRODUCE,FETCH
GATEWAY_LIFECYCLE_METRICS_INTERCEPTORS=encrypt,decrypt
```

## Configure with the Helm chart

If you deploy Gateway with the [Conduktor Helm chart](https://github.com/conduktor/conduktor-public-charts/tree/main/charts/gateway), enable and scope the metrics through the `metrics.lifecycle` values rather than setting the feature flag or environment variables yourself:

```yaml theme={null}
metrics:
  lifecycle:
    enable: true
    # Optional cardinality scoping — both default to ALL:
    # apiKeys: PRODUCE,FETCH
    # interceptors: encrypt,decrypt
```

Setting `metrics.lifecycle.enable: true` emits `GATEWAY_FEATURE_FLAGS_LIFECYCLE_METRICS=true` for you. The chart emits `GATEWAY_LIFECYCLE_METRICS_API_KEYS` and `GATEWAY_LIFECYCLE_METRICS_INTERCEPTORS` only when you set `apiKeys` or `interceptors` to a value other than `ALL`.

| Helm value | Description | Default value |
| - | - | - |
| `metrics.lifecycle.enable` | Enable the request lifecycle metrics feature. | `false` |
| `metrics.lifecycle.apiKeys` | Comma-separated list of Kafka API key names to record (for example `PRODUCE,FETCH`), or `ALL`. Scopes the `api_key` dimension of all 11 metrics. | `ALL` |
| `metrics.lifecycle.interceptors` | Comma-separated list of Interceptor names to record, or `ALL`. Scopes the two per-Interceptor metrics only. | `ALL` |

`metrics.lifecycle.enable` works independently from `metrics.grafana.enable`. When you enable both, the chart deploys the three [Grafana dashboards](#grafana-dashboards) automatically, alongside the main Gateway dashboard.

## The request lifecycle

The following diagrams show the lifecycle of a single request/response inside Gateway. Each arrow is one timed segment, labeled with the metric that records it and the boundary it spans.

The edge labels drop the common metric prefix and suffix: prepend `gateway.request.` to each segment label in the request path and `gateway.response.` in the response path, and append `.duration` to all of them (so `preprocess` is `gateway.request.preprocess.duration`).

The request path runs from the client read (`T0`) through the round-trip to Kafka (`T6`):

```mermaid theme={null}
flowchart TB
    C1([Client]) -->|network read| T0(("T0"))
    T0 --- S1["preprocess (T0 -> T1)"] --> T1(("T1"))
    T1 --- S2["authorization (T1 -> T2)"] --> T2(("T2"))
    T2 --- S3["interceptor (T2 -> T3)"] --> T3(("T3"))
    T3 --- S4["rebuilder (T3 -> T4)"] --> T4(("T4"))
    T4 --- S5["upstreamqueue (T4 -> T5)"] --> T5(("T5"))
    T5 --- S6["upstreamwait (T5 -> T6)"] -->|request| Kafka(["Kafka cluster"])
    Kafka -->|response| T6(("T6"))
```

The response path runs from the Kafka response (`T6`) back to the client write (`T10`):

```mermaid theme={null}
flowchart TB
    Kafka(["Kafka cluster"]) -->|response| T6(("T6"))
    T6(("T6")) --- S7["rebuilder (T6 -> T7)"] --> T7(("T7"))
    T7 --- S8["interceptor (T7 -> T8)"] --> T8(("T8"))
    T8 --- S9["authorization (T8 -> T9)"] --> T9(("T9"))
    T9 --- S10["send (T9 -> T10)"] --> T10(("T10"))
    T10 -->|network write| C2([Client])
```

`gateway.request.total.duration` spans the whole path, `T0 -> T10`.

Each boundary marks a lifecycle event:

| Boundary | Event |
| - | - |
| `T0` | Request received — Gateway reads it off the client socket |
| `T1` | Request preprocessing done (request parsing, deserialization, and Gateway authentication) |
| `T2` | Request authorization (ACL check) done |
| `T3` | Request Interceptors done |
| `T4` | Request rebuilt, ready to serialize and send to Kafka |
| `T5` | Request sent to the Kafka wire |
| `T6` | Kafka response received |
| `T7` | Response rebuilt |
| `T8` | Response Interceptors done |
| `T9` | Response authorization (ACL check) done |
| `T10` | Response written back to the client socket |

<Note>
  Request/response rebuilding refers to unpacking Kafka requests and responses, then modifying them for Gateway features such as encryption, Virtual Clusters, and cluster switching. Gateway then sends the modified request to the Kafka cluster, or the modified response back to the client.
</Note>

## Lifecycle stages and metrics

In Prometheus, these metrics use underscores instead of dots and end in `_seconds`, so `gateway.request.preprocess.duration` is scraped as `gateway_request_preprocess_duration_seconds` with `_bucket`, `_sum`, `_count`, and `_max` series. Following are the metric names in Prometheus:

| Prometheus metric | Segment | What it measures | Labels |
| - | - | - | - |
| `gateway_request_preprocess_duration_seconds` | T0 -> T1 | Request parsing, deserialization, and Gateway authentication (principal extraction) | `api_key` |
| `gateway_request_authorization_duration_seconds` | T1 -> T2 | The request ACL check | `api_key` |
| `gateway_request_interceptor_duration_seconds` | T2 -> T3 | One sample per request Interceptor. `_count` equals Interceptors run, not requests | `api_key`, `interceptor` |
| `gateway_request_rebuilder_duration_seconds` | T3 -> T4 | Rebuilding the request after Interceptors | `api_key` |
| `gateway_request_upstreamqueue_duration_seconds` | T4 -> T5 | Waiting for a writable broker connection, then writing the request to the Kafka wire | `api_key` |
| `gateway_request_upstreamwait_duration_seconds` | T5 -> T6 | Waiting for Kafka to respond — the upstream round trip | `api_key` |
| `gateway_response_rebuilder_duration_seconds` | T6 -> T7 | Deserializing and rebuilding the Kafka response | `api_key` |
| `gateway_response_interceptor_duration_seconds` | T7 -> T8 | One sample per response Interceptor. `_count` equals Interceptors run, not requests | `api_key`, `interceptor` |
| `gateway_response_authorization_duration_seconds` | T8 -> T9 | The response ACL check | `api_key` |
| `gateway_response_send_duration_seconds` | T9 -> T10 | Serializing the response and writing it back to the client socket | `api_key` |
| `gateway_request_total_duration_seconds` | T0 -> T10 | The full pipeline, end to end | `api_key` |

<Note>
  The `interceptor` label is the Interceptor's configured name (for example `encrypt` or `guard-schema-payload-validate`), not its plugin class.
</Note>

## How to read the metrics

Three useful ways to read each metric:

* **Mean**: `rate(_sum) / rate(_count)`. `_sum` is the total time of all requests and `_count` is how many requests there were.
* **Worst case**: `_max`, the slowest recent request.
* **Share of requests over a threshold**: for example, `1 - rate(_bucket{le="0.1"}) / rate(_count)` gives the share that took longer than 100ms. This is the form the dashboards use. `le` ("less than or equal to") is in seconds.

The threshold must be one of the histogram buckets, which are the same for all 11 metrics: **1ms, 10ms, 50ms, 100ms, 500ms, 1s, 5s, 30s**. They are chosen to show which stage has crossed into latency that clients notice.

<Tip>
  Reading `gateway_response_send_duration_seconds` (T9 -> T10): Kafka requires in-order responses per connection, so a spike here usually means a slow request ahead of it on the same connection is holding up the line, not a slow client socket. Check the other stages to find the actual bottleneck. `T10` also fires when the write completes whether or not the client is still connected, so neither `response.send` nor `request.total` confirms the response reached the client.
</Tip>

## Grafana dashboards

Conduktor ships three ready-made dashboards based on these request lifecycle metrics, at [charts/gateway/grafana-dashboards](https://github.com/conduktor/conduktor-public-charts/tree/main/charts/gateway/grafana-dashboards).

If you deploy Gateway with the [Conduktor Helm chart](https://github.com/conduktor/conduktor-public-charts/tree/main/charts/gateway) and enable both `metrics.lifecycle.enable` and `metrics.grafana.enable`, the chart deploys these three dashboards for you, alongside the main Gateway dashboard. See [Configure with the Helm chart](#configure-with-the-helm-chart).

Otherwise, import them into Grafana manually and select your data source, `job`, and `pod` at the top of each.

| Dashboard | Use it to | Filter |
| - | - | - |
| **Conduktor Gateway Lifecycle - Overview** | See the whole pipeline at a glance and find which stage owns the latency | `api_key` |
| **Conduktor Gateway Lifecycle - by API key** | Compare a single stage across API keys | `api_key` |
| **Conduktor Gateway Lifecycle - by Interceptor** | Compare Interceptor latency across Interceptors | `interceptor` |

### Conduktor Gateway Lifecycle - Overview

This dashboard has four sections. The first two give an overview, and the last two show each stage in detail.

#### Overview per API key

Total request latency (`T0 -> T10`), with one line per API key. The four panels show the mean, the max, the throughput, and the share of requests slower than 1s.

In the following screenshot, `DESCRIBE_LOG_DIRS` stands out: its mean latency spikes to several seconds, and at times all of its requests take longer than 1s. The other API keys stay under a second.

<img src="https://mintcdn.com/conduktor/CLmvmqMPOlxwR4Mx/images/gateway-lifecycle-overview-per-api-key.png?fit=max&auto=format&n=CLmvmqMPOlxwR4Mx&q=85&s=2ba9480c15d94846a283338d1ad57859" alt="Overview per API key section, showing total request latency per API key" width="3024" height="1294" data-path="images/gateway-lifecycle-overview-per-api-key.png" />

#### Overview per stage

The same four panels for each stage, with one line per stage. Here the threshold is 100ms.

In the following screenshot, most latency sits in the `upstreamwait` stage of `FETCH` requests — the time Gateway waits for Kafka. Its mean runs 100–500ms: Kafka holds a `FETCH` request for up to `fetch.max.wait.ms` (default 500ms) when no new message arrives. `request.interceptor` spikes occasionally, and the other stages stay near zero.

<img src="https://mintcdn.com/conduktor/CLmvmqMPOlxwR4Mx/images/gateway-lifecycle-overview-per-stage.png?fit=max&auto=format&n=CLmvmqMPOlxwR4Mx&q=85&s=e0044221f152ac107a485963de95a13d" alt="Overview per stage section, showing latency per lifecycle stage" width="3024" height="1290" data-path="images/gateway-lifecycle-overview-per-stage.png" />

#### Request path and Response path

Each stage has its own panel showing the mean, the max, and the share of requests slower than 500ms. A "% in bucket" panel next to it shows how that stage's latency spreads across the histogram buckets.

In the following screenshot, which shows the first two stages of the request path, nearly all `request.preprocess` requests take under 1ms, though the max reaches about 60ms. Every `request.authorization` request takes under 1ms.

<img src="https://mintcdn.com/conduktor/CLmvmqMPOlxwR4Mx/images/gateway-lifecycle-per-stage.png?fit=max&auto=format&n=CLmvmqMPOlxwR4Mx&q=85&s=7c8dc0d1fdc26e2878bca763505f1654" alt="Request path section, with a panel and a % in bucket panel for each stage" width="3000" height="1298" data-path="images/gateway-lifecycle-per-stage.png" />

### Conduktor Gateway Lifecycle - by API key

The dashboard has one section for each stage, plus a **total (T0 -> T10)** section at the end. Each has three panels, with one line per API key: the mean, the max, and the rate. Use it to spot an API key that's slow in one stage, for example a `FETCH` that lags in `upstreamwait` while the other API keys stay flat. Use the API key dropdown to focus on a single API key. A section shows "No data" when no API key runs at that stage.

In this example, which shows the first two sections, every API key is fast in both stages: the **preprocess (T0 -> T1)** mean stays under 3ms and the **authorization (T1 -> T2)** mean under about 120µs. The max has occasional spikes, up to about 65ms for preprocess and 4ms for authorization.

<img src="https://mintcdn.com/conduktor/CLmvmqMPOlxwR4Mx/images/gateway-lifecycle-by-apikey.png?fit=max&auto=format&n=CLmvmqMPOlxwR4Mx&q=85&s=f7592b51fe60e2d7b0944b4df433086a" alt="Lifecycle by API key dashboard, with preprocess and authorization sections split by API key" width="3012" height="1472" data-path="images/gateway-lifecycle-by-apikey.png" />

### Conduktor Gateway Lifecycle - by Interceptor

This dashboard has two sections, one for each Interceptor stage: **request.interceptor (T2 -> T3)** and **response.interceptor (T7 -> T8)**. Each has three panels, with one line per Interceptor: the mean, the max, and the rate of Interceptor runs. Use the `interceptor` dropdown to focus on a single Interceptor. A section shows "No data" when no Interceptor runs at that stage.

In this example, `simulator-benchmark` stands out on the request path: its mean spikes to about 800ms and its max stays between 1s and 2s, while the other Interceptors stay close to zero. None of the Interceptors run on responses, so the response section shows "No data".

<img src="https://mintcdn.com/conduktor/CLmvmqMPOlxwR4Mx/images/gateway-lifecycle-by-interceptor.png?fit=max&auto=format&n=CLmvmqMPOlxwR4Mx&q=85&s=3954c658e9a999ec9cb57af16fb0d02e" alt="Lifecycle by Interceptor dashboard, with request.interceptor and response.interceptor sections split by Interceptor" width="3014" height="1578" data-path="images/gateway-lifecycle-by-interceptor.png" />

## When metrics don't fire

Not every request runs the full `T0 -> T10`. Denials (from authorization checks or Interceptors), failures, timeouts, and cache hits cause the request to leave the lifecycle at different points. Therefore, a stage with no samples likely means the traffic never reached that stage. The common cases:

| Scenario | Metrics that fire | Metrics skipped |
| - | - | - |
| Authorization denial | `request.preprocess`, `request.total` | Everything else |
| Request Interceptor short-circuit or timeout | Request-path metrics up to the failing Interceptor, plus `request.total` | `request.rebuilder`, `request.upstreamqueue`, `request.upstreamwait`, all `response.*` |
| Produce with `acks=0` (fire-and-forget) | All request-path metrics through `request.upstreamqueue` | `request.upstreamwait`, all `response.*`, and `request.total` (no response is sent, so `T10` never fires) |
| In-flight timeout — Gateway expires the request after `GATEWAY_INFLIGHT_REQUEST_EXPIRY_MS` (default 330s) | All request-path metrics through `request.upstreamqueue`, plus `request.total`, which records at least \~330s — above the top 30s bucket, so it lands only in `+Inf` | `request.upstreamwait`, all `response.*` |
| Lost upstream connection | Request-path metrics through `request.rebuilder`, plus `request.upstreamqueue` for requests already sent upstream | `request.upstreamwait`, all `response.*`, and `request.total` (Gateway closes the client connection, so `T10` never fires) |
| Cache hit | Request-path metrics, plus `response.interceptor`, `response.authorization`, `response.send`, and `request.total` | `request.rebuilder`, `request.upstreamqueue`, `request.upstreamwait`, `response.rebuilder` |
| TLS, authentication, or malformed-request failure | None — the request fails before Gateway builds its metrics carrier | All 11 metrics |

Examples of how to use the above table:

* A very low mean for `gateway_request_total_duration_seconds` points to denials, short-circuits, or cache hits, which complete quickly with no Kafka round trip. A cache hit still records `response.send`, but a denial or short-circuit doesn't — so `rate(response.send)` close to `rate(request.total)` means cache hits, while a gap between them counts denials, short-circuits, timeouts, and response-side failures.
* A rising `_max`, or a growing share of requests over 30s (`1 - rate(gateway_request_total_duration_seconds_bucket{le="30"}) / rate(gateway_request_total_duration_seconds_count)`), points to in-flight timeouts waiting for a response from upstream Kafka. Gateway expires these after `GATEWAY_INFLIGHT_REQUEST_EXPIRY_MS` ([default 330s](/guide/conduktor-in-production/deploy-artifacts/deploy-gateway/environment-variables#gateway-internal-timeout)) and `request.total` records the full wait. A few timeouts barely move the mean, so read the tail rather than the average. A broker disconnect looks different — the abandoned requests record no `request.total` and instead show up as a delayed rise in `gateway.request_expired` about 330s later.

## Control cardinality

Each metric carries tags, and every tag combination creates a separate metric instance — the metric's cardinality. For example, `gateway_request_total_duration_seconds` with `api_key=PRODUCE` and with `api_key=FETCH` counts as two instances. The `api_key` tag is bounded (77 client-facing Kafka API keys, around 25 in common use), but the `interceptor` tag isn't, so your Interceptor count drives cardinality the most.

Each instance also generates 12 Prometheus series — nine latency buckets plus `_count`, `_sum`, and `_max` — and every series consumes memory.

The table below shows the cardinality for 25 or 77 API keys and 0, 1, or 10 Interceptors. Each Interceptor is counted on the request path only; one that also runs on the response path counts twice.

| API keys used | Request-path Interceptors | Metric instances | Series count | Approx. scrape body size |
| - | - | - | - | - |
| 25 | 0 | 225 | 2,700 | \~0.3 MB |
| 77 | 0 | 693 | 8,316 | \~1 MB |
| 25 | 1 | 250 | 3,000 | \~0.36 MB |
| 77 | 1 | 770 | 9,240 | \~1.1 MB |
| 25 | 10 | 475 | 5,700 | \~0.68 MB |
| 77 | 10 | 1,463 | 17,556 | \~2.1 MB |

The Interceptor tag drives growth fastest, so to reduce cardinality [scope the metrics](#scope-the-metrics) with `GATEWAY_LIFECYCLE_METRICS_INTERCEPTORS` first, then narrow the API keys with `GATEWAY_LIFECYCLE_METRICS_API_KEYS`.

## Relationship to existing metrics

`gateway_latency_request_response` and `gateway_apiKeys_latency_request_response` measure a narrower window, roughly `T5 -> T9`, than `gateway_request_total_duration_seconds`, which covers the full `T0 -> T10` pipeline. Use the lifecycle metrics for per-stage analysis and these two for the high-level view.

## Related resources

* [Gateway monitoring and alerting recommendations](/guide/conduktor-in-production/monitor/gateway_jmx_recommendations)
* [Gateway metrics reference](/guide/reference/gateway-metrics)
* [Gateway environment variables](/guide/conduktor-in-production/deploy-artifacts/deploy-gateway/environment-variables)
* [Set up monitoring](/guide/conduktor-in-production/monitor)


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.