# How Should EKS Teams Route OpenTelemetry Telemetry by Tenant in 2026?

thane.zone · September 28, 2026

> Direct Answer EKS OpenTelemetry tenant routing is primarily an application-to-collector routing problem, not something that Cilium or Kubernetes...

## Direct Answer

EKS OpenTelemetry tenant routing is primarily an application-to-collector routing problem, not something that Cilium or Kubernetes performs automatically. The reliable pattern is to send each workload’s OTLP traffic through a controlled gateway, authenticate the caller, attach a trusted tenant identity, and then route the signal to a tenant-specific OpenTelemetry Collector pipeline or backend. Kubernetes namespaces, pod labels, and Cilium network policies can enforce boundaries around that path, but they do not by themselves decide which commercial telemetry account receives a span or metric. For multi-team operations, use a shared EKS cluster and shared observability control plane only when tenant identity, quotas, retention, and access controls are explicit; otherwise, separate collector gateways, namespaces, or AWS accounts provide simpler isolation. As of 29 September 2026, teams should treat tenant routing as a zero-trust data-plane design rather than a naming convention.

**Also worth reading:** [How Do Teams Set Up OpenTelemetry Tail Sampling Without Losing Important Traces?](https://thane.zone/knowledge/how_do_teams_set_up_opentelemetry_tail_sampling_without_losing_important_traces.php) · [How Should Engineering Leaders Architect Multi-Tenant Operations Telemetry Ingestion Pipelines in 2026?](https://thane.zone/knowledge/how_should_engineering_leaders_architect_multi-tenant_operations_telemetry_ingestion_pipelines_in_2026.php) · [What Are Agent Telemetry Controls, and How Should B2B Teams Set Them Up in 2026?](https://thane.zone/knowledge/what_are_agent_telemetry_controls_and_how_should_b2b_teams_set_them_up_in_2026.php)

The standard starting architecture uses the AWS Distro for OpenTelemetry Collector, or an upstream OpenTelemetry Collector, as an OTLP gateway. Application SDKs export to that gateway over gRPC or HTTP, the gateway authenticates Kubernetes workload identity or a signed tenant token, and a routing mechanism selects a tenant-specific pipeline. The downstream pipeline can add or validate resource attributes such as tenant.id, export to a dedicated observability vendor account, or feed a shared backend with enforced partitioning. Because OTLP does not define a universally trusted tenant field, accepting X-Tenant-ID directly from an untrusted application is unsafe; any client-provided header must be treated as a claim and checked against workload identity, namespace, service account, or another authoritative source.

## How Routing Actually Works

OpenTelemetry telemetry normally arrives as OTLP traces, metrics, or logs carrying service and resource metadata, but a resource attribute is not an authorization boundary by itself. A collector routing connector, load-balancing exporter, or separate collector deployment can use the authenticated identity to select an exporter and pipeline. With AWS Load Balancer Controller or Kubernetes Gateway API, teams can also separate gateways by tenant or trust class before requests reach a collector. This two-stage model is useful: the edge validates identity and limits traffic, while the collector decides the destination and enriches accepted telemetry. Cilium can restrict which pods and ports may communicate, but it cannot infer the commercial ownership of an OTLP request merely from packet metadata.

A practical request path has four decisions: which workload sent the signal, whether that workload is allowed to send telemetry, which tenant it belongs to, and where the resulting data may be stored. Kubernetes service accounts and IAM roles are often better identity inputs than mutable pod labels because they can be bound to AWS IAM roles through IAM Roles for Service Accounts. A collector can compare an incoming token or authenticated principal with an approved tenant mapping, reject mismatches, and set tenant.id on exported records. For high-volume production systems, define an explicit fallback behavior: fail closed when identity cannot be verified, return 401 or 403 rather than silently routing to a default tenant, and send a small number of internal metrics about rejected requests without echoing customer payloads.

The routing decision should happen before data reaches a shared indexing or storage layer. Sending every namespace to one collector and adding a tenant label afterward may appear cheaper, but it creates a risk that one exporter misconfiguration can mix streams or bypass quotas. It also makes deletion requests harder because the operational team must trust downstream filters for isolation. A separate pipeline for each regulated tenant is more expensive to operate, yet it provides a clearer audit boundary and allows retention, sampling, encryption keys, and regional storage to differ by contract. The correct balance depends on whether teams are merely sharing compute or must demonstrate hard data separation.

## Reference Architecture on Amazon EKS

A common design runs platform collectors in a dedicated namespace, application collectors in tenant namespaces, and a gateway service in front of both. The gateway enforces TLS, authentication, request-size limits, and rate limits; tenant collectors apply organization-specific transformations and exporters. A NetworkPolicy limits the gateway to approved source namespaces, while Cilium can provide additional identity-aware, egress-deny-by-default controls. Cilium’s policy model and observability features are well suited to this network containment role, and public examples from organizations such as The New York Times, Trip.com, and OpenAI show that Cilium has been used for demanding EKS or cloud networking environments. Those examples support Cilium’s suitability for shared-cluster networking, but they do not prove that Cilium handles OpenTelemetry account routing.

The architecture should distinguish three trust levels. Public-facing or low-trust services send only redacted telemetry and receive strict rate limits; ordinary internal services may send richer attributes after workload identity is verified; privileged platform services can use a separate high-throughput gateway with a different sampling policy. Each level should map to a different collector Service or Gateway API listener rather than sharing one overloaded endpoint. This prevents an internal platform component from using a broad route to bypass tenant export restrictions. It also gives SRE leaders measurable indicators such as accepted spans per second, rejected requests, exporter errors, and queue depth for each tenant rather than one aggregate number.

| Feature | Shared collector, shared backend | Dedicated tenant pipeline |
| --- | --- | --- |
| Isolation | Policy, labels, and access controls in one system | Separate exporters, credentials, and often indexes |
| Typical tenant count | Hundreds to low thousands when mappings are stable | Tens, or selected regulated tenants |
| Incremental EKS cost | Lower node and gateway overhead | More collectors, endpoints, and capacity planning |
| Failure containment | A collector can affect many tenants | A pipeline failure is easier to bound to one tenant |
| Retention control | Commonly uniform | Configurable per tenant |
| Operational complexity | Lower initially, higher as mapping grows | Higher initially, clearer ownership later |
| Best fit | Internal SaaS with contractual logical separation | Regulated, high-value, or strict residency tenants |

## Practical Implementation Steps
Begin by inventorying telemetry producers, destinations, and contractual requirements before creating gateways. Record the EKS cluster and region, namespace, Kubernetes service account, telemetry types, expected spans or bytes per second, data residency, retention period, and whether a tenant requires a separate vendor account. A useful pilot contains no more than 2 to 5 representative services and should run through at least 7 days, including normal peak traffic. During that period, measure peak OTLP request rate, compressed payload size, collector CPU, memory, queue depth, and exporter latency. These figures prevent teams from sizing a gateway from namespace count alone, because one service can produce more telemetry than 100 small batch jobs.

Next, establish identity and a tenant registry. Map the Kubernetes service account or another non-forgeable workload identity to a tenant identifier, and require the collector to derive the route from that mapping instead of trusting an arbitrary header. Configure the gateway with mutually authenticated TLS where appropriate, restrict access to OTLP ports 4317 for gRPC and 4318 for HTTP, and avoid exposing the collector directly to the public internet. Add per-tenant rate limits based on measured baselines; an initial ceiling equal to 2 times the observed 95th-percentile traffic rate is often safer than an unlimited endpoint, then adjust after 14 to 30 days. Preserve a small emergency allowance, perhaps 10%, for releases that legitimately increase traffic.

After identity is verified, create a narrow collector pipeline for each trust class or high-risk tenant. The pipeline should remove unapproved attributes, set canonical tenant metadata, select the correct backend credentials, and apply tenant-specific sampling. Traces and metrics can require different policies: retain every authentication event or error trace for one team while sampling successful checkout traces at 5% for another. A sensible default is to retain 100% of errors and critical spans initially, then sample successful high-volume paths between 1% and 10% only after reviewing cost and diagnostic value. Finally, test wrong-tenant access, missing identity, stale mappings, backend outages, and collector restarts; a design that works only when every client behaves correctly is not tenant isolation.

## Cilium, Namespaces, and Gateway Alternatives

Cilium is valuable when the problem is network-level containment, especially in a shared EKS cluster where platform teams want identity-aware policies and egress restrictions. CiliumNetworkPolicy can limit which workloads may reach collector services, and Cilium’s networking stack can make the permitted path visible when it is deployed as the cluster CNI. Cilium also supports network observability features that can help explain connection failures. However, OpenTelemetry tenant routing still requires an application-aware or collector-aware mapping. A packet that reaches a permitted collector endpoint has not yet been assigned to a tenant, and network policy cannot determine whether its payload should enter Tenant A’s vendor account rather than Tenant B’s.

Kubernetes namespaces provide a useful administrative boundary but should not be treated as perfect tenant identity. Names can be reused, service accounts can be misconfigured, and a shared gateway may intentionally accept traffic from many namespaces. Stronger designs use namespace plus service account, AWS IAM role, and a centrally managed mapping, with deployment policy preventing tenants from editing gateway configuration. If tenants can deploy arbitrary pods, platform teams should deny direct egress to public observability endpoints and permit only approved internal routes. This reduces exfiltration and prevents a tenant from bypassing its sampler, redaction policy, or billing account. Cilium can enforce much of that egress boundary, while the collector and cloud IAM settings enforce the data destination.

Other options include separate EKS clusters, separate AWS accounts, a managed observability platform with native workspace isolation, or a message broker between producers and collectors. A message broker, such as an appropriately secured Kafka or MSK deployment, can buffer bursts and decouple ingestion from routing, but it introduces another system to secure, partition, and monitor. Separate clusters offer strong operational isolation but duplicate node capacity, add upgrades, and make cross-cluster fleet management harder. Managed multitenancy products can reduce collector engineering, yet teams must verify whether isolation is logical or whether the provider supplies separate encryption, retention, and regional controls. A hybrid pattern is common: one shared cluster for ordinary teams, dedicated pipelines for regulated customers, and separate accounts for the highest contract requirements.

## Common Mistakes and Failure Modes

The most common error is trusting X-Tenant-ID, tenant.id, or a pod label as if a client had proved ownership of that tenant. Those values are convenient metadata, not authentication. A compromised SDK, sidecar, or pod can change them and send data to another route unless the gateway validates them against a protected mapping. The second error is using one default exporter for unknown identities, because this turns configuration errors into cross-tenant leakage. The third is exposing OTLP port 4317 or 4318 through a public load balancer without authentication, rate limiting, or request-size controls. The fourth is applying sampling before authentication and later assuming rejected traffic has no cost, since parsing and buffering still consume compute and network.

Teams also underestimate redaction and attribute cardinality. Tenant IDs should be low-cardinality and controlled, while request bodies, user emails, full URLs, and arbitrary exception strings may contain customer data or create expensive backend indexes. A 200-byte span can expand substantially after indexing, and 1 million spans per second at even a modest indexed cost can become a material cloud bill. A collector does not automatically remove sensitive fields merely because OpenTelemetry supports attributes. Use processors or an approved gateway policy to remove prohibited fields, cap attribute and tag lengths, and test that errors do not serialize request bodies. Keep telemetry routing configuration in version control, review changes as production changes, and maintain a record of which identity can select which exporter.

Operational failures need explicit limits and alerts. Set queue and retry policies so a failing tenant backend does not cause unbounded memory growth in a shared collector. Use separate concurrency limits and circuit breakers for each exporter, and route only low-cardinality operational metrics to the central monitoring account. Alert when rejection rates exceed 1% over 15 minutes, when queue depth remains above the tested capacity for 5 minutes, or when exporter errors exceed 5% during a sustained period; these are starting thresholds, not universal standards. A tenant-routing test should include an intentionally wrong identity, an expired credential, a deleted service account, and a backend outage. Verify that failures are visible to platform operators without exposing the rejected customer’s payload.

## Cost, Timing, and When to Act

OpenTelemetry software is generally free to use, but the bill is driven by the EKS nodes, gateways, storage, network transfer, and observability vendor’s ingestion or query pricing. A small internal pilot may use 2 to 4 modest gateway replicas across 2 availability zones, but production sizing should follow measured OTLP load and the collector’s tail-latency behavior. Kubernetes cluster management charges, AWS Load Balancer Controller usage, NAT gateway traffic, and cross-region replication can add recurring cost. A dedicated pipeline for 100 tenants does not necessarily require 100 sets of nodes; it may require 100 exporter configurations and credential sets, but high-volume tenants can still need separate capacity. Calculate cost using peak sustained load and failure headroom rather than average traffic.

Many teams should act when a second or third business unit needs a different billing, retention, or regional destination, especially if the current setup uses one shared endpoint. Acting earlier is justified when contracts require auditable separation, when tenants can deploy code independently, or when a single incident could expose another customer’s telemetry. Waiting is reasonable for a single internal product using one backend, one retention period, and one team that controls every collector configuration. In that case, excessive gateway tiers may add operational burden without reducing meaningful risk. A practical review point is every 90 days, with an immediate review after a new region, acquisition, regulated-customer onboarding, or material change in telemetry volume.

The decision can be expressed as a control-strength threshold rather than a tenant-count rule. If isolation means only separate labels in one index, a shared collector is usually adequate for a small internal deployment. If isolation means separate credentials, deletion guarantees, regional storage, and incident containment, create dedicated pipelines or accounts before onboarding the tenant. If tenants are actively adversarial, use separate EKS clusters or accounts when network and IAM policy cannot provide a sufficient boundary. This is why the final architecture should be documented in terms of failure consequences: a mistaken metric label is an inconvenience, while a cross-tenant trace containing regulated data is a contractual and security event. Make the stronger design the default only where the consequence warrants it.

## Recommended Decision Standard

For leadership teams operating several business functions on EKS, the recommended standard is a shared cluster with a centrally operated, authenticated OpenTelemetry gateway, per-tenant route mappings, and dedicated pipelines for customers whose contracts require stronger boundaries. Begin with one gateway per trust class, not one gateway per namespace. Use IAM Roles for Service Accounts, TLS, Cilium or Kubernetes network policy, collector-side authorization, and an approved mapping registry. Keep ordinary tenants in a shared backend only when the observability platform demonstrably enforces row, workspace, or key-level isolation and the team accepts the residual risk. Revisit the architecture at 50, 200, and 1,000 active tenant routes because credential and policy maintenance costs tend to grow faster than raw traffic.

Measure success through four outcomes: fewer than 1% unauthorized-route attempts, no confirmed cross-tenant payload exposure, bounded exporter queues during backend failure, and predictable cost per team. Review these at weekly operational meetings and monthly leadership meetings rather than waiting for an audit. The key phrase “tenant routing” should refer to a documented decision process, not merely a header added in code. That process identifies the tenant, verifies the sender, selects the destination, applies retention and sampling, and records an auditable result. If any of those steps depend on manual trust, the system is not ready for independent multi-team operation. This standard supports shared infrastructure while keeping customer boundaries visible, testable, and proportionate to the risk.

## Quick answers

### Does Cilium automatically route OpenTelemetry data to the correct tenant?

No. Cilium can restrict network paths, provide identity-aware policy, and expose connection observability, but it does not infer the commercial tenant from an OTLP payload. The collector or an application-aware gateway must authenticate the workload and choose the exporter or backend route.

### Can an X-Tenant-ID header route telemetry on Amazon EKS?

It can be used as a routing input only after the gateway verifies the caller and checks the header against an authoritative tenant mapping. Accepting the header directly from an untrusted SDK or pod allows one tenant to claim another tenant’s identity.

### Should every Kubernetes tenant use a separate OpenTelemetry Collector?

Not usually. Shared collectors are practical for ordinary internal teams when identity, quotas, sampling, and backend access are centrally enforced. Dedicated collectors, pipelines, vendor accounts, or clusters are more appropriate for regulated tenants, strict residency, or hard failure-containment requirements.

### What is the safest default when a workload has no verified tenant identity?

Reject or quarantine the telemetry rather than sending it to a default tenant. A fail-closed policy avoids cross-tenant leakage, while operators can inspect low-cardinality rejection metrics and the workload identity that caused the failure.

### How much does EKS OpenTelemetry tenant routing cost?

OpenTelemetry and its collector are generally free, but EKS nodes, load balancers, NAT or inter-region transfer, storage, and observability ingestion dominate cost. Actual expense depends on peak spans or metrics per second, retention, sampling, and whether pipelines share infrastructure.

Canonical: https://thane.zone/knowledge/how_should_eks_teams_route_opentelemetry_telemetry_by_tenant_in_2026.php
Markdown: https://thane.zone/knowledge/how_should_eks_teams_route_opentelemetry_telemetry_by_tenant_in_2026.php/index.md
