Skip to main content

Ship Kubernetes Logs to ClusterNest Managed OpenSearch Without Static Credentials

Pods authenticate to OpenSearch with their own ServiceAccount token. There is no password to distribute or rotate: kubelet mints and refreshes the token, and OpenSearch validates it against an OIDC issuer. This guide uses Vector as the log forwarder, with a small sidecar that keeps its configuration in step with the rotating token.

There are four parts. Parts 1 to 3 are three ways to ship logs without static credentials, and Part 4 covers how people read them. Each part is complete on its own, so you can follow any of them without the others:

  • Part 1 trusts one Kubernetes cluster's issuer directly.
  • Part 2 puts Dex in front, so several clusters can ship logs to one OpenSearch cluster, each limited to its own index pattern.
  • Part 3 gives every cluster its own Dex and its own auth source, so one cluster going down does not affect the others.
  • Part 4 lets people read the logs by signing in to Dashboards with SAML, with access limited to each cluster's index pattern.

Why Vector and not Fluent Bit​

The forwarder has to send the ServiceAccount token (or a token derived from it) to OpenSearch as Authorization: Bearer <token>. Fluent Bit is the usual choice for Kubernetes logs, but it cannot do this with its OpenSearch output:

  • Fluent Bit 5.1.3's opensearch output has no option to send a custom header or a Bearer token. Its authentication options are http_user, http_passwd and the AWS ones, and an unknown property such as header stops the output from initialising.
  • Fluent Bit's generic http output does accept headers, but it posts plain records. OpenSearch's bulk API needs an action line before every document, so you would have to build that yourself, for example with a Lua filter.

Vector's elasticsearch sink takes arbitrary request headers and speaks the bulk API natively, so it can send the token directly. It also reloads its own config when the file changes, which is how the guide keeps up with a token that rotates every few minutes.

Part 1: One cluster, direct​

How it works​

  1. Your Kubernetes cluster publishes an OIDC issuer (ServiceAccount issuer discovery).
  2. The OpenSearch cluster trusts that issuer through an oidc auth source.
  3. A pod projects a ServiceAccount token and sends it as Authorization: Bearer <token>.
  4. OpenSearch takes the token's sub claim (system:serviceaccount:<namespace>:<name>) as the username. A role mapping on that username decides what the pod may write.

Requirements​

  • The issuer URL is served over HTTPS and its /.well-known/openid-configuration and JWKS (/openid/v1/jwks) are reachable from the internet. ClusterNest rejects an auth source whose host is not publicly resolvable.

Find the issuer of your cluster by decoding any ServiceAccount token and reading its iss claim:

kubectl create token default --duration=10m \
| cut -d. -f2 | tr '_-' '/+' \
| awk '{while (length($0) % 4) $0 = $0 "="; print}' | base64 --decode | jq -r .iss

1. Create the cluster with an OIDC auth source​

Set audience to the audience your tokens are issued for (the --audience you pass when minting a token below). client_id is optional: it is only used for Dashboards login, so leave it out. subject_key must be sub.

  1. Open the ClusterNest Console, go to OpenSearch under Services and click Launch Cluster.
  2. Enter a name and pick a tier.
  3. Under Authentication, click Add OIDC source and fill in:
    • Name: k8s
    • Connect URL: https://oidc.example.com/.well-known/openid-configuration
    • Audience: https://oidc.example.com
    • Subject Key: sub
    • Leave Offer on OpenSearch Dashboards off.
  4. Click Create Cluster. The Console shows the admin credentials once, so save them.
  5. Wait until the cluster's state is available.

The API endpoint is on the cluster page.

2. Create a role and map the ServiceAccount​

ClusterNest does not create role mappings. Create them with the cluster's admin credential. Replace logging and vector below with the namespace and ServiceAccount your forwarder runs as.

  1. Sign in to OpenSearch Dashboards as admin, with the credential the Console showed when it created the cluster.
  2. Open Security > Roles and click Create role. Name it logs_writer.
  3. Under Cluster permissions, add indices:data/write/bulk.
  4. Under Index permissions, set the index pattern to logs-* and the permissions to create_index, indices:data/write/bulk* and indices:data/write/index. Click Create.
  5. Open the role's Mapped users tab and click Manage mapping.
  6. Under Users, add system:serviceaccount:logging:vector and click Map.

users accepts wildcards, so system:serviceaccount:logging:* covers every ServiceAccount in a namespace. To add a ServiceAccount later, open Manage mapping again and add another user.

note

With audience set, OpenSearch rejects a token whose aud claim does not match, even if the trusted issuer signed it. Without it, a token for any audience from that issuer is accepted and access is decided by the sub mapping alone, so keep the mapping as narrow as the workloads that need it.

3. Check authentication​

Mint a token for the ServiceAccount and call the cluster. A ServiceAccount with no mapping authenticates but is denied, which proves the issuer is trusted and the sub becomes the username:

kubectl create token default -n default --audience https://oidc.example.com --duration=10m \
| sed 's/^/Authorization: Bearer /' \
| curl -H @- "$OS_URL/_cluster/health"
{"error":{"root_cause":[{"type":"security_exception","reason":"no permissions for [cluster:monitor/health] and User [name=system:serviceaccount:default:default, backend_roles=[], requestedTenant=null]"}],"type":"security_exception","reason":"no permissions for [cluster:monitor/health] and User [name=system:serviceaccount:default:default, backend_roles=[], requestedTenant=null]"},"status":403}

A 401 means OpenSearch could not validate the token, usually because the issuer is not reachable from the internet.

4. Deploy Vector​

Vector reads header values when it starts and does not re-read a file. The sidecar renders the Vector config with the current token and rewrites it whenever the token changes. Vector watches the config file and reloads itself.

The pod needs a projected token, a ServiceAccount, and read access to pod metadata:

apiVersion: v1
kind: ServiceAccount
metadata:
name: vector
namespace: logging
---
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRole
metadata:
name: vector-logging
rules:
- apiGroups: [""]
resources: ["pods", "namespaces", "nodes"]
verbs: ["get", "list", "watch"]
---
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRoleBinding
metadata:
name: vector-logging
roleRef:
apiGroup: rbac.authorization.k8s.io
kind: ClusterRole
name: vector-logging
subjects:
- kind: ServiceAccount
name: vector
namespace: logging

The config template and renderer script. @@TOKEN@@ is replaced with the token on every render:

apiVersion: v1
kind: ConfigMap
metadata:
name: vector
namespace: logging
data:
vector.yaml.tpl: |
data_dir: /vector-data
sources:
logs:
type: kubernetes_logs
extra_field_selector: metadata.namespace=logging
exclude_paths_glob_patterns:
- "/var/log/pods/logging_vector-*/**"
transforms:
limit:
type: throttle
inputs: [logs]
threshold: 50
window_secs: 1
sinks:
opensearch:
type: elasticsearch
inputs: [limit]
endpoints:
- https://logs-123-os.c9t.io
api_version: v7
mode: bulk
bulk:
index: "logs-k8s-%Y.%m.%d"
healthcheck:
enabled: false
request:
headers:
Authorization: "Bearer @@TOKEN@@"
buffer:
type: memory
max_events: 500
render.sh: |
#!/bin/sh
tpl=/tpl/vector.yaml.tpl
tok=/var/run/secrets/opensearch/token
out=/conf/vector.yaml
render() {
sed "s|@@TOKEN@@|$(cat "$tok")|" "$tpl" > "$out.tmp" && mv "$out.tmp" "$out"
}
render
[ "$1" = once ] && exit 0
last=$(cksum < "$tok")
while true; do
sleep 5
cur=$(cksum < "$tok")
if [ "$cur" != "$last" ]; then
render && last=$cur
fi
done

The exclude_paths_glob_patterns entry keeps Vector from shipping its own logs, which would otherwise feed its own errors back into the pipeline. The sink's health check is disabled because the logs_writer role cannot read cluster information, and api_version is set explicitly for the same reason: Vector would otherwise probe the cluster root. To collect more than one namespace, widen extra_field_selector and the role mapping together.

The DaemonSet:

apiVersion: apps/v1
kind: DaemonSet
metadata:
name: vector
namespace: logging
spec:
selector:
matchLabels:
app: vector
template:
metadata:
labels:
app: vector
spec:
serviceAccountName: vector
initContainers:
- name: render-config
image: busybox:1.36
command: ["/bin/sh", "/tpl/render.sh", "once"]
volumeMounts:
- {name: tpl, mountPath: /tpl}
- {name: conf, mountPath: /conf}
- {name: token, mountPath: /var/run/secrets/opensearch, readOnly: true}
containers:
- name: vector
image: timberio/vector:0.58.0-alpine
args:
- --config=/conf/vector.yaml
- --watch-config
- --watch-config-method=poll
- --watch-config-poll-interval-seconds=5
env:
- name: VECTOR_SELF_NODE_NAME
valueFrom:
fieldRef: {fieldPath: spec.nodeName}
resources:
requests: {cpu: 20m, memory: 48Mi}
limits: {cpu: 150m, memory: 128Mi}
volumeMounts:
- {name: conf, mountPath: /conf, readOnly: true}
- {name: data, mountPath: /vector-data}
- {name: varlog, mountPath: /var/log, readOnly: true}
- name: token-renderer
image: busybox:1.36
command: ["/bin/sh", "/tpl/render.sh"]
volumeMounts:
- {name: tpl, mountPath: /tpl}
- {name: conf, mountPath: /conf}
- {name: token, mountPath: /var/run/secrets/opensearch, readOnly: true}
volumes:
- name: tpl
configMap: {name: vector, defaultMode: 0555}
- name: conf
emptyDir: {medium: Memory}
- name: data
emptyDir: {}
- name: varlog
hostPath: {path: /var/log}
- name: token
projected:
sources:
- serviceAccountToken:
path: token
audience: https://oidc.example.com
expirationSeconds: 600

The config lives on a memory-backed emptyDir, so the token is never written to disk on the node. Set audience to your issuer URL.

Token rotation​

The projected token is valid for expirationSeconds (here 10 minutes, the minimum). Kubelet replaces the file at about 80% of that lifetime. The token-renderer sidecar notices the new content within 5 seconds and rewrites the Vector config, and Vector reloads without dropping events. Each pod logs a reload once at startup and once per rotation:

kubectl -n logging logs ds/vector -c vector | grep -E 'Configuration file changed|has reloaded|ERROR'
INFO vector::config::watcher: Configuration file changed.
INFO vector: Vector has reloaded. path=[File("/conf/vector.yaml", None)]

With a log line written every 5 seconds, delivery stayed at a steady 6 documents per 30 seconds across the rotation of all three pods, with no ERROR or 401 lines.

Pods restarted by a rollout start with a new token and rotate about 8 minutes later, so reloads in the first minutes after a rollout are only the startup render.

5. Expire old indices​

A daily index is created every day and never removed on its own. An index state management (ISM) policy deletes each index once it is old enough. This one deletes after 7 days. ism_template attaches it automatically to every new logs-k8s-* index:

  1. In OpenSearch Dashboards, open Management > Index Management > State management policies.

  2. Click Create policy and choose the JSON editor.

  3. Set the policy ID to logs-retention and paste:

    {
    "policy": {
    "description": "Delete daily log indices after 7 days",
    "default_state": "hot",
    "states": [
    {"name": "hot", "actions": [], "transitions": [{"state_name": "delete", "conditions": {"min_index_age": "7d"}}]},
    {"name": "delete", "actions": [{"delete": {}}], "transitions": []}
    ],
    "ism_template": [{"index_patterns": ["logs-k8s-*"], "priority": 100}]
    }
    }
  4. Click Create.

  5. To attach the policy to indices that already exist, open Indices, select them, and choose Actions > Apply policy. The template only applies to indices created after the policy exists.

  6. Open Managed indices to check that each index shows logs-retention.

Change min_index_age to keep logs for longer or shorter.

6. Verify delivery​

Once a pod writes a log line, the document appears in the day's index, logs-k8s-YYYY.MM.DD:

curl -u admin:$ADMIN_PASSWORD "$OS_URL/logs-k8s-*/_search?size=1&sort=timestamp:desc"

Each document carries the message and Kubernetes metadata (kubernetes.pod_namespace, kubernetes.pod_name, container and node details). In OpenSearch Dashboards, create an index pattern logs-k8s-* with timestamp as the time field. The role from step 2 allows every logs-* index, so no permission change is needed.

7. Clean up​

Remove Vector from the Kubernetes cluster. Deleting the namespace removes the DaemonSet, ConfigMap and ServiceAccount, and the ClusterRole and ClusterRoleBinding are cluster-wide, so they go separately. Only delete the namespace if you created it for this guide:

kubectl delete clusterrolebinding/vector-logging clusterrole/vector-logging
kubectl delete namespace logging

Then delete the OpenSearch cluster. This also deletes its indices, the role and mapping, and the retention policy:

  1. Open the cluster in the ClusterNest Console.
  2. Click Delete.
  3. Type the cluster name to confirm.
  4. Click Delete Cluster.

The cluster shows "state": "deleting" until it is gone.

Part 2: Several clusters through Dex​

With several Kubernetes clusters, trusting each cluster's issuer directly (Part 1) is not enough. The sub claim of a ServiceAccount token is system:serviceaccount:<namespace>:<name> in every cluster, so logging/vector in prod-eu is indistinguishable from logging/vector in staging. A role mapping cannot require both a cluster and a namespace.

Dex fixes this. It trusts each cluster's issuer through its own connector, exchanges a ServiceAccount token for a Dex token, and prefixes a groups claim with the connector's id. OpenSearch trusts only Dex and maps that group to a role limited to the cluster's own index pattern. There is still no static credential: the only secret in the flow is the pod's own short-lived ServiceAccount token.

This part is complete on its own. You do not need Part 1. One Dex serves any number of clusters, and you can run more than one Dex, each added to OpenSearch as its own auth source.

It uses three example clusters:

ClusterConnector idGroupIndex pattern
prod-euprod-euprod-eu:system:serviceaccount:logging:vectorlogs-k8s-prod-eu-*
prod-usprod-usprod-us:system:serviceaccount:logging:vectorlogs-k8s-prod-us-*
stagingstagingstaging:system:serviceaccount:logging:vectorlogs-k8s-staging-*

How it works​

  1. A pod projects a ServiceAccount token with audience dex.
  2. A sidecar posts it to Dex with the connector id of the pod's cluster (prod-eu).
  3. Dex verifies the token against that cluster's issuer and returns an ID token whose groups claim is prod-eu:system:serviceaccount:logging:vector.
  4. The forwarder sends the Dex token to OpenSearch, which takes the group as a backend role.
  5. The role mapped to that group can only write logs-k8s-prod-eu-*.

Requirements​

  • Dex is served over HTTPS on a public hostname (here https://dex.example.com). ClusterNest rejects an auth source whose host is not publicly resolvable.
  • Dex can reach each cluster's issuer over HTTPS: the issuer's /.well-known/openid-configuration and its JWKS (/openid/v1/jwks).
  • The OpenSearch cluster can reach Dex.

Find a cluster's issuer by decoding any of its ServiceAccount tokens and reading the iss claim:

kubectl create token default --duration=10m \
| cut -d. -f2 | tr '_-' '/+' \
| awk '{while (length($0) % 4) $0 = $0 "="; print}' | base64 --decode | jq -r .iss

1. Configure Dex​

One public client is shared by every shipper, and one connector per cluster verifies that cluster's tokens. The client is public: true, so the exchange needs no client secret.

issuer: https://dex.example.com
storage:
type: memory
web:
http: 0.0.0.0:5556
oauth2:
grantTypes:
- urn:ietf:params:oauth:grant-type:token-exchange
skipApprovalScreen: true
expiry:
idTokens: 15m
staticClients:
- id: logs-shipper
name: Log shippers
public: true
redirectURIs:
- http://localhost/unused
connectors:
- type: oidc
id: prod-eu
name: prod-eu
config:
issuer: https://oidc.prod-eu.example.com
clientID: dex
clientSecret: unused
redirectURI: https://dex.example.com/callback
scopes:
- openid
userNameKey: sub
insecureSkipEmailVerified: true
claimModifications:
newGroupFromClaims:
- claims:
- sub
delimiter: ":"
prefix: "prod-eu"
- type: oidc
id: prod-us
name: prod-us
config:
issuer: https://oidc.prod-us.example.com
clientID: dex
clientSecret: unused
redirectURI: https://dex.example.com/callback
scopes:
- openid
userNameKey: sub
insecureSkipEmailVerified: true
claimModifications:
newGroupFromClaims:
- claims:
- sub
delimiter: ":"
prefix: "prod-us"
- type: oidc
id: staging
name: staging
config:
issuer: https://oidc.staging.example.com
clientID: dex
clientSecret: unused
redirectURI: https://dex.example.com/callback
scopes:
- openid
userNameKey: sub
insecureSkipEmailVerified: true
claimModifications:
newGroupFromClaims:
- claims:
- sub
delimiter: ":"
prefix: "staging"
  • clientID: dex on a connector is the audience Dex requires on the incoming ServiceAccount token. Project tokens with audience: dex.
  • userNameKey: sub makes the Dex name claim the ServiceAccount identity. scopes: [openid] with insecureSkipEmailVerified: true is needed because ServiceAccount tokens carry no email. The connector's clientSecret is a required field that is never used.
  • newGroupFromClaims builds the group from sub. With prefix: "prod-eu" and delimiter: ":" the group reads prod-eu:system:serviceaccount:logging:vector. A prefix that already ends in a colon produces a doubled colon.
  • The exchange request must carry client_id=logs-shipper. A request that names no client is rejected.
warning

With storage: type: memory, Dex generates new signing keys every time it restarts. OpenSearch re-fetches Dex's keys only at a limited rate, so for a few minutes after a restart it rejects tokens signed by the new keys with 401 Authentication finally failed, and Vector drops the events it could not send in that window. In testing this lasted about 4 minutes after a restart that followed several earlier restarts. Use a persistent Dex storage backend so the keys survive restarts.

2. Deploy Dex​

Save the config from step 1 as config.yaml, then run Dex in a cluster that can publish HTTPS services. It does not have to be one of the clusters that ship logs.

kubectl create namespace dex
kubectl -n dex create configmap dex --from-file=config.yaml

A Deployment and a Service. The readiness probe uses the discovery document:

apiVersion: apps/v1
kind: Deployment
metadata:
name: dex
namespace: dex
spec:
replicas: 1
selector:
matchLabels:
app: dex
template:
metadata:
labels:
app: dex
spec:
containers:
- name: dex
image: ghcr.io/dexidp/dex:v2.45.1
args: ["dex", "serve", "/etc/dex/config.yaml"]
ports:
- {name: http, containerPort: 5556}
readinessProbe:
httpGet: {path: /.well-known/openid-configuration, port: 5556}
resources:
requests: {cpu: 10m, memory: 32Mi}
limits: {cpu: 100m, memory: 96Mi}
volumeMounts:
- {name: config, mountPath: /etc/dex}
volumes:
- name: config
configMap: {name: dex}
---
apiVersion: v1
kind: Service
metadata:
name: dex
namespace: dex
spec:
selector:
app: dex
ports:
- {name: http, port: 5556, targetPort: 5556}

Publish the Service at the issuer hostname over HTTPS, with a certificate that matches it. This example uses a Gateway API HTTPRoute on an existing Gateway with an https listener. If you use an Ingress instead, route dex.example.com to the dex Service on port 5556:

apiVersion: gateway.networking.k8s.io/v1
kind: HTTPRoute
metadata:
name: dex
namespace: dex
spec:
hostnames:
- dex.example.com
parentRefs:
- group: gateway.networking.k8s.io
kind: Gateway
name: gateway
namespace: gateway-system
sectionName: https
rules:
- backendRefs:
- {group: "", kind: Service, name: dex, port: 5556}
matches:
- path: {type: PathPrefix, value: /}

Check that Dex answers from the internet:

curl https://dex.example.com/.well-known/openid-configuration

The response lists "issuer": "https://dex.example.com" and a jwks_uri of https://dex.example.com/keys. A 503 no healthy upstream just after the deploy means the pod is not Ready yet. It clears within seconds once the readiness probe passes. To change the config later, update the ConfigMap and restart the Deployment. Read the storage warning in step 1 before restarting Dex.

3. Create the OpenSearch cluster with Dex as the auth source​

subject_key is name and roles_key is groups. client_id and audience are the Dex client. With audience set, OpenSearch rejects a token Dex issued for any other client:

  1. Open the ClusterNest Console, go to OpenSearch under Services and click Launch Cluster.
  2. Enter a name and pick a tier.
  3. Under Authentication, click Add OIDC source and fill in:
    • Name: dex
    • Connect URL: https://dex.example.com/.well-known/openid-configuration
    • Client ID: logs-shipper
    • Audience: logs-shipper
    • Subject Key: name
    • Roles Key: groups
    • Leave Offer on OpenSearch Dashboards off.
  4. Click Create Cluster. The Console shows the admin credentials once, so save them.
  5. Wait until the cluster's state is available.

4. Create one role per cluster​

Each role can write only its cluster's indices. Use the OpenSearch security API with the cluster's admin credential, and map each role by backend role, which is the group Dex emits.

Do not map users for these ServiceAccounts. The username (system:serviceaccount:logging:vector) is the same in every cluster, so a username mapping on any role grants that role's access to all of them.

Repeat steps 2 to 7 for each cluster (prod-eu, prod-us, staging), replacing <cluster>.

  1. Sign in to OpenSearch Dashboards as admin, with the credential the Console showed when it created the cluster.
  2. Open Security > Roles and click Create role. Name it logs_writer_<cluster>.
  3. Under Cluster permissions, add indices:data/write/bulk.
  4. Under Index permissions, set the index pattern to logs-k8s-<cluster>-* and the permissions to create_index, indices:data/write/bulk* and indices:data/write/index. Click Create.
  5. Open the role's Mapped users tab and click Manage mapping.
  6. Under Backend roles, add <cluster>:system:serviceaccount:logging:vector.
  7. Click Map.

To allow another ServiceAccount from one cluster, add its group (prod-eu:system:serviceaccount:logging:other) as another backend role on that cluster's mapping.

5. Check the exchange​

Mint a token for the ServiceAccount, exchange it, and read the claims of the result:

SA_TOKEN=$(kubectl create token vector -n logging --audience dex --duration=10m)

curl https://dex.example.com/token \
-d client_id=logs-shipper \
-d grant_type=urn:ietf:params:oauth:grant-type:token-exchange \
-d connector_id=prod-eu \
--data-urlencode "subject_token=$SA_TOKEN" \
-d subject_token_type=urn:ietf:params:oauth:token-type:id_token \
-d requested_token_type=urn:ietf:params:oauth:token-type:id_token \
--data-urlencode "scope=openid profile groups"

The access_token field of the response holds a JWT. Its claims include:

{
"iss": "https://dex.example.com",
"aud": "logs-shipper",
"name": "system:serviceaccount:logging:vector",
"groups": ["prod-eu:system:serviceaccount:logging:vector"]
}

requested_token_type must be id_token to get a JWT that OpenSearch can validate. The token expires after expiry.idTokens.

Use it as a Bearer token. The bulk API answers 200 and reports each write in the response items. A write to the cluster's own index succeeds, a write to another cluster's index is denied, and reads are denied because the role is write-only:

# own index: the item status is 201
curl -X POST "$OS_URL/_bulk" -H "Authorization: Bearer $DEX_TOKEN" \
-H "Content-Type: application/x-ndjson" \
--data-binary $'{"index":{"_index":"logs-k8s-prod-eu-2026.10.04"}}\n{"message":"hello"}\n'

The same token writing to another cluster's index, for example logs-k8s-prod-us-2026.10.04, is denied: the request returns 200, and the item in the response carries "status":403 with a security_exception.

A token that Dex issued for a different client, so with another aud, is rejected with 401 before any role is checked.

6. Deploy Vector in each cluster​

Vector reads header values only when it starts, so a sidecar keeps its config in step with the token. It exchanges the pod's ServiceAccount token at Dex every 5 minutes (Dex tokens last 15), renders the Dex token into the Vector config, and Vector reloads when the file changes. The examples are for prod-eu. For another cluster, change the connector id, the index prefix and nothing else.

The ServiceAccount and read access to pod metadata:

apiVersion: v1
kind: Namespace
metadata:
name: logging
---
apiVersion: v1
kind: ServiceAccount
metadata:
name: vector
namespace: logging
---
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRole
metadata:
name: vector-logging
rules:
- apiGroups: [""]
resources: ["pods", "namespaces", "nodes"]
verbs: ["get", "list", "watch"]
---
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRoleBinding
metadata:
name: vector-logging
roleRef:
apiGroup: rbac.authorization.k8s.io
kind: ClusterRole
name: vector-logging
subjects:
- kind: ServiceAccount
name: vector
namespace: logging

The Vector config template and the exchange script. @@TOKEN@@ is replaced with the Dex token on every render:

apiVersion: v1
kind: ConfigMap
metadata:
name: vector
namespace: logging
data:
vector.yaml.tpl: |
data_dir: /vector-data
sources:
logs:
type: kubernetes_logs
extra_field_selector: metadata.namespace=logging
exclude_paths_glob_patterns:
- "/var/log/pods/logging_vector-*/**"
transforms:
limit:
type: throttle
inputs: [logs]
threshold: 50
window_secs: 1
sinks:
opensearch:
type: elasticsearch
inputs: [limit]
endpoints:
- https://logs-123-os.c9t.io
api_version: v7
mode: bulk
bulk:
index: "logs-k8s-prod-eu-%Y.%m.%d"
healthcheck:
enabled: false
request:
headers:
Authorization: "Bearer @@TOKEN@@"
buffer:
type: memory
max_events: 500
render.sh: |
#!/bin/sh
tpl=/tpl/vector.yaml.tpl
sa=/var/run/secrets/dex-exchange/token
out=/conf/vector.yaml
exchange() {
curl -fsS -m 20 https://dex.example.com/token \
--data-urlencode client_id=logs-shipper \
--data-urlencode grant_type=urn:ietf:params:oauth:grant-type:token-exchange \
--data-urlencode connector_id=prod-eu \
--data-urlencode "subject_token=$(cat "$sa")" \
--data-urlencode subject_token_type=urn:ietf:params:oauth:token-type:id_token \
--data-urlencode requested_token_type=urn:ietf:params:oauth:token-type:id_token \
--data-urlencode "scope=openid profile groups" \
| sed -n 's/.*"access_token":"\([^"]*\)".*/\1/p'
}
render() {
t=$(exchange)
[ -n "$t" ] || return 1
sed "s|@@TOKEN@@|$t|" "$tpl" > "$out.tmp" && mv "$out.tmp" "$out"
}
if [ "$1" = once ]; then render; exit $?; fi
while true; do
sleep 300
render || echo "token exchange failed, keeping previous token"
done
  • The exclude_paths_glob_patterns entry keeps Vector from shipping its own logs, which would otherwise feed its own errors back into the pipeline.
  • The healthcheck is off because the per-cluster role cannot read cluster information. api_version is set explicitly for the same reason: Vector would otherwise probe the cluster root.
  • To collect more than one namespace, widen extra_field_selector and add the namespace's ServiceAccounts to the role mapping.

The DaemonSet. The init container renders the first config so Vector never starts without one. If an exchange fails later, the sidecar keeps the previous token and retries after the next interval.

apiVersion: apps/v1
kind: DaemonSet
metadata:
name: vector
namespace: logging
spec:
selector:
matchLabels:
app: vector
template:
metadata:
labels:
app: vector
spec:
serviceAccountName: vector
initContainers:
- name: render-config
image: curlimages/curl:8.10.1
command: ["/bin/sh", "/tpl/render.sh", "once"]
volumeMounts:
- {name: tpl, mountPath: /tpl}
- {name: conf, mountPath: /conf}
- {name: sa-token, mountPath: /var/run/secrets/dex-exchange, readOnly: true}
containers:
- name: vector
image: timberio/vector:0.58.0-alpine
args:
- --config=/conf/vector.yaml
- --watch-config
- --watch-config-method=poll
- --watch-config-poll-interval-seconds=5
env:
- name: VECTOR_SELF_NODE_NAME
valueFrom:
fieldRef: {fieldPath: spec.nodeName}
resources:
requests: {cpu: 20m, memory: 48Mi}
limits: {cpu: 150m, memory: 128Mi}
volumeMounts:
- {name: conf, mountPath: /conf, readOnly: true}
- {name: data, mountPath: /vector-data}
- {name: varlog, mountPath: /var/log, readOnly: true}
- name: token-exchanger
image: curlimages/curl:8.10.1
command: ["/bin/sh", "/tpl/render.sh"]
volumeMounts:
- {name: tpl, mountPath: /tpl}
- {name: conf, mountPath: /conf}
- {name: sa-token, mountPath: /var/run/secrets/dex-exchange, readOnly: true}
volumes:
- name: tpl
configMap: {name: vector, defaultMode: 0555}
- name: conf
emptyDir: {medium: Memory}
- name: data
emptyDir: {}
- name: varlog
hostPath: {path: /var/log}
- name: sa-token
projected:
sources:
- serviceAccountToken:
path: token
audience: dex
expirationSeconds: 600

The rendered config lives on a memory-backed emptyDir, so the Dex token is never written to the node's disk. Check that the sidecar keeps working. Each pod logs one reload at startup and one about every 5 minutes:

kubectl -n logging logs ds/vector -c vector | grep -E 'Configuration file changed|ERROR'
kubectl -n logging logs ds/vector -c token-exchanger

7. Expire old indices​

Every cluster writes a new daily index. An index state management (ISM) policy deletes each index once it is old enough. This one deletes after 7 days, and ism_template attaches it automatically to every new logs-*-* index:

  1. In OpenSearch Dashboards, open Management > Index Management > State management policies.

  2. Click Create policy and choose the JSON editor.

  3. Set the policy ID to logs-retention and paste:

    {
    "policy": {
    "description": "Delete daily log indices after 7 days",
    "default_state": "hot",
    "states": [
    {"name": "hot", "actions": [], "transitions": [{"state_name": "delete", "conditions": {"min_index_age": "7d"}}]},
    {"name": "delete", "actions": [{"delete": {}}], "transitions": []}
    ],
    "ism_template": [{"index_patterns": ["logs-*-*"], "priority": 100}]
    }
    }
  4. Click Create.

  5. To attach the policy to indices that already exist, open Indices, select them, and choose Actions > Apply policy. The template only applies to indices created after the policy exists.

  6. Open Managed indices to check that each index shows logs-retention.

Change min_index_age to keep logs for longer or shorter.

8. Verify delivery​

Once a pod writes a log line, the document appears in the cluster's daily index. Each cluster's documents live only in its own pattern:

curl -u admin:$ADMIN_PASSWORD "$OS_URL/logs-k8s-prod-eu-*/_search?size=1&sort=timestamp:desc"

Each document carries the message and Kubernetes metadata (kubernetes.pod_namespace, kubernetes.pod_name, container and node details). In OpenSearch Dashboards, create one index pattern per cluster, for example logs-k8s-prod-eu-*, with timestamp as the time field.

9. Clean up​

In every Kubernetes cluster, remove Vector. Deleting the namespace removes the DaemonSet, ConfigMap and ServiceAccount, and the ClusterRole and ClusterRoleBinding are cluster-wide, so they go separately. Only delete the namespace if you created it for this guide:

kubectl delete clusterrolebinding/vector-logging clusterrole/vector-logging
kubectl delete namespace logging

In the cluster that runs Dex, remove Dex, its ConfigMap and its HTTPRoute:

kubectl delete namespace dex

Then delete the OpenSearch cluster. This also deletes its indices, roles and mappings, and the retention policy:

  1. Open the cluster in the ClusterNest Console.
  2. Click Delete.
  3. Type the cluster name to confirm.
  4. Click Delete Cluster.

Part 3: One Dex per cluster​

Part 2 puts every cluster behind a single Dex. That Dex is then shared: if it goes down, no cluster can get a new token, and OpenSearch can no longer validate the tokens it issued. This part removes the shared component. Each cluster has its own Dex with its own issuer, and OpenSearch trusts each one through its own auth source. When a cluster, and with it its Dex, goes down, only that cluster's logging stops.

This part is complete on its own. You do not need Part 1 or Part 2.

It uses three example clusters, each with its own Dex:

ClusterDex issuerConnector idGroupIndex pattern
prod-euhttps://dex.prod-eu.example.comprod-euprod-eu:system:serviceaccount:logging:vectorlogs-k8s-prod-eu-*
prod-ushttps://dex.prod-us.example.comprod-usprod-us:system:serviceaccount:logging:vectorlogs-k8s-prod-us-*
staginghttps://dex.staging.example.comstagingstaging:system:serviceaccount:logging:vectorlogs-k8s-staging-*

How it works​

  1. A pod projects a ServiceAccount token with audience dex.
  2. A sidecar posts it to its own cluster's Dex, with that cluster's connector id.
  3. That Dex verifies the token against its cluster's issuer and returns an ID token whose groups claim is prod-eu:system:serviceaccount:logging:vector.
  4. The forwarder sends the Dex token to OpenSearch. OpenSearch validates it against the issuer of the matching auth source and takes the group as a backend role.
  5. The role mapped to that group can only write logs-k8s-prod-eu-*.

Run each Dex in the cluster it serves, so that an outage takes the Dex down together with the cluster. OpenSearch contacts each Dex's issuer to validate its tokens, so each Dex must be reachable over HTTPS from the internet, and each Dex must reach its own cluster's issuer.

What one Dex going down does​

With three Dexes configured as separate auth sources, one was stopped while the others kept serving. The results:

  • Tokens from the running Dex kept writing to their own index (201) at the same speed as before, about 0.2 seconds a write. The log shippers using it had no errors and no gap.
  • The stopped Dex answered 503, so its cluster could not get a new token. A token minted just before the outage kept working while OpenSearch had Dex's keys cached, and was still accepted 5 seconds into the outage.
  • After the stopped Dex came back, its cluster's tokens were rejected with 401 for about 8 minutes, then accepted again. A Dex with in-memory storage generates new signing keys when it starts, and OpenSearch picks them up on its next key fetch, which can take several minutes. The other clusters were not affected.
  • A token from one Dex could not write the other cluster's index, in either direction (403).

1. Configure each Dex​

Each Dex has one public client and one connector, for its own cluster. The client is public: true, so the exchange needs no client secret. This is the prod-eu Dex:

issuer: https://dex.prod-eu.example.com
storage:
type: memory
web:
http: 0.0.0.0:5556
oauth2:
grantTypes:
- urn:ietf:params:oauth:grant-type:token-exchange
skipApprovalScreen: true
expiry:
idTokens: 15m
staticClients:
- id: logs-shipper
name: Log shippers
public: true
redirectURIs:
- http://localhost/unused
connectors:
- type: oidc
id: prod-eu
name: prod-eu
config:
issuer: https://oidc.prod-eu.example.com
clientID: dex
clientSecret: unused
redirectURI: https://dex.prod-eu.example.com/callback
scopes:
- openid
userNameKey: sub
insecureSkipEmailVerified: true
claimModifications:
newGroupFromClaims:
- claims:
- sub
delimiter: ":"
prefix: "prod-eu"

For prod-us and staging, change the three values that name the cluster: issuer and redirectURI (https://dex.prod-us.example.com), the connector's id, name and prefix, and the connector's issuer (https://oidc.prod-us.example.com).

Find a cluster's own issuer by decoding any of its ServiceAccount tokens and reading the iss claim:

kubectl create token default --duration=10m \
| cut -d. -f2 | tr '_-' '/+' \
| awk '{while (length($0) % 4) $0 = $0 "="; print}' | base64 --decode | jq -r .iss
  • clientID: dex on the connector is the audience Dex requires on the incoming ServiceAccount token. Project tokens with audience: dex.
  • userNameKey: sub makes the Dex name claim the ServiceAccount identity. scopes: [openid] with insecureSkipEmailVerified: true is needed because ServiceAccount tokens carry no email. The connector's clientSecret is a required field that is never used.
  • newGroupFromClaims builds the group from sub. With prefix: "prod-eu" and delimiter: ":" the group reads prod-eu:system:serviceaccount:logging:vector. A prefix that already ends in a colon produces a doubled colon.
  • The exchange request must carry client_id=logs-shipper. A request that names no client is rejected.
  • With storage: type: memory, a Dex generates new signing keys every time it restarts, and its cluster's tokens are rejected for several minutes afterwards (about 8 in testing). Use a persistent storage backend so the keys survive restarts.

2. Deploy Dex​

Save the config from step 1 as config.yaml, then run Dex in the cluster whose logs it serves, and publish it at that cluster's own hostname. Repeat for every cluster, each with its own config.

kubectl create namespace logging
kubectl -n logging create configmap dex --from-file=config.yaml

A Deployment and a Service. The readiness probe uses the discovery document:

apiVersion: apps/v1
kind: Deployment
metadata:
name: dex
namespace: logging
spec:
replicas: 1
selector:
matchLabels:
app: dex
template:
metadata:
labels:
app: dex
spec:
containers:
- name: dex
image: ghcr.io/dexidp/dex:v2.45.1
args: ["dex", "serve", "/etc/dex/config.yaml"]
ports:
- {name: http, containerPort: 5556}
readinessProbe:
httpGet: {path: /.well-known/openid-configuration, port: 5556}
resources:
requests: {cpu: 10m, memory: 32Mi}
limits: {cpu: 100m, memory: 96Mi}
volumeMounts:
- {name: config, mountPath: /etc/dex}
volumes:
- name: config
configMap: {name: dex}
---
apiVersion: v1
kind: Service
metadata:
name: dex
namespace: logging
spec:
selector:
app: dex
ports:
- {name: http, port: 5556, targetPort: 5556}

Publish the Service at the issuer hostname over HTTPS, with a certificate that matches it. This example uses a Gateway API HTTPRoute on an existing Gateway with an https listener. If you use an Ingress instead, route dex.prod-eu.example.com to the dex Service on port 5556:

apiVersion: gateway.networking.k8s.io/v1
kind: HTTPRoute
metadata:
name: dex
namespace: logging
spec:
hostnames:
- dex.prod-eu.example.com
parentRefs:
- group: gateway.networking.k8s.io
kind: Gateway
name: gateway
namespace: gateway-system
sectionName: https
rules:
- backendRefs:
- {group: "", kind: Service, name: dex, port: 5556}
matches:
- path: {type: PathPrefix, value: /}

Check that Dex answers from the internet:

curl https://dex.prod-eu.example.com/.well-known/openid-configuration

The response lists "issuer": "https://dex.prod-eu.example.com" and a jwks_uri of https://dex.prod-eu.example.com/keys. A 503 no healthy upstream just after the deploy means the pod is not Ready yet. It clears within seconds once the readiness probe passes. To change the config later, update the ConfigMap and restart the Deployment. Read the storage warning in step 1 before restarting Dex.

3. Create the OpenSearch cluster with one auth source per Dex​

Add one OIDC auth source for each Dex. Give each a unique name. subject_key is name, roles_key is groups, and client_id and audience are the Dex client. With audience set, OpenSearch rejects a token Dex issued for any other client:

  1. Open the ClusterNest Console, go to OpenSearch under Services and click Launch Cluster.
  2. Enter a name and pick a tier.
  3. Under Authentication, click Add OIDC source once for each Dex and fill in:
    • Name: dex-prod-eu
    • Connect URL: https://dex.prod-eu.example.com/.well-known/openid-configuration
    • Client ID: logs-shipper
    • Audience: logs-shipper
    • Subject Key: name
    • Roles Key: groups
    • Leave Offer on OpenSearch Dashboards off.
  4. Click Create Cluster. The Console shows the admin credentials once, so save them.
  5. Wait until the cluster's state is available.

Change the name and the URL for each Dex. ClusterNest rejects an auth source whose host is not publicly resolvable.

To add a cluster later:

  1. Open the cluster in the Console and click Edit.
  2. Click Add OIDC source and fill it in the same way.
  3. Save, and wait for the state to go from updating back to available. The change takes a few minutes.

The cluster is updating for a few minutes while the change applies. Adding the second Dex to a cluster whose log shippers were already writing through the first did not interrupt them: they kept delivering at a steady rate with no gap.

4. Create one role per cluster​

Each role can write only its cluster's indices. Use the OpenSearch security API with the cluster's admin credential, and map each role by backend role, which is the group Dex emits.

Do not map users for these ServiceAccounts. The username (system:serviceaccount:logging:vector) is the same in every cluster, so a username mapping on any role grants that role's access to all of them.

Repeat steps 2 to 7 for each cluster (prod-eu, prod-us, staging), replacing <cluster>.

  1. Sign in to OpenSearch Dashboards as admin, with the credential the Console showed when it created the cluster.
  2. Open Security > Roles and click Create role. Name it logs_writer_<cluster>.
  3. Under Cluster permissions, add indices:data/write/bulk.
  4. Under Index permissions, set the index pattern to logs-k8s-<cluster>-* and the permissions to create_index, indices:data/write/bulk* and indices:data/write/index. Click Create.
  5. Open the role's Mapped users tab and click Manage mapping.
  6. Under Backend roles, add <cluster>:system:serviceaccount:logging:vector.
  7. Click Map.

To allow another ServiceAccount from one cluster, add its group (prod-eu:system:serviceaccount:logging:other) as another backend role on that cluster's mapping.

5. Check the exchange​

Mint a token for the ServiceAccount, exchange it, and read the claims of the result:

SA_TOKEN=$(kubectl create token vector -n logging --audience dex --duration=10m)

curl https://dex.prod-eu.example.com/token \
-d client_id=logs-shipper \
-d grant_type=urn:ietf:params:oauth:grant-type:token-exchange \
-d connector_id=prod-eu \
--data-urlencode "subject_token=$SA_TOKEN" \
-d subject_token_type=urn:ietf:params:oauth:token-type:id_token \
-d requested_token_type=urn:ietf:params:oauth:token-type:id_token \
--data-urlencode "scope=openid profile groups"

The access_token field of the response holds a JWT. Its claims include:

{
"iss": "https://dex.prod-eu.example.com",
"aud": "logs-shipper",
"name": "system:serviceaccount:logging:vector",
"groups": ["prod-eu:system:serviceaccount:logging:vector"]
}

requested_token_type must be id_token to get a JWT that OpenSearch can validate. The token expires after expiry.idTokens.

Use it as a Bearer token. The bulk API answers 200 and reports each write in the response items. A write to the cluster's own index succeeds, a write to another cluster's index is denied, and reads are denied because the role is write-only:

# own index: the item status is 201
curl -X POST "$OS_URL/_bulk" -H "Authorization: Bearer $DEX_TOKEN" \
-H "Content-Type: application/x-ndjson" \
--data-binary $'{"index":{"_index":"logs-k8s-prod-eu-2026.10.04"}}\n{"message":"hello"}\n'

The same token writing to another cluster's index, for example logs-k8s-prod-us-2026.10.04, is denied: the request returns 200, and the item in the response carries "status":403 with a security_exception.

6. Deploy Vector in each cluster​

Vector reads header values only when it starts, so a sidecar keeps its config in step with the token. It exchanges the pod's ServiceAccount token at Dex every 5 minutes (Dex tokens last 15), renders the Dex token into the Vector config, and Vector reloads when the file changes. The examples are for prod-eu. For another cluster, change the connector id, the index prefix and nothing else.

The ServiceAccount and read access to pod metadata. The logging namespace already exists from step 2, where Dex runs:

apiVersion: v1
kind: ServiceAccount
metadata:
name: vector
namespace: logging
---
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRole
metadata:
name: vector-logging
rules:
- apiGroups: [""]
resources: ["pods", "namespaces", "nodes"]
verbs: ["get", "list", "watch"]
---
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRoleBinding
metadata:
name: vector-logging
roleRef:
apiGroup: rbac.authorization.k8s.io
kind: ClusterRole
name: vector-logging
subjects:
- kind: ServiceAccount
name: vector
namespace: logging

The Vector config template and the exchange script. @@TOKEN@@ is replaced with the Dex token on every render:

apiVersion: v1
kind: ConfigMap
metadata:
name: vector
namespace: logging
data:
vector.yaml.tpl: |
data_dir: /vector-data
sources:
logs:
type: kubernetes_logs
extra_field_selector: metadata.namespace=logging
exclude_paths_glob_patterns:
- "/var/log/pods/logging_vector-*/**"
transforms:
limit:
type: throttle
inputs: [logs]
threshold: 50
window_secs: 1
sinks:
opensearch:
type: elasticsearch
inputs: [limit]
endpoints:
- https://logs-123-os.c9t.io
api_version: v7
mode: bulk
bulk:
index: "logs-k8s-prod-eu-%Y.%m.%d"
healthcheck:
enabled: false
request:
headers:
Authorization: "Bearer @@TOKEN@@"
buffer:
type: memory
max_events: 500
render.sh: |
#!/bin/sh
tpl=/tpl/vector.yaml.tpl
sa=/var/run/secrets/dex-exchange/token
out=/conf/vector.yaml
exchange() {
curl -fsS -m 20 https://dex.prod-eu.example.com/token \
--data-urlencode client_id=logs-shipper \
--data-urlencode grant_type=urn:ietf:params:oauth:grant-type:token-exchange \
--data-urlencode connector_id=prod-eu \
--data-urlencode "subject_token=$(cat "$sa")" \
--data-urlencode subject_token_type=urn:ietf:params:oauth:token-type:id_token \
--data-urlencode requested_token_type=urn:ietf:params:oauth:token-type:id_token \
--data-urlencode "scope=openid profile groups" \
| sed -n 's/.*"access_token":"\([^"]*\)".*/\1/p'
}
render() {
t=$(exchange)
[ -n "$t" ] || return 1
sed "s|@@TOKEN@@|$t|" "$tpl" > "$out.tmp" && mv "$out.tmp" "$out"
}
if [ "$1" = once ]; then render; exit $?; fi
while true; do
sleep 300
render || echo "token exchange failed, keeping previous token"
done
  • The exclude_paths_glob_patterns entry keeps Vector from shipping its own logs, which would otherwise feed its own errors back into the pipeline.
  • Dex runs in the same logging namespace, so its logs are shipped too. Add "/var/log/pods/logging_dex-*/**" to exclude_paths_glob_patterns to leave them out.
  • The healthcheck is off because the per-cluster role cannot read cluster information. api_version is set explicitly for the same reason: Vector would otherwise probe the cluster root.
  • To collect more than one namespace, widen extra_field_selector and add the namespace's ServiceAccounts to the role mapping.

The DaemonSet. The init container renders the first config so Vector never starts without one. If an exchange fails later, the sidecar keeps the previous token and retries after the next interval.

apiVersion: apps/v1
kind: DaemonSet
metadata:
name: vector
namespace: logging
spec:
selector:
matchLabels:
app: vector
template:
metadata:
labels:
app: vector
spec:
serviceAccountName: vector
initContainers:
- name: render-config
image: curlimages/curl:8.10.1
command: ["/bin/sh", "/tpl/render.sh", "once"]
volumeMounts:
- {name: tpl, mountPath: /tpl}
- {name: conf, mountPath: /conf}
- {name: sa-token, mountPath: /var/run/secrets/dex-exchange, readOnly: true}
containers:
- name: vector
image: timberio/vector:0.58.0-alpine
args:
- --config=/conf/vector.yaml
- --watch-config
- --watch-config-method=poll
- --watch-config-poll-interval-seconds=5
env:
- name: VECTOR_SELF_NODE_NAME
valueFrom:
fieldRef: {fieldPath: spec.nodeName}
resources:
requests: {cpu: 20m, memory: 48Mi}
limits: {cpu: 150m, memory: 128Mi}
volumeMounts:
- {name: conf, mountPath: /conf, readOnly: true}
- {name: data, mountPath: /vector-data}
- {name: varlog, mountPath: /var/log, readOnly: true}
- name: token-exchanger
image: curlimages/curl:8.10.1
command: ["/bin/sh", "/tpl/render.sh"]
volumeMounts:
- {name: tpl, mountPath: /tpl}
- {name: conf, mountPath: /conf}
- {name: sa-token, mountPath: /var/run/secrets/dex-exchange, readOnly: true}
volumes:
- name: tpl
configMap: {name: vector, defaultMode: 0555}
- name: conf
emptyDir: {medium: Memory}
- name: data
emptyDir: {}
- name: varlog
hostPath: {path: /var/log}
- name: sa-token
projected:
sources:
- serviceAccountToken:
path: token
audience: dex
expirationSeconds: 600

The rendered config lives on a memory-backed emptyDir, so the Dex token is never written to the node's disk. Check that the sidecar keeps working. Each pod logs one reload at startup and one about every 5 minutes:

kubectl -n logging logs ds/vector -c vector | grep -E 'Configuration file changed|ERROR'
kubectl -n logging logs ds/vector -c token-exchanger

7. Expire old indices​

Every cluster writes a new daily index. An index state management (ISM) policy deletes each index once it is old enough. This one deletes after 7 days, and ism_template attaches it automatically to every new logs-*-* index:

  1. In OpenSearch Dashboards, open Management > Index Management > State management policies.

  2. Click Create policy and choose the JSON editor.

  3. Set the policy ID to logs-retention and paste:

    {
    "policy": {
    "description": "Delete daily log indices after 7 days",
    "default_state": "hot",
    "states": [
    {"name": "hot", "actions": [], "transitions": [{"state_name": "delete", "conditions": {"min_index_age": "7d"}}]},
    {"name": "delete", "actions": [{"delete": {}}], "transitions": []}
    ],
    "ism_template": [{"index_patterns": ["logs-*-*"], "priority": 100}]
    }
    }
  4. Click Create.

  5. To attach the policy to indices that already exist, open Indices, select them, and choose Actions > Apply policy. The template only applies to indices created after the policy exists.

  6. Open Managed indices to check that each index shows logs-retention.

Change min_index_age to keep logs for longer or shorter.

8. Verify delivery​

Once a pod writes a log line, the document appears in the cluster's daily index. Each cluster's documents live only in its own pattern:

curl -u admin:$ADMIN_PASSWORD "$OS_URL/logs-k8s-prod-eu-*/_search?size=1&sort=timestamp:desc"

Each document carries the message and Kubernetes metadata (kubernetes.pod_namespace, kubernetes.pod_name, container and node details). In OpenSearch Dashboards, create one index pattern per cluster, for example logs-k8s-prod-eu-*, with timestamp as the time field.

9. Clean up​

In every Kubernetes cluster, remove Vector and the Dex that runs there. Deleting the logging namespace removes the Dex Deployment, Service, ConfigMap and HTTPRoute together with the Vector DaemonSet, ConfigMap and ServiceAccount. The ClusterRole and ClusterRoleBinding are cluster-wide, so they go separately. Only delete the namespace if you created it for this guide:

kubectl delete clusterrolebinding/vector-logging clusterrole/vector-logging
kubectl delete namespace logging

Then delete the OpenSearch cluster. This also deletes its indices, roles and mappings, and the retention policy:

  1. Open the cluster in the ClusterNest Console.
  2. Click Delete.
  3. Type the cluster name to confirm.
  4. Click Delete Cluster.

Part 4: Let people sign in with SAML​

The earlier parts let workloads write logs. This part lets people read them: users sign in to OpenSearch Dashboards through a SAML identity provider and see only the indices their group is allowed to read. It is complete on its own, so it works with any cluster that has logs in daily indices, whichever way they got there. The examples use per-cluster index patterns such as logs-k8s-prod-eu-* and Authentik as the identity provider.

A cluster accepts one SAML source, and it is always offered on the Dashboards login page as Login with followed by the source's name. It sits next to any other auth source, such as the OIDC source for Dex, and neither affects the other.

How it works​

  1. A user opens Dashboards and clicks Login with authentik.
  2. The identity provider authenticates them and returns a signed SAML assertion that carries their username and groups.
  3. OpenSearch takes the username from subject_key and the groups from roles_key as backend roles.
  4. A role mapping turns each group into a read-only role for one cluster's indices.

1. Register the application in your identity provider​

You need the cluster's Dashboards address, https://<name>-<organization_id>-osd.c9t.io (the opensearch_dashboards_url field of the cluster). The settings OpenSearch requires of any identity provider:

SettingValue
ACS URLhttps://<dashboards-hostname>/_opendistro/_security/saml/acs
Audiencethe cluster source's sp_entity_id, here opensearch-dashboards
BindingPOST
Signingsign both the assertion and the response
Attributesthe username, and the groups

In Authentik, create a SAML Provider with those values: set Service Provider Binding to Post, pick a Signing Certificate and enable Sign assertions and Sign responses, and attach the username, email, name and groups property mappings. Then create an Application that uses the provider.

Unsigned assertions are rejected, and a provider left on the redirect binding returns the response as a GET, which Dashboards answers with a 401.

Next, bind the groups that may sign in to the application, under Policy / Group / User Bindings. An Authentik application with no bindings can refuse every user, and the sign-in then ends on Authentik's own Permission denied / Request has been denied page without ever reaching OpenSearch. Bind exactly the groups you map in step 3.

Other identity providers follow the same settings. See the SSO overview for provider-specific guides.

2. Add the SAML source to the cluster​

idp_entity_id must match the entity ID in the identity provider's metadata exactly. For Authentik it is https://<authentik-host>/application/saml/<application-slug>/metadata/. subject_key is the attribute that holds the username, and roles_key is the attribute that holds the groups. For Authentik the defaults are:

  • subject_key: http://schemas.goauthentik.io/2021/02/saml/username
  • roles_key: http://schemas.xmlsoap.org/claims/Group
  1. Open the cluster in the ClusterNest Console and click Edit.
  2. Under Authentication, click Add SAML source and fill in:
    • Name: authentik (it becomes the text of the Dashboards button, "Login with" followed by the name)
    • IDP Metadata URL: https://auth.example.com/api/v3/providers/saml/12/metadata/?download
    • IDP Entity ID: https://auth.example.com/application/saml/opensearch/metadata/
    • SP Entity ID: opensearch-dashboards
    • Subject Key: http://schemas.goauthentik.io/2021/02/saml/username
    • Roles Key: http://schemas.xmlsoap.org/claims/Group
  3. Save.
  4. Wait for the state to go from updating back to available. The change takes a few minutes.

3. Map each group to a read-only role​

A user who signs in has no permissions until you map their groups to roles. Create one reader role per cluster, limited to that cluster's index pattern, and map an identity provider group to it. This example uses a group logs-prod-eu:

Repeat steps 2 to 7 for each cluster's group.

  1. Sign in to OpenSearch Dashboards as admin.
  2. Open Security > Roles and click Create role. Name it logs_reader_prod-eu.
  3. Under Cluster permissions, add cluster_composite_ops_ro.
  4. Under Index permissions, set the index pattern to logs-k8s-prod-eu-* and the permissions to read, indices:admin/mappings/get and indices:admin/resolve/index. Click Create.
  5. Open the role's Mapped users tab and click Manage mapping.
  6. Under Backend roles, add logs-prod-eu and click Map.
  7. Open Security > Roles, choose kibana_user, open Mapped users > Manage mapping, add logs-prod-eu as a backend role and click Map.

The kibana_user mapping gives the group the basic Dashboards access that every user needs. When several groups share it, list all of them in that one mapping.

Repeat the role and its mapping with prod-us and staging for the other clusters, so a person in logs-prod-eu can read prod-eu logs and nothing else.

4. Sign in and check​

Open the Dashboards address and click Login with authentik. After signing in, the user has the roles kibana_user and logs_reader_prod-eu (Dashboards also adds own_index, a built-in role for each user's own tenant).

Query the logs in Dev Tools (Management > Dev Tools). It sends each request as the signed-in user, so there is no cookie or password to handle. Read the newest documents of the cluster's own indices:

GET logs-k8s-prod-eu-*/_search
{
"size": 5,
"sort": [{"timestamp": "desc"}]
}

Everything outside the role's index pattern is denied with a security_exception, whether it is another cluster's logs, the audit log, or the security index:

GET logs-k8s-prod-us-*/_count
GET security-auditlog-*/_count
GET .opendistro_security/_count

To browse the logs, create an index pattern named exactly logs-k8s-prod-eu-*, with timestamp as the time field, and open Discover. An index pattern such as * fails, because it includes indices the role cannot see. Always create patterns that match the role's pattern.

Remove access​

To remove someone's access, take them out of the group in the identity provider. To remove a group's access, unbind it from the application (the sign-in is then refused before it reaches OpenSearch) and delete its mapping:

  1. Sign in to OpenSearch Dashboards as admin.
  2. Open Security > Roles > logs_reader_prod-eu > Mapped users and click Manage mapping.
  3. Remove the group from Backend roles and click Map.
  4. Open Security > Roles > kibana_user > Mapped users, click Manage mapping, remove the group there too and click Map.