Ship Kubernetes Logs to ClusterNest Managed OpenSearch Without Static Credentials
Pods authenticate to OpenSearch with their own ServiceAccount token. There is no password to distribute or rotate: kubelet mints and refreshes the token, and OpenSearch validates it against an OIDC issuer. This guide uses Vector as the log forwarder, with a small sidecar that keeps its configuration in step with the rotating token.
There are four parts. Parts 1 to 3 are three ways to ship logs without static credentials, and Part 4 covers how people read them. Each part is complete on its own, so you can follow any of them without the others:
- Part 1 trusts one Kubernetes cluster's issuer directly.
- Part 2 puts Dex in front, so several clusters can ship logs to one OpenSearch cluster, each limited to its own index pattern.
- Part 3 gives every cluster its own Dex and its own auth source, so one cluster going down does not affect the others.
- Part 4 lets people read the logs by signing in to Dashboards with SAML, with access limited to each cluster's index pattern.
Why Vector and not Fluent Bit
The forwarder has to send the ServiceAccount token (or a token derived from it) to OpenSearch as Authorization: Bearer <token>. Fluent Bit is the usual choice for Kubernetes logs, but it cannot do this with its OpenSearch output:
- Fluent Bit 5.1.3's
opensearchoutput has no option to send a custom header or a Bearer token. Its authentication options arehttp_user,http_passwdand the AWS ones, and an unknown property such asheaderstops the output from initialising. - Fluent Bit's generic
httpoutput does accept headers, but it posts plain records. OpenSearch's bulk API needs an action line before every document, so you would have to build that yourself, for example with a Lua filter.
Vector's elasticsearch sink takes arbitrary request headers and speaks the bulk API natively, so it can send the token directly. It also reloads its own config when the file changes, which is how the guide keeps up with a token that rotates every few minutes.
Part 1: One cluster, direct
How it works
- Your Kubernetes cluster publishes an OIDC issuer (ServiceAccount issuer discovery).
- The OpenSearch cluster trusts that issuer through an
oidcauth source. - A pod projects a ServiceAccount token and sends it as
Authorization: Bearer <token>. - OpenSearch takes the token's
subclaim (system:serviceaccount:<namespace>:<name>) as the username. A role mapping on that username decides what the pod may write.
Requirements
- The issuer URL is served over HTTPS and its
/.well-known/openid-configurationand JWKS (/openid/v1/jwks) are reachable from the internet. ClusterNest rejects an auth source whose host is not publicly resolvable.
Find the issuer of your cluster by decoding any ServiceAccount token and reading its iss claim:
kubectl create token default --duration=10m \
| cut -d. -f2 | tr '_-' '/+' \
| awk '{while (length($0) % 4) $0 = $0 "="; print}' | base64 --decode | jq -r .iss
1. Create the cluster with an OIDC auth source
Set audience to the audience your tokens are issued for (the --audience you pass when minting a token below). client_id is optional: it is only used for Dashboards login, so leave it out. subject_key must be sub.
- Console
- API
- Terraform
- Open the ClusterNest Console, go to OpenSearch under Services and click Launch Cluster.
- Enter a name and pick a tier.
- Under Authentication, click Add OIDC source and fill in:
- Name:
k8s - Connect URL:
https://oidc.example.com/.well-known/openid-configuration - Audience:
https://oidc.example.com - Subject Key:
sub - Leave Offer on OpenSearch Dashboards off.
- Name:
- Click Create Cluster. The Console shows the admin credentials once, so save them.
- Wait until the cluster's state is
available.
The API endpoint is on the cluster page.
Every request to the ClusterNest API carries an access token. Generate an app password from the MFA settings page of your account (your regular password does not work), then exchange it for a token:
export ACCESS_TOKEN=$(curl -sS -X POST https://api.clusternest.com/auth/token \
-H "Content-Type: application/json" \
-d '{"email": "[email protected]", "password": "<app password>"}' \
| jq -r .access_token)
The token expires, so run this again when requests start returning 401. See Get an access token for the full request and response.
curl -X POST "https://api.clusternest.com/cluster/opensearch/" \
-H "Authorization: Bearer $ACCESS_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"name": "logs",
"organization_id": 123,
"tier": "basic",
"auth_sources": [
{
"type": "oidc",
"name": "k8s",
"connect_url": "https://oidc.example.com/.well-known/openid-configuration",
"audience": "https://oidc.example.com",
"subject_key": "sub",
"dashboards_login": false
}
]
}'
Poll GET /cluster/opensearch/$CLUSTER_ID until state is available. The API endpoint is https://<name>-<organization_id>-os.c9t.io.
terraform {
required_providers {
clusternest = {
source = "tf.clusternest.com/clusternest/clusternest"
version = ">=1.1.0"
}
opensearch = {
source = "opensearch-project/opensearch"
version = ">= 2.2.0"
}
}
}
provider "clusternest" {}
resource "clusternest_opensearch" "logs" {
name = "logs"
tier = "basic"
organization_id = 123
auth_sources = [
{
name = "k8s"
type = "oidc"
connect_url = "https://oidc.example.com/.well-known/openid-configuration"
audience = "https://oidc.example.com"
subject_key = "sub"
dashboards_login = false
},
]
}
output "opensearch_url" {
value = clusternest_opensearch.logs.url
}
The provider reads your email and app password from CLUSTERNEST_EMAIL and CLUSTERNEST_APP_PASSWORD. terraform apply returns once the cluster is available, and opensearch_url is the API endpoint.
2. Create a role and map the ServiceAccount
ClusterNest does not create role mappings. Create them with the cluster's admin credential. Replace logging and vector below with the namespace and ServiceAccount your forwarder runs as.
- Dashboards
- API
- Terraform
- Sign in to OpenSearch Dashboards as
admin, with the credential the Console showed when it created the cluster. - Open Security > Roles and click Create role. Name it
logs_writer. - Under Cluster permissions, add
indices:data/write/bulk. - Under Index permissions, set the index pattern to
logs-*and the permissions tocreate_index,indices:data/write/bulk*andindices:data/write/index. Click Create. - Open the role's Mapped users tab and click Manage mapping.
- Under Users, add
system:serviceaccount:logging:vectorand click Map.
users accepts wildcards, so system:serviceaccount:logging:* covers every ServiceAccount in a namespace. To add a ServiceAccount later, open Manage mapping again and add another user.
Fetch the admin credential from the ClusterNest API into a file, so the password is never printed:
(umask 077; curl -sS --fail-with-body \
-H "Authorization: Bearer $ACCESS_TOKEN" \
"https://api.clusternest.com/cluster/opensearch/$CLUSTER_ID/credentials" \
-o credentials.json)
export ADMIN_PASSWORD=$(jq -r .password credentials.json)
Then create the role and the mapping:
export OS_URL=https://logs-123-os.c9t.io
curl -u admin:$ADMIN_PASSWORD -X PUT "$OS_URL/_plugins/_security/api/roles/logs_writer" \
-H "Content-Type: application/json" -d '{
"cluster_permissions": ["indices:data/write/bulk"],
"index_permissions": [{
"index_patterns": ["logs-*"],
"allowed_actions": ["create_index", "indices:data/write/bulk*", "indices:data/write/index"]
}]
}'
curl -u admin:$ADMIN_PASSWORD -X PUT "$OS_URL/_plugins/_security/api/rolesmapping/logs_writer" \
-H "Content-Type: application/json" \
-d '{"users": ["system:serviceaccount:logging:vector"]}'
users accepts wildcards, so system:serviceaccount:logging:* covers every ServiceAccount in a namespace. To add a ServiceAccount later without replacing the list, use PATCH:
curl -u admin:$ADMIN_PASSWORD -X PATCH "$OS_URL/_plugins/_security/api/rolesmapping/logs_writer" \
-H "Content-Type: application/json" \
-d '[{"op": "add", "path": "/users/-", "value": "system:serviceaccount:logging:other"}]'
The opensearch provider signs in with the cluster's admin credential, which the clusternest_opensearch_credentials data source reads. Add this to the configuration from step 1:
data "clusternest_opensearch_credentials" "logs" {
cluster_id = clusternest_opensearch.logs.id
}
provider "opensearch" {
url = clusternest_opensearch.logs.url
username = data.clusternest_opensearch_credentials.logs.username
password = data.clusternest_opensearch_credentials.logs.password
}
resource "opensearch_role" "logs_writer" {
role_name = "logs_writer"
cluster_permissions = ["indices:data/write/bulk"]
index_permissions {
index_patterns = ["logs-*"]
allowed_actions = ["create_index", "indices:data/write/bulk*", "indices:data/write/index"]
}
}
resource "opensearch_roles_mapping" "logs_writer" {
role_name = opensearch_role.logs_writer.role_name
users = ["system:serviceaccount:logging:vector"]
}
users accepts wildcards, so system:serviceaccount:logging:* covers every ServiceAccount in a namespace. To add a ServiceAccount later, add it to the list.
With audience set, OpenSearch rejects a token whose aud claim does not match, even if the trusted issuer signed it. Without it, a token for any audience from that issuer is accepted and access is decided by the sub mapping alone, so keep the mapping as narrow as the workloads that need it.
3. Check authentication
Mint a token for the ServiceAccount and call the cluster. A ServiceAccount with no mapping authenticates but is denied, which proves the issuer is trusted and the sub becomes the username:
kubectl create token default -n default --audience https://oidc.example.com --duration=10m \
| sed 's/^/Authorization: Bearer /' \
| curl -H @- "$OS_URL/_cluster/health"
{"error":{"root_cause":[{"type":"security_exception","reason":"no permissions for [cluster:monitor/health] and User [name=system:serviceaccount:default:default, backend_roles=[], requestedTenant=null]"}],"type":"security_exception","reason":"no permissions for [cluster:monitor/health] and User [name=system:serviceaccount:default:default, backend_roles=[], requestedTenant=null]"},"status":403}
A 401 means OpenSearch could not validate the token, usually because the issuer is not reachable from the internet.
4. Deploy Vector
Vector reads header values when it starts and does not re-read a file. The sidecar renders the Vector config with the current token and rewrites it whenever the token changes. Vector watches the config file and reloads itself.
The pod needs a projected token, a ServiceAccount, and read access to pod metadata:
apiVersion: v1
kind: ServiceAccount
metadata:
name: vector
namespace: logging
---
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRole
metadata:
name: vector-logging
rules:
- apiGroups: [""]
resources: ["pods", "namespaces", "nodes"]
verbs: ["get", "list", "watch"]
---
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRoleBinding
metadata:
name: vector-logging
roleRef:
apiGroup: rbac.authorization.k8s.io
kind: ClusterRole
name: vector-logging
subjects:
- kind: ServiceAccount
name: vector
namespace: logging
The config template and renderer script. @@TOKEN@@ is replaced with the token on every render:
apiVersion: v1
kind: ConfigMap
metadata:
name: vector
namespace: logging
data:
vector.yaml.tpl: |
data_dir: /vector-data
sources:
logs:
type: kubernetes_logs
extra_field_selector: metadata.namespace=logging
exclude_paths_glob_patterns:
- "/var/log/pods/logging_vector-*/**"
transforms:
limit:
type: throttle
inputs: [logs]
threshold: 50
window_secs: 1
sinks:
opensearch:
type: elasticsearch
inputs: [limit]
endpoints:
- https://logs-123-os.c9t.io
api_version: v7
mode: bulk
bulk:
index: "logs-k8s-%Y.%m.%d"
healthcheck:
enabled: false
request:
headers:
Authorization: "Bearer @@TOKEN@@"
buffer:
type: memory
max_events: 500
render.sh: |
#!/bin/sh
tpl=/tpl/vector.yaml.tpl
tok=/var/run/secrets/opensearch/token
out=/conf/vector.yaml
render() {
sed "s|@@TOKEN@@|$(cat "$tok")|" "$tpl" > "$out.tmp" && mv "$out.tmp" "$out"
}
render
[ "$1" = once ] && exit 0
last=$(cksum < "$tok")
while true; do
sleep 5
cur=$(cksum < "$tok")
if [ "$cur" != "$last" ]; then
render && last=$cur
fi
done
The exclude_paths_glob_patterns entry keeps Vector from shipping its own logs, which would otherwise feed its own errors back into the pipeline. The sink's health check is disabled because the logs_writer role cannot read cluster information, and api_version is set explicitly for the same reason: Vector would otherwise probe the cluster root. To collect more than one namespace, widen extra_field_selector and the role mapping together.
The DaemonSet:
apiVersion: apps/v1
kind: DaemonSet
metadata:
name: vector
namespace: logging
spec:
selector:
matchLabels:
app: vector
template:
metadata:
labels:
app: vector
spec:
serviceAccountName: vector
initContainers:
- name: render-config
image: busybox:1.36
command: ["/bin/sh", "/tpl/render.sh", "once"]
volumeMounts:
- {name: tpl, mountPath: /tpl}
- {name: conf, mountPath: /conf}
- {name: token, mountPath: /var/run/secrets/opensearch, readOnly: true}
containers:
- name: vector
image: timberio/vector:0.58.0-alpine
args:
- --config=/conf/vector.yaml
- --watch-config
- --watch-config-method=poll
- --watch-config-poll-interval-seconds=5
env:
- name: VECTOR_SELF_NODE_NAME
valueFrom:
fieldRef: {fieldPath: spec.nodeName}
resources:
requests: {cpu: 20m, memory: 48Mi}
limits: {cpu: 150m, memory: 128Mi}
volumeMounts:
- {name: conf, mountPath: /conf, readOnly: true}
- {name: data, mountPath: /vector-data}
- {name: varlog, mountPath: /var/log, readOnly: true}
- name: token-renderer
image: busybox:1.36
command: ["/bin/sh", "/tpl/render.sh"]
volumeMounts:
- {name: tpl, mountPath: /tpl}
- {name: conf, mountPath: /conf}
- {name: token, mountPath: /var/run/secrets/opensearch, readOnly: true}
volumes:
- name: tpl
configMap: {name: vector, defaultMode: 0555}
- name: conf
emptyDir: {medium: Memory}
- name: data
emptyDir: {}
- name: varlog
hostPath: {path: /var/log}
- name: token
projected:
sources:
- serviceAccountToken:
path: token
audience: https://oidc.example.com
expirationSeconds: 600
The config lives on a memory-backed emptyDir, so the token is never written to disk on the node. Set audience to your issuer URL.
Token rotation
The projected token is valid for expirationSeconds (here 10 minutes, the minimum). Kubelet replaces the file at about 80% of that lifetime. The token-renderer sidecar notices the new content within 5 seconds and rewrites the Vector config, and Vector reloads without dropping events. Each pod logs a reload once at startup and once per rotation:
kubectl -n logging logs ds/vector -c vector | grep -E 'Configuration file changed|has reloaded|ERROR'
INFO vector::config::watcher: Configuration file changed.
INFO vector: Vector has reloaded. path=[File("/conf/vector.yaml", None)]
With a log line written every 5 seconds, delivery stayed at a steady 6 documents per 30 seconds across the rotation of all three pods, with no ERROR or 401 lines.
Pods restarted by a rollout start with a new token and rotate about 8 minutes later, so reloads in the first minutes after a rollout are only the startup render.
5. Expire old indices
A daily index is created every day and never removed on its own. An index state management (ISM) policy deletes each index once it is old enough. This one deletes after 7 days. ism_template attaches it automatically to every new logs-k8s-* index:
- Dashboards
- API
- Terraform
-
In OpenSearch Dashboards, open Management > Index Management > State management policies.
-
Click Create policy and choose the JSON editor.
-
Set the policy ID to
logs-retentionand paste:{"policy": {"description": "Delete daily log indices after 7 days","default_state": "hot","states": [{"name": "hot", "actions": [], "transitions": [{"state_name": "delete", "conditions": {"min_index_age": "7d"}}]},{"name": "delete", "actions": [{"delete": {}}], "transitions": []}],"ism_template": [{"index_patterns": ["logs-k8s-*"], "priority": 100}]}} -
Click Create.
-
To attach the policy to indices that already exist, open Indices, select them, and choose Actions > Apply policy. The template only applies to indices created after the policy exists.
-
Open Managed indices to check that each index shows
logs-retention.
Change min_index_age to keep logs for longer or shorter.
curl -u admin:$ADMIN_PASSWORD -X PUT "$OS_URL/_plugins/_ism/policies/logs-retention" \
-H "Content-Type: application/json" -d '{
"policy": {
"description": "Delete daily log indices after 7 days",
"default_state": "hot",
"states": [
{"name": "hot", "actions": [], "transitions": [{"state_name": "delete", "conditions": {"min_index_age": "7d"}}]},
{"name": "delete", "actions": [{"delete": {}}], "transitions": []}
],
"ism_template": [{"index_patterns": ["logs-k8s-*"], "priority": 100}]
}
}'
The template only applies to indices created after the policy exists. Attach it to indices that already exist:
curl -u admin:$ADMIN_PASSWORD -X POST "$OS_URL/_plugins/_ism/add/logs-k8s-*" \
-H "Content-Type: application/json" -d '{"policy_id": "logs-retention"}'
Check that an index is managed:
curl -u admin:$ADMIN_PASSWORD "$OS_URL/_plugins/_ism/explain/logs-k8s-*"
Each managed index lists "policy_id": "logs-retention". Change min_index_age to keep logs for longer or shorter.
Add the policy to the configuration from the previous steps. The opensearch provider block is the one from step 2.
resource "opensearch_ism_policy" "logs_retention" {
policy_id = "logs-retention"
body = jsonencode({
policy = {
description = "Delete daily log indices after 7 days"
default_state = "hot"
states = [
{ name = "hot", actions = [], transitions = [{ state_name = "delete", conditions = { min_index_age = "7d" } }] },
{ name = "delete", actions = [{ delete = {} }], transitions = [] },
]
ism_template = [{ index_patterns = ["logs-k8s-*"], priority = 100 }]
}
})
}
The template only applies to indices created after the policy exists. Apply the policy to existing indices in Dashboards or with the API.
6. Verify delivery
Once a pod writes a log line, the document appears in the day's index, logs-k8s-YYYY.MM.DD:
curl -u admin:$ADMIN_PASSWORD "$OS_URL/logs-k8s-*/_search?size=1&sort=timestamp:desc"
Each document carries the message and Kubernetes metadata (kubernetes.pod_namespace, kubernetes.pod_name, container and node details). In OpenSearch Dashboards, create an index pattern logs-k8s-* with timestamp as the time field. The role from step 2 allows every logs-* index, so no permission change is needed.
7. Clean up
Remove Vector from the Kubernetes cluster. Deleting the namespace removes the DaemonSet, ConfigMap and ServiceAccount, and the ClusterRole and ClusterRoleBinding are cluster-wide, so they go separately. Only delete the namespace if you created it for this guide:
kubectl delete clusterrolebinding/vector-logging clusterrole/vector-logging
kubectl delete namespace logging
Then delete the OpenSearch cluster. This also deletes its indices, the role and mapping, and the retention policy:
- Console
- API
- Terraform
- Open the cluster in the ClusterNest Console.
- Click Delete.
- Type the cluster name to confirm.
- Click Delete Cluster.
curl -X DELETE "https://api.clusternest.com/cluster/opensearch/$CLUSTER_ID" \
-H "Authorization: Bearer $ACCESS_TOKEN"
rm -f credentials.json
terraform destroy
The cluster shows "state": "deleting" until it is gone.
Part 2: Several clusters through Dex
With several Kubernetes clusters, trusting each cluster's issuer directly (Part 1) is not enough. The sub claim of a ServiceAccount token is system:serviceaccount:<namespace>:<name> in every cluster, so logging/vector in prod-eu is indistinguishable from logging/vector in staging. A role mapping cannot require both a cluster and a namespace.
Dex fixes this. It trusts each cluster's issuer through its own connector, exchanges a ServiceAccount token for a Dex token, and prefixes a groups claim with the connector's id. OpenSearch trusts only Dex and maps that group to a role limited to the cluster's own index pattern. There is still no static credential: the only secret in the flow is the pod's own short-lived ServiceAccount token.
This part is complete on its own. You do not need Part 1. One Dex serves any number of clusters, and you can run more than one Dex, each added to OpenSearch as its own auth source.
It uses three example clusters:
| Cluster | Connector id | Group | Index pattern |
|---|---|---|---|
prod-eu | prod-eu | prod-eu:system:serviceaccount:logging:vector | logs-k8s-prod-eu-* |
prod-us | prod-us | prod-us:system:serviceaccount:logging:vector | logs-k8s-prod-us-* |
staging | staging | staging:system:serviceaccount:logging:vector | logs-k8s-staging-* |
How it works
- A pod projects a ServiceAccount token with audience
dex. - A sidecar posts it to Dex with the connector id of the pod's cluster (
prod-eu). - Dex verifies the token against that cluster's issuer and returns an ID token whose
groupsclaim isprod-eu:system:serviceaccount:logging:vector. - The forwarder sends the Dex token to OpenSearch, which takes the group as a backend role.
- The role mapped to that group can only write
logs-k8s-prod-eu-*.
Requirements
- Dex is served over HTTPS on a public hostname (here
https://dex.example.com). ClusterNest rejects an auth source whose host is not publicly resolvable. - Dex can reach each cluster's issuer over HTTPS: the issuer's
/.well-known/openid-configurationand its JWKS (/openid/v1/jwks). - The OpenSearch cluster can reach Dex.
Find a cluster's issuer by decoding any of its ServiceAccount tokens and reading the iss claim:
kubectl create token default --duration=10m \
| cut -d. -f2 | tr '_-' '/+' \
| awk '{while (length($0) % 4) $0 = $0 "="; print}' | base64 --decode | jq -r .iss
1. Configure Dex
One public client is shared by every shipper, and one connector per cluster verifies that cluster's tokens. The client is public: true, so the exchange needs no client secret.
issuer: https://dex.example.com
storage:
type: memory
web:
http: 0.0.0.0:5556
oauth2:
grantTypes:
- urn:ietf:params:oauth:grant-type:token-exchange
skipApprovalScreen: true
expiry:
idTokens: 15m
staticClients:
- id: logs-shipper
name: Log shippers
public: true
redirectURIs:
- http://localhost/unused
connectors:
- type: oidc
id: prod-eu
name: prod-eu
config:
issuer: https://oidc.prod-eu.example.com
clientID: dex
clientSecret: unused
redirectURI: https://dex.example.com/callback
scopes:
- openid
userNameKey: sub
insecureSkipEmailVerified: true
claimModifications:
newGroupFromClaims:
- claims:
- sub
delimiter: ":"
prefix: "prod-eu"
- type: oidc
id: prod-us
name: prod-us
config:
issuer: https://oidc.prod-us.example.com
clientID: dex
clientSecret: unused
redirectURI: https://dex.example.com/callback
scopes:
- openid
userNameKey: sub
insecureSkipEmailVerified: true
claimModifications:
newGroupFromClaims:
- claims:
- sub
delimiter: ":"
prefix: "prod-us"
- type: oidc
id: staging
name: staging
config:
issuer: https://oidc.staging.example.com
clientID: dex
clientSecret: unused
redirectURI: https://dex.example.com/callback
scopes:
- openid
userNameKey: sub
insecureSkipEmailVerified: true
claimModifications:
newGroupFromClaims:
- claims:
- sub
delimiter: ":"
prefix: "staging"
clientID: dexon a connector is the audience Dex requires on the incoming ServiceAccount token. Project tokens withaudience: dex.userNameKey: submakes the Dexnameclaim the ServiceAccount identity.scopes: [openid]withinsecureSkipEmailVerified: trueis needed because ServiceAccount tokens carry no email. The connector'sclientSecretis a required field that is never used.newGroupFromClaimsbuilds the group fromsub. Withprefix: "prod-eu"anddelimiter: ":"the group readsprod-eu:system:serviceaccount:logging:vector. A prefix that already ends in a colon produces a doubled colon.- The exchange request must carry
client_id=logs-shipper. A request that names no client is rejected.
With storage: type: memory, Dex generates new signing keys every time it restarts. OpenSearch re-fetches Dex's keys only at a limited rate, so for a few minutes after a restart it rejects tokens signed by the new keys with 401 Authentication finally failed, and Vector drops the events it could not send in that window. In testing this lasted about 4 minutes after a restart that followed several earlier restarts. Use a persistent Dex storage backend so the keys survive restarts.
2. Deploy Dex
Save the config from step 1 as config.yaml, then run Dex in a cluster that can publish HTTPS services. It does not have to be one of the clusters that ship logs.
kubectl create namespace dex
kubectl -n dex create configmap dex --from-file=config.yaml
A Deployment and a Service. The readiness probe uses the discovery document:
apiVersion: apps/v1
kind: Deployment
metadata:
name: dex
namespace: dex
spec:
replicas: 1
selector:
matchLabels:
app: dex
template:
metadata:
labels:
app: dex
spec:
containers:
- name: dex
image: ghcr.io/dexidp/dex:v2.45.1
args: ["dex", "serve", "/etc/dex/config.yaml"]
ports:
- {name: http, containerPort: 5556}
readinessProbe:
httpGet: {path: /.well-known/openid-configuration, port: 5556}
resources:
requests: {cpu: 10m, memory: 32Mi}
limits: {cpu: 100m, memory: 96Mi}
volumeMounts:
- {name: config, mountPath: /etc/dex}
volumes:
- name: config
configMap: {name: dex}
---
apiVersion: v1
kind: Service
metadata:
name: dex
namespace: dex
spec:
selector:
app: dex
ports:
- {name: http, port: 5556, targetPort: 5556}
Publish the Service at the issuer hostname over HTTPS, with a certificate that matches it. This example uses a Gateway API HTTPRoute on an existing Gateway with an https listener. If you use an Ingress instead, route dex.example.com to the dex Service on port 5556:
apiVersion: gateway.networking.k8s.io/v1
kind: HTTPRoute
metadata:
name: dex
namespace: dex
spec:
hostnames:
- dex.example.com
parentRefs:
- group: gateway.networking.k8s.io
kind: Gateway
name: gateway
namespace: gateway-system
sectionName: https
rules:
- backendRefs:
- {group: "", kind: Service, name: dex, port: 5556}
matches:
- path: {type: PathPrefix, value: /}
Check that Dex answers from the internet:
curl https://dex.example.com/.well-known/openid-configuration
The response lists "issuer": "https://dex.example.com" and a jwks_uri of https://dex.example.com/keys. A 503 no healthy upstream just after the deploy means the pod is not Ready yet. It clears within seconds once the readiness probe passes. To change the config later, update the ConfigMap and restart the Deployment. Read the storage warning in step 1 before restarting Dex.
3. Create the OpenSearch cluster with Dex as the auth source
subject_key is name and roles_key is groups. client_id and audience are the Dex client. With audience set, OpenSearch rejects a token Dex issued for any other client:
- Console
- API
- Terraform
- Open the ClusterNest Console, go to OpenSearch under Services and click Launch Cluster.
- Enter a name and pick a tier.
- Under Authentication, click Add OIDC source and fill in:
- Name:
dex - Connect URL:
https://dex.example.com/.well-known/openid-configuration - Client ID:
logs-shipper - Audience:
logs-shipper - Subject Key:
name - Roles Key:
groups - Leave Offer on OpenSearch Dashboards off.
- Name:
- Click Create Cluster. The Console shows the admin credentials once, so save them.
- Wait until the cluster's state is
available.
Every request to the ClusterNest API carries an access token. Generate an app password from the MFA settings page of your account (your regular password does not work), then exchange it for a token:
export ACCESS_TOKEN=$(curl -sS -X POST https://api.clusternest.com/auth/token \
-H "Content-Type: application/json" \
-d '{"email": "[email protected]", "password": "<app password>"}' \
| jq -r .access_token)
The token expires, so run this again when requests start returning 401. See Get an access token for the full request and response.
curl -X POST "https://api.clusternest.com/cluster/opensearch/" \
-H "Authorization: Bearer $ACCESS_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"name": "logs",
"organization_id": 123,
"tier": "basic",
"auth_sources": [
{
"type": "oidc",
"name": "dex",
"connect_url": "https://dex.example.com/.well-known/openid-configuration",
"client_id": "logs-shipper",
"audience": "logs-shipper",
"subject_key": "name",
"roles_key": "groups",
"dashboards_login": false
}
]
}'
Poll GET /cluster/opensearch/$CLUSTER_ID until state is available. The API endpoint is https://<name>-<organization_id>-os.c9t.io.
terraform {
required_providers {
clusternest = {
source = "tf.clusternest.com/clusternest/clusternest"
version = ">=1.1.0"
}
opensearch = {
source = "opensearch-project/opensearch"
version = ">= 2.2.0"
}
}
}
provider "clusternest" {}
locals {
clusters = ["prod-eu", "prod-us", "staging"]
}
resource "clusternest_opensearch" "logs" {
name = "logs"
tier = "basic"
organization_id = 123
auth_sources = [
{
name = "dex"
type = "oidc"
connect_url = "https://dex.example.com/.well-known/openid-configuration"
client_id = "logs-shipper"
audience = "logs-shipper"
subject_key = "name"
roles_key = "groups"
dashboards_login = false
},
]
}
output "opensearch_url" {
value = clusternest_opensearch.logs.url
}
The provider reads your email and app password from CLUSTERNEST_EMAIL and CLUSTERNEST_APP_PASSWORD. terraform apply returns once the cluster is available, and opensearch_url is the API endpoint.
4. Create one role per cluster
Each role can write only its cluster's indices. Use the OpenSearch security API with the cluster's admin credential, and map each role by backend role, which is the group Dex emits.
Do not map users for these ServiceAccounts. The username (system:serviceaccount:logging:vector) is the same in every cluster, so a username mapping on any role grants that role's access to all of them.
- Dashboards
- API
- Terraform
Repeat steps 2 to 7 for each cluster (prod-eu, prod-us, staging), replacing <cluster>.
- Sign in to OpenSearch Dashboards as
admin, with the credential the Console showed when it created the cluster. - Open Security > Roles and click Create role. Name it
logs_writer_<cluster>. - Under Cluster permissions, add
indices:data/write/bulk. - Under Index permissions, set the index pattern to
logs-k8s-<cluster>-*and the permissions tocreate_index,indices:data/write/bulk*andindices:data/write/index. Click Create. - Open the role's Mapped users tab and click Manage mapping.
- Under Backend roles, add
<cluster>:system:serviceaccount:logging:vector. - Click Map.
To allow another ServiceAccount from one cluster, add its group (prod-eu:system:serviceaccount:logging:other) as another backend role on that cluster's mapping.
Fetch the admin credential from the ClusterNest API into a file, so the password is never printed:
(umask 077; curl -sS --fail-with-body \
-H "Authorization: Bearer $ACCESS_TOKEN" \
"https://api.clusternest.com/cluster/opensearch/$CLUSTER_ID/credentials" \
-o credentials.json)
export ADMIN_PASSWORD=$(jq -r .password credentials.json)
export OS_URL=https://logs-123-os.c9t.io
for cluster in prod-eu prod-us staging; do
curl -u admin:$ADMIN_PASSWORD -X PUT "$OS_URL/_plugins/_security/api/roles/logs_writer_$cluster" \
-H "Content-Type: application/json" -d '{
"cluster_permissions": ["indices:data/write/bulk"],
"index_permissions": [{
"index_patterns": ["logs-k8s-'"$cluster"'-*"],
"allowed_actions": ["create_index", "indices:data/write/bulk*", "indices:data/write/index"]
}]
}'
curl -u admin:$ADMIN_PASSWORD -X PUT "$OS_URL/_plugins/_security/api/rolesmapping/logs_writer_$cluster" \
-H "Content-Type: application/json" \
-d '{"backend_roles": ["'"$cluster"':system:serviceaccount:logging:vector"]}'
done
To allow another ServiceAccount from one cluster, add its group to that cluster's mapping:
curl -u admin:$ADMIN_PASSWORD -X PATCH "$OS_URL/_plugins/_security/api/rolesmapping/logs_writer_prod-eu" \
-H "Content-Type: application/json" \
-d '[{"op": "add", "path": "/backend_roles/-", "value": "prod-eu:system:serviceaccount:logging:other"}]'
The opensearch provider signs in with the cluster's admin credential, which the clusternest_opensearch_credentials data source reads. Add this to the configuration from step 3:
data "clusternest_opensearch_credentials" "logs" {
cluster_id = clusternest_opensearch.logs.id
}
provider "opensearch" {
url = clusternest_opensearch.logs.url
username = data.clusternest_opensearch_credentials.logs.username
password = data.clusternest_opensearch_credentials.logs.password
}
resource "opensearch_role" "logs_writer" {
for_each = toset(local.clusters)
role_name = "logs_writer_${each.key}"
cluster_permissions = ["indices:data/write/bulk"]
index_permissions {
index_patterns = ["logs-k8s-${each.key}-*"]
allowed_actions = ["create_index", "indices:data/write/bulk*", "indices:data/write/index"]
}
}
resource "opensearch_roles_mapping" "logs_writer" {
for_each = toset(local.clusters)
role_name = opensearch_role.logs_writer[each.key].role_name
backend_roles = ["${each.key}:system:serviceaccount:logging:vector"]
}
To allow another ServiceAccount from one cluster, add its group to that cluster's backend_roles.
5. Check the exchange
Mint a token for the ServiceAccount, exchange it, and read the claims of the result:
SA_TOKEN=$(kubectl create token vector -n logging --audience dex --duration=10m)
curl https://dex.example.com/token \
-d client_id=logs-shipper \
-d grant_type=urn:ietf:params:oauth:grant-type:token-exchange \
-d connector_id=prod-eu \
--data-urlencode "subject_token=$SA_TOKEN" \
-d subject_token_type=urn:ietf:params:oauth:token-type:id_token \
-d requested_token_type=urn:ietf:params:oauth:token-type:id_token \
--data-urlencode "scope=openid profile groups"
The access_token field of the response holds a JWT. Its claims include:
{
"iss": "https://dex.example.com",
"aud": "logs-shipper",
"name": "system:serviceaccount:logging:vector",
"groups": ["prod-eu:system:serviceaccount:logging:vector"]
}
requested_token_type must be id_token to get a JWT that OpenSearch can validate. The token expires after expiry.idTokens.
Use it as a Bearer token. The bulk API answers 200 and reports each write in the response items. A write to the cluster's own index succeeds, a write to another cluster's index is denied, and reads are denied because the role is write-only:
# own index: the item status is 201
curl -X POST "$OS_URL/_bulk" -H "Authorization: Bearer $DEX_TOKEN" \
-H "Content-Type: application/x-ndjson" \
--data-binary $'{"index":{"_index":"logs-k8s-prod-eu-2026.10.04"}}\n{"message":"hello"}\n'
The same token writing to another cluster's index, for example logs-k8s-prod-us-2026.10.04, is denied: the request returns 200, and the item in the response carries "status":403 with a security_exception.
A token that Dex issued for a different client, so with another aud, is rejected with 401 before any role is checked.
6. Deploy Vector in each cluster
Vector reads header values only when it starts, so a sidecar keeps its config in step with the token. It exchanges the pod's ServiceAccount token at Dex every 5 minutes (Dex tokens last 15), renders the Dex token into the Vector config, and Vector reloads when the file changes. The examples are for prod-eu. For another cluster, change the connector id, the index prefix and nothing else.
The ServiceAccount and read access to pod metadata:
apiVersion: v1
kind: Namespace
metadata:
name: logging
---
apiVersion: v1
kind: ServiceAccount
metadata:
name: vector
namespace: logging
---
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRole
metadata:
name: vector-logging
rules:
- apiGroups: [""]
resources: ["pods", "namespaces", "nodes"]
verbs: ["get", "list", "watch"]
---
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRoleBinding
metadata:
name: vector-logging
roleRef:
apiGroup: rbac.authorization.k8s.io
kind: ClusterRole
name: vector-logging
subjects:
- kind: ServiceAccount
name: vector
namespace: logging
The Vector config template and the exchange script. @@TOKEN@@ is replaced with the Dex token on every render:
apiVersion: v1
kind: ConfigMap
metadata:
name: vector
namespace: logging
data:
vector.yaml.tpl: |
data_dir: /vector-data
sources:
logs:
type: kubernetes_logs
extra_field_selector: metadata.namespace=logging
exclude_paths_glob_patterns:
- "/var/log/pods/logging_vector-*/**"
transforms:
limit:
type: throttle
inputs: [logs]
threshold: 50
window_secs: 1
sinks:
opensearch:
type: elasticsearch
inputs: [limit]
endpoints:
- https://logs-123-os.c9t.io
api_version: v7
mode: bulk
bulk:
index: "logs-k8s-prod-eu-%Y.%m.%d"
healthcheck:
enabled: false
request:
headers:
Authorization: "Bearer @@TOKEN@@"
buffer:
type: memory
max_events: 500
render.sh: |
#!/bin/sh
tpl=/tpl/vector.yaml.tpl
sa=/var/run/secrets/dex-exchange/token
out=/conf/vector.yaml
exchange() {
curl -fsS -m 20 https://dex.example.com/token \
--data-urlencode client_id=logs-shipper \
--data-urlencode grant_type=urn:ietf:params:oauth:grant-type:token-exchange \
--data-urlencode connector_id=prod-eu \
--data-urlencode "subject_token=$(cat "$sa")" \
--data-urlencode subject_token_type=urn:ietf:params:oauth:token-type:id_token \
--data-urlencode requested_token_type=urn:ietf:params:oauth:token-type:id_token \
--data-urlencode "scope=openid profile groups" \
| sed -n 's/.*"access_token":"\([^"]*\)".*/\1/p'
}
render() {
t=$(exchange)
[ -n "$t" ] || return 1
sed "s|@@TOKEN@@|$t|" "$tpl" > "$out.tmp" && mv "$out.tmp" "$out"
}
if [ "$1" = once ]; then render; exit $?; fi
while true; do
sleep 300
render || echo "token exchange failed, keeping previous token"
done
- The
exclude_paths_glob_patternsentry keeps Vector from shipping its own logs, which would otherwise feed its own errors back into the pipeline. - The healthcheck is off because the per-cluster role cannot read cluster information.
api_versionis set explicitly for the same reason: Vector would otherwise probe the cluster root. - To collect more than one namespace, widen
extra_field_selectorand add the namespace's ServiceAccounts to the role mapping.
The DaemonSet. The init container renders the first config so Vector never starts without one. If an exchange fails later, the sidecar keeps the previous token and retries after the next interval.
apiVersion: apps/v1
kind: DaemonSet
metadata:
name: vector
namespace: logging
spec:
selector:
matchLabels:
app: vector
template:
metadata:
labels:
app: vector
spec:
serviceAccountName: vector
initContainers:
- name: render-config
image: curlimages/curl:8.10.1
command: ["/bin/sh", "/tpl/render.sh", "once"]
volumeMounts:
- {name: tpl, mountPath: /tpl}
- {name: conf, mountPath: /conf}
- {name: sa-token, mountPath: /var/run/secrets/dex-exchange, readOnly: true}
containers:
- name: vector
image: timberio/vector:0.58.0-alpine
args:
- --config=/conf/vector.yaml
- --watch-config
- --watch-config-method=poll
- --watch-config-poll-interval-seconds=5
env:
- name: VECTOR_SELF_NODE_NAME
valueFrom:
fieldRef: {fieldPath: spec.nodeName}
resources:
requests: {cpu: 20m, memory: 48Mi}
limits: {cpu: 150m, memory: 128Mi}
volumeMounts:
- {name: conf, mountPath: /conf, readOnly: true}
- {name: data, mountPath: /vector-data}
- {name: varlog, mountPath: /var/log, readOnly: true}
- name: token-exchanger
image: curlimages/curl:8.10.1
command: ["/bin/sh", "/tpl/render.sh"]
volumeMounts:
- {name: tpl, mountPath: /tpl}
- {name: conf, mountPath: /conf}
- {name: sa-token, mountPath: /var/run/secrets/dex-exchange, readOnly: true}
volumes:
- name: tpl
configMap: {name: vector, defaultMode: 0555}
- name: conf
emptyDir: {medium: Memory}
- name: data
emptyDir: {}
- name: varlog
hostPath: {path: /var/log}
- name: sa-token
projected:
sources:
- serviceAccountToken:
path: token
audience: dex
expirationSeconds: 600
The rendered config lives on a memory-backed emptyDir, so the Dex token is never written to the node's disk. Check that the sidecar keeps working. Each pod logs one reload at startup and one about every 5 minutes:
kubectl -n logging logs ds/vector -c vector | grep -E 'Configuration file changed|ERROR'
kubectl -n logging logs ds/vector -c token-exchanger
7. Expire old indices
Every cluster writes a new daily index. An index state management (ISM) policy deletes each index once it is old enough. This one deletes after 7 days, and ism_template attaches it automatically to every new logs-*-* index:
- Dashboards
- API
- Terraform
-
In OpenSearch Dashboards, open Management > Index Management > State management policies.
-
Click Create policy and choose the JSON editor.
-
Set the policy ID to
logs-retentionand paste:{"policy": {"description": "Delete daily log indices after 7 days","default_state": "hot","states": [{"name": "hot", "actions": [], "transitions": [{"state_name": "delete", "conditions": {"min_index_age": "7d"}}]},{"name": "delete", "actions": [{"delete": {}}], "transitions": []}],"ism_template": [{"index_patterns": ["logs-*-*"], "priority": 100}]}} -
Click Create.
-
To attach the policy to indices that already exist, open Indices, select them, and choose Actions > Apply policy. The template only applies to indices created after the policy exists.
-
Open Managed indices to check that each index shows
logs-retention.
Change min_index_age to keep logs for longer or shorter.
curl -u admin:$ADMIN_PASSWORD -X PUT "$OS_URL/_plugins/_ism/policies/logs-retention" \
-H "Content-Type: application/json" -d '{
"policy": {
"description": "Delete daily log indices after 7 days",
"default_state": "hot",
"states": [
{"name": "hot", "actions": [], "transitions": [{"state_name": "delete", "conditions": {"min_index_age": "7d"}}]},
{"name": "delete", "actions": [{"delete": {}}], "transitions": []}
],
"ism_template": [{"index_patterns": ["logs-*-*"], "priority": 100}]
}
}'
The template only applies to indices created after the policy exists. Attach it to indices that already exist:
curl -u admin:$ADMIN_PASSWORD -X POST "$OS_URL/_plugins/_ism/add/logs-*-*" \
-H "Content-Type: application/json" -d '{"policy_id": "logs-retention"}'
Check that an index is managed:
curl -u admin:$ADMIN_PASSWORD "$OS_URL/_plugins/_ism/explain/logs-*-*"
Each managed index lists "policy_id": "logs-retention". Change min_index_age to keep logs for longer or shorter.
Add the policy to the configuration from the previous steps. The opensearch provider block is the one from step 4.
resource "opensearch_ism_policy" "logs_retention" {
policy_id = "logs-retention"
body = jsonencode({
policy = {
description = "Delete daily log indices after 7 days"
default_state = "hot"
states = [
{ name = "hot", actions = [], transitions = [{ state_name = "delete", conditions = { min_index_age = "7d" } }] },
{ name = "delete", actions = [{ delete = {} }], transitions = [] },
]
ism_template = [{ index_patterns = ["logs-*-*"], priority = 100 }]
}
})
}
The template only applies to indices created after the policy exists. Apply the policy to existing indices in Dashboards or with the API.
8. Verify delivery
Once a pod writes a log line, the document appears in the cluster's daily index. Each cluster's documents live only in its own pattern:
curl -u admin:$ADMIN_PASSWORD "$OS_URL/logs-k8s-prod-eu-*/_search?size=1&sort=timestamp:desc"
Each document carries the message and Kubernetes metadata (kubernetes.pod_namespace, kubernetes.pod_name, container and node details). In OpenSearch Dashboards, create one index pattern per cluster, for example logs-k8s-prod-eu-*, with timestamp as the time field.
9. Clean up
In every Kubernetes cluster, remove Vector. Deleting the namespace removes the DaemonSet, ConfigMap and ServiceAccount, and the ClusterRole and ClusterRoleBinding are cluster-wide, so they go separately. Only delete the namespace if you created it for this guide:
kubectl delete clusterrolebinding/vector-logging clusterrole/vector-logging
kubectl delete namespace logging
In the cluster that runs Dex, remove Dex, its ConfigMap and its HTTPRoute:
kubectl delete namespace dex
Then delete the OpenSearch cluster. This also deletes its indices, roles and mappings, and the retention policy:
- Console
- API
- Terraform
- Open the cluster in the ClusterNest Console.
- Click Delete.
- Type the cluster name to confirm.
- Click Delete Cluster.
curl -X DELETE "https://api.clusternest.com/cluster/opensearch/$CLUSTER_ID" \
-H "Authorization: Bearer $ACCESS_TOKEN"
rm -f credentials.json
terraform destroy
Part 3: One Dex per cluster
Part 2 puts every cluster behind a single Dex. That Dex is then shared: if it goes down, no cluster can get a new token, and OpenSearch can no longer validate the tokens it issued. This part removes the shared component. Each cluster has its own Dex with its own issuer, and OpenSearch trusts each one through its own auth source. When a cluster, and with it its Dex, goes down, only that cluster's logging stops.
This part is complete on its own. You do not need Part 1 or Part 2.
It uses three example clusters, each with its own Dex:
| Cluster | Dex issuer | Connector id | Group | Index pattern |
|---|---|---|---|---|
prod-eu | https://dex.prod-eu.example.com | prod-eu | prod-eu:system:serviceaccount:logging:vector | logs-k8s-prod-eu-* |
prod-us | https://dex.prod-us.example.com | prod-us | prod-us:system:serviceaccount:logging:vector | logs-k8s-prod-us-* |
staging | https://dex.staging.example.com | staging | staging:system:serviceaccount:logging:vector | logs-k8s-staging-* |
How it works
- A pod projects a ServiceAccount token with audience
dex. - A sidecar posts it to its own cluster's Dex, with that cluster's connector id.
- That Dex verifies the token against its cluster's issuer and returns an ID token whose
groupsclaim isprod-eu:system:serviceaccount:logging:vector. - The forwarder sends the Dex token to OpenSearch. OpenSearch validates it against the issuer of the matching auth source and takes the group as a backend role.
- The role mapped to that group can only write
logs-k8s-prod-eu-*.
Run each Dex in the cluster it serves, so that an outage takes the Dex down together with the cluster. OpenSearch contacts each Dex's issuer to validate its tokens, so each Dex must be reachable over HTTPS from the internet, and each Dex must reach its own cluster's issuer.
What one Dex going down does
With three Dexes configured as separate auth sources, one was stopped while the others kept serving. The results:
- Tokens from the running Dex kept writing to their own index (
201) at the same speed as before, about 0.2 seconds a write. The log shippers using it had no errors and no gap. - The stopped Dex answered
503, so its cluster could not get a new token. A token minted just before the outage kept working while OpenSearch had Dex's keys cached, and was still accepted 5 seconds into the outage. - After the stopped Dex came back, its cluster's tokens were rejected with
401for about 8 minutes, then accepted again. A Dex with in-memory storage generates new signing keys when it starts, and OpenSearch picks them up on its next key fetch, which can take several minutes. The other clusters were not affected. - A token from one Dex could not write the other cluster's index, in either direction (
403).
1. Configure each Dex
Each Dex has one public client and one connector, for its own cluster. The client is public: true, so the exchange needs no client secret. This is the prod-eu Dex:
issuer: https://dex.prod-eu.example.com
storage:
type: memory
web:
http: 0.0.0.0:5556
oauth2:
grantTypes:
- urn:ietf:params:oauth:grant-type:token-exchange
skipApprovalScreen: true
expiry:
idTokens: 15m
staticClients:
- id: logs-shipper
name: Log shippers
public: true
redirectURIs:
- http://localhost/unused
connectors:
- type: oidc
id: prod-eu
name: prod-eu
config:
issuer: https://oidc.prod-eu.example.com
clientID: dex
clientSecret: unused
redirectURI: https://dex.prod-eu.example.com/callback
scopes:
- openid
userNameKey: sub
insecureSkipEmailVerified: true
claimModifications:
newGroupFromClaims:
- claims:
- sub
delimiter: ":"
prefix: "prod-eu"
For prod-us and staging, change the three values that name the cluster: issuer and redirectURI (https://dex.prod-us.example.com), the connector's id, name and prefix, and the connector's issuer (https://oidc.prod-us.example.com).
Find a cluster's own issuer by decoding any of its ServiceAccount tokens and reading the iss claim:
kubectl create token default --duration=10m \
| cut -d. -f2 | tr '_-' '/+' \
| awk '{while (length($0) % 4) $0 = $0 "="; print}' | base64 --decode | jq -r .iss
clientID: dexon the connector is the audience Dex requires on the incoming ServiceAccount token. Project tokens withaudience: dex.userNameKey: submakes the Dexnameclaim the ServiceAccount identity.scopes: [openid]withinsecureSkipEmailVerified: trueis needed because ServiceAccount tokens carry no email. The connector'sclientSecretis a required field that is never used.newGroupFromClaimsbuilds the group fromsub. Withprefix: "prod-eu"anddelimiter: ":"the group readsprod-eu:system:serviceaccount:logging:vector. A prefix that already ends in a colon produces a doubled colon.- The exchange request must carry
client_id=logs-shipper. A request that names no client is rejected. - With
storage: type: memory, a Dex generates new signing keys every time it restarts, and its cluster's tokens are rejected for several minutes afterwards (about 8 in testing). Use a persistent storage backend so the keys survive restarts.
2. Deploy Dex
Save the config from step 1 as config.yaml, then run Dex in the cluster whose logs it serves, and publish it at that cluster's own hostname. Repeat for every cluster, each with its own config.
kubectl create namespace logging
kubectl -n logging create configmap dex --from-file=config.yaml
A Deployment and a Service. The readiness probe uses the discovery document:
apiVersion: apps/v1
kind: Deployment
metadata:
name: dex
namespace: logging
spec:
replicas: 1
selector:
matchLabels:
app: dex
template:
metadata:
labels:
app: dex
spec:
containers:
- name: dex
image: ghcr.io/dexidp/dex:v2.45.1
args: ["dex", "serve", "/etc/dex/config.yaml"]
ports:
- {name: http, containerPort: 5556}
readinessProbe:
httpGet: {path: /.well-known/openid-configuration, port: 5556}
resources:
requests: {cpu: 10m, memory: 32Mi}
limits: {cpu: 100m, memory: 96Mi}
volumeMounts:
- {name: config, mountPath: /etc/dex}
volumes:
- name: config
configMap: {name: dex}
---
apiVersion: v1
kind: Service
metadata:
name: dex
namespace: logging
spec:
selector:
app: dex
ports:
- {name: http, port: 5556, targetPort: 5556}
Publish the Service at the issuer hostname over HTTPS, with a certificate that matches it. This example uses a Gateway API HTTPRoute on an existing Gateway with an https listener. If you use an Ingress instead, route dex.prod-eu.example.com to the dex Service on port 5556:
apiVersion: gateway.networking.k8s.io/v1
kind: HTTPRoute
metadata:
name: dex
namespace: logging
spec:
hostnames:
- dex.prod-eu.example.com
parentRefs:
- group: gateway.networking.k8s.io
kind: Gateway
name: gateway
namespace: gateway-system
sectionName: https
rules:
- backendRefs:
- {group: "", kind: Service, name: dex, port: 5556}
matches:
- path: {type: PathPrefix, value: /}
Check that Dex answers from the internet:
curl https://dex.prod-eu.example.com/.well-known/openid-configuration
The response lists "issuer": "https://dex.prod-eu.example.com" and a jwks_uri of https://dex.prod-eu.example.com/keys. A 503 no healthy upstream just after the deploy means the pod is not Ready yet. It clears within seconds once the readiness probe passes. To change the config later, update the ConfigMap and restart the Deployment. Read the storage warning in step 1 before restarting Dex.
3. Create the OpenSearch cluster with one auth source per Dex
Add one OIDC auth source for each Dex. Give each a unique name. subject_key is name, roles_key is groups, and client_id and audience are the Dex client. With audience set, OpenSearch rejects a token Dex issued for any other client:
- Console
- API
- Terraform
- Open the ClusterNest Console, go to OpenSearch under Services and click Launch Cluster.
- Enter a name and pick a tier.
- Under Authentication, click Add OIDC source once for each Dex and fill in:
- Name:
dex-prod-eu - Connect URL:
https://dex.prod-eu.example.com/.well-known/openid-configuration - Client ID:
logs-shipper - Audience:
logs-shipper - Subject Key:
name - Roles Key:
groups - Leave Offer on OpenSearch Dashboards off.
- Name:
- Click Create Cluster. The Console shows the admin credentials once, so save them.
- Wait until the cluster's state is
available.
Change the name and the URL for each Dex. ClusterNest rejects an auth source whose host is not publicly resolvable.
To add a cluster later:
- Open the cluster in the Console and click Edit.
- Click Add OIDC source and fill it in the same way.
- Save, and wait for the state to go from
updatingback toavailable. The change takes a few minutes.
Every request to the ClusterNest API carries an access token. Generate an app password from the MFA settings page of your account (your regular password does not work), then exchange it for a token:
export ACCESS_TOKEN=$(curl -sS -X POST https://api.clusternest.com/auth/token \
-H "Content-Type: application/json" \
-d '{"email": "[email protected]", "password": "<app password>"}' \
| jq -r .access_token)
The token expires, so run this again when requests start returning 401. See Get an access token for the full request and response.
curl -X POST "https://api.clusternest.com/cluster/opensearch/" \
-H "Authorization: Bearer $ACCESS_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"name": "logs",
"organization_id": 123,
"tier": "basic",
"auth_sources": [
{
"type": "oidc",
"name": "dex-prod-eu",
"connect_url": "https://dex.prod-eu.example.com/.well-known/openid-configuration",
"client_id": "logs-shipper",
"audience": "logs-shipper",
"subject_key": "name",
"roles_key": "groups",
"dashboards_login": false
},
{
"type": "oidc",
"name": "dex-prod-us",
"connect_url": "https://dex.prod-us.example.com/.well-known/openid-configuration",
"client_id": "logs-shipper",
"audience": "logs-shipper",
"subject_key": "name",
"roles_key": "groups",
"dashboards_login": false
},
{
"type": "oidc",
"name": "dex-staging",
"connect_url": "https://dex.staging.example.com/.well-known/openid-configuration",
"client_id": "logs-shipper",
"audience": "logs-shipper",
"subject_key": "name",
"roles_key": "groups",
"dashboards_login": false
}
]
}'
ClusterNest rejects an auth source whose host is not publicly resolvable. Poll GET /cluster/opensearch/$CLUSTER_ID until state is available. The API endpoint is https://<name>-<organization_id>-os.c9t.io.
To add a cluster later, read the cluster, append its auth source, and write the result back. A PUT replaces the whole cluster, so send the merged result:
curl "https://api.clusternest.com/cluster/opensearch/$CLUSTER_ID" \
-H "Authorization: Bearer $ACCESS_TOKEN" \
| jq '.auth_sources += [{
"type": "oidc",
"name": "dex-prod-asia",
"connect_url": "https://dex.prod-asia.example.com/.well-known/openid-configuration",
"client_id": "logs-shipper",
"audience": "logs-shipper",
"subject_key": "name",
"roles_key": "groups",
"dashboards_login": false
}]' > cluster.json
curl -X PUT "https://api.clusternest.com/cluster/opensearch/$CLUSTER_ID" \
-H "Authorization: Bearer $ACCESS_TOKEN" \
-H "Content-Type: application/json" \
--data-binary @cluster.json
terraform {
required_providers {
clusternest = {
source = "tf.clusternest.com/clusternest/clusternest"
version = ">=1.1.0"
}
opensearch = {
source = "opensearch-project/opensearch"
version = ">= 2.2.0"
}
}
}
provider "clusternest" {}
locals {
clusters = ["prod-eu", "prod-us", "staging"]
}
resource "clusternest_opensearch" "logs" {
name = "logs"
tier = "basic"
organization_id = 123
auth_sources = [
for cluster in local.clusters : {
name = "dex-${cluster}"
type = "oidc"
connect_url = "https://dex.${cluster}.example.com/.well-known/openid-configuration"
client_id = "logs-shipper"
audience = "logs-shipper"
subject_key = "name"
roles_key = "groups"
dashboards_login = false
}
]
}
output "opensearch_url" {
value = clusternest_opensearch.logs.url
}
The provider reads your email and app password from CLUSTERNEST_EMAIL and CLUSTERNEST_APP_PASSWORD. terraform apply returns once the cluster is available, and opensearch_url is the API endpoint. To add a cluster later, add its name to local.clusters and apply again.
The cluster is updating for a few minutes while the change applies. Adding the second Dex to a cluster whose log shippers were already writing through the first did not interrupt them: they kept delivering at a steady rate with no gap.
4. Create one role per cluster
Each role can write only its cluster's indices. Use the OpenSearch security API with the cluster's admin credential, and map each role by backend role, which is the group Dex emits.
Do not map users for these ServiceAccounts. The username (system:serviceaccount:logging:vector) is the same in every cluster, so a username mapping on any role grants that role's access to all of them.
- Dashboards
- API
- Terraform
Repeat steps 2 to 7 for each cluster (prod-eu, prod-us, staging), replacing <cluster>.
- Sign in to OpenSearch Dashboards as
admin, with the credential the Console showed when it created the cluster. - Open Security > Roles and click Create role. Name it
logs_writer_<cluster>. - Under Cluster permissions, add
indices:data/write/bulk. - Under Index permissions, set the index pattern to
logs-k8s-<cluster>-*and the permissions tocreate_index,indices:data/write/bulk*andindices:data/write/index. Click Create. - Open the role's Mapped users tab and click Manage mapping.
- Under Backend roles, add
<cluster>:system:serviceaccount:logging:vector. - Click Map.
To allow another ServiceAccount from one cluster, add its group (prod-eu:system:serviceaccount:logging:other) as another backend role on that cluster's mapping.
Fetch the admin credential from the ClusterNest API into a file, so the password is never printed:
(umask 077; curl -sS --fail-with-body \
-H "Authorization: Bearer $ACCESS_TOKEN" \
"https://api.clusternest.com/cluster/opensearch/$CLUSTER_ID/credentials" \
-o credentials.json)
export ADMIN_PASSWORD=$(jq -r .password credentials.json)
export OS_URL=https://logs-123-os.c9t.io
for cluster in prod-eu prod-us staging; do
curl -u admin:$ADMIN_PASSWORD -X PUT "$OS_URL/_plugins/_security/api/roles/logs_writer_$cluster" \
-H "Content-Type: application/json" -d '{
"cluster_permissions": ["indices:data/write/bulk"],
"index_permissions": [{
"index_patterns": ["logs-k8s-'"$cluster"'-*"],
"allowed_actions": ["create_index", "indices:data/write/bulk*", "indices:data/write/index"]
}]
}'
curl -u admin:$ADMIN_PASSWORD -X PUT "$OS_URL/_plugins/_security/api/rolesmapping/logs_writer_$cluster" \
-H "Content-Type: application/json" \
-d '{"backend_roles": ["'"$cluster"':system:serviceaccount:logging:vector"]}'
done
To allow another ServiceAccount from one cluster, add its group to that cluster's mapping:
curl -u admin:$ADMIN_PASSWORD -X PATCH "$OS_URL/_plugins/_security/api/rolesmapping/logs_writer_prod-eu" \
-H "Content-Type: application/json" \
-d '[{"op": "add", "path": "/backend_roles/-", "value": "prod-eu:system:serviceaccount:logging:other"}]'
The opensearch provider signs in with the cluster's admin credential, which the clusternest_opensearch_credentials data source reads. Add this to the configuration from step 3:
data "clusternest_opensearch_credentials" "logs" {
cluster_id = clusternest_opensearch.logs.id
}
provider "opensearch" {
url = clusternest_opensearch.logs.url
username = data.clusternest_opensearch_credentials.logs.username
password = data.clusternest_opensearch_credentials.logs.password
}
resource "opensearch_role" "logs_writer" {
for_each = toset(local.clusters)
role_name = "logs_writer_${each.key}"
cluster_permissions = ["indices:data/write/bulk"]
index_permissions {
index_patterns = ["logs-k8s-${each.key}-*"]
allowed_actions = ["create_index", "indices:data/write/bulk*", "indices:data/write/index"]
}
}
resource "opensearch_roles_mapping" "logs_writer" {
for_each = toset(local.clusters)
role_name = opensearch_role.logs_writer[each.key].role_name
backend_roles = ["${each.key}:system:serviceaccount:logging:vector"]
}
To allow another ServiceAccount from one cluster, add its group to that cluster's backend_roles.
5. Check the exchange
Mint a token for the ServiceAccount, exchange it, and read the claims of the result:
SA_TOKEN=$(kubectl create token vector -n logging --audience dex --duration=10m)
curl https://dex.prod-eu.example.com/token \
-d client_id=logs-shipper \
-d grant_type=urn:ietf:params:oauth:grant-type:token-exchange \
-d connector_id=prod-eu \
--data-urlencode "subject_token=$SA_TOKEN" \
-d subject_token_type=urn:ietf:params:oauth:token-type:id_token \
-d requested_token_type=urn:ietf:params:oauth:token-type:id_token \
--data-urlencode "scope=openid profile groups"
The access_token field of the response holds a JWT. Its claims include:
{
"iss": "https://dex.prod-eu.example.com",
"aud": "logs-shipper",
"name": "system:serviceaccount:logging:vector",
"groups": ["prod-eu:system:serviceaccount:logging:vector"]
}
requested_token_type must be id_token to get a JWT that OpenSearch can validate. The token expires after expiry.idTokens.
Use it as a Bearer token. The bulk API answers 200 and reports each write in the response items. A write to the cluster's own index succeeds, a write to another cluster's index is denied, and reads are denied because the role is write-only:
# own index: the item status is 201
curl -X POST "$OS_URL/_bulk" -H "Authorization: Bearer $DEX_TOKEN" \
-H "Content-Type: application/x-ndjson" \
--data-binary $'{"index":{"_index":"logs-k8s-prod-eu-2026.10.04"}}\n{"message":"hello"}\n'
The same token writing to another cluster's index, for example logs-k8s-prod-us-2026.10.04, is denied: the request returns 200, and the item in the response carries "status":403 with a security_exception.
6. Deploy Vector in each cluster
Vector reads header values only when it starts, so a sidecar keeps its config in step with the token. It exchanges the pod's ServiceAccount token at Dex every 5 minutes (Dex tokens last 15), renders the Dex token into the Vector config, and Vector reloads when the file changes. The examples are for prod-eu. For another cluster, change the connector id, the index prefix and nothing else.
The ServiceAccount and read access to pod metadata. The logging namespace already exists from step 2, where Dex runs:
apiVersion: v1
kind: ServiceAccount
metadata:
name: vector
namespace: logging
---
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRole
metadata:
name: vector-logging
rules:
- apiGroups: [""]
resources: ["pods", "namespaces", "nodes"]
verbs: ["get", "list", "watch"]
---
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRoleBinding
metadata:
name: vector-logging
roleRef:
apiGroup: rbac.authorization.k8s.io
kind: ClusterRole
name: vector-logging
subjects:
- kind: ServiceAccount
name: vector
namespace: logging
The Vector config template and the exchange script. @@TOKEN@@ is replaced with the Dex token on every render:
apiVersion: v1
kind: ConfigMap
metadata:
name: vector
namespace: logging
data:
vector.yaml.tpl: |
data_dir: /vector-data
sources:
logs:
type: kubernetes_logs
extra_field_selector: metadata.namespace=logging
exclude_paths_glob_patterns:
- "/var/log/pods/logging_vector-*/**"
transforms:
limit:
type: throttle
inputs: [logs]
threshold: 50
window_secs: 1
sinks:
opensearch:
type: elasticsearch
inputs: [limit]
endpoints:
- https://logs-123-os.c9t.io
api_version: v7
mode: bulk
bulk:
index: "logs-k8s-prod-eu-%Y.%m.%d"
healthcheck:
enabled: false
request:
headers:
Authorization: "Bearer @@TOKEN@@"
buffer:
type: memory
max_events: 500
render.sh: |
#!/bin/sh
tpl=/tpl/vector.yaml.tpl
sa=/var/run/secrets/dex-exchange/token
out=/conf/vector.yaml
exchange() {
curl -fsS -m 20 https://dex.prod-eu.example.com/token \
--data-urlencode client_id=logs-shipper \
--data-urlencode grant_type=urn:ietf:params:oauth:grant-type:token-exchange \
--data-urlencode connector_id=prod-eu \
--data-urlencode "subject_token=$(cat "$sa")" \
--data-urlencode subject_token_type=urn:ietf:params:oauth:token-type:id_token \
--data-urlencode requested_token_type=urn:ietf:params:oauth:token-type:id_token \
--data-urlencode "scope=openid profile groups" \
| sed -n 's/.*"access_token":"\([^"]*\)".*/\1/p'
}
render() {
t=$(exchange)
[ -n "$t" ] || return 1
sed "s|@@TOKEN@@|$t|" "$tpl" > "$out.tmp" && mv "$out.tmp" "$out"
}
if [ "$1" = once ]; then render; exit $?; fi
while true; do
sleep 300
render || echo "token exchange failed, keeping previous token"
done
- The
exclude_paths_glob_patternsentry keeps Vector from shipping its own logs, which would otherwise feed its own errors back into the pipeline. - Dex runs in the same
loggingnamespace, so its logs are shipped too. Add"/var/log/pods/logging_dex-*/**"toexclude_paths_glob_patternsto leave them out. - The healthcheck is off because the per-cluster role cannot read cluster information.
api_versionis set explicitly for the same reason: Vector would otherwise probe the cluster root. - To collect more than one namespace, widen
extra_field_selectorand add the namespace's ServiceAccounts to the role mapping.
The DaemonSet. The init container renders the first config so Vector never starts without one. If an exchange fails later, the sidecar keeps the previous token and retries after the next interval.
apiVersion: apps/v1
kind: DaemonSet
metadata:
name: vector
namespace: logging
spec:
selector:
matchLabels:
app: vector
template:
metadata:
labels:
app: vector
spec:
serviceAccountName: vector
initContainers:
- name: render-config
image: curlimages/curl:8.10.1
command: ["/bin/sh", "/tpl/render.sh", "once"]
volumeMounts:
- {name: tpl, mountPath: /tpl}
- {name: conf, mountPath: /conf}
- {name: sa-token, mountPath: /var/run/secrets/dex-exchange, readOnly: true}
containers:
- name: vector
image: timberio/vector:0.58.0-alpine
args:
- --config=/conf/vector.yaml
- --watch-config
- --watch-config-method=poll
- --watch-config-poll-interval-seconds=5
env:
- name: VECTOR_SELF_NODE_NAME
valueFrom:
fieldRef: {fieldPath: spec.nodeName}
resources:
requests: {cpu: 20m, memory: 48Mi}
limits: {cpu: 150m, memory: 128Mi}
volumeMounts:
- {name: conf, mountPath: /conf, readOnly: true}
- {name: data, mountPath: /vector-data}
- {name: varlog, mountPath: /var/log, readOnly: true}
- name: token-exchanger
image: curlimages/curl:8.10.1
command: ["/bin/sh", "/tpl/render.sh"]
volumeMounts:
- {name: tpl, mountPath: /tpl}
- {name: conf, mountPath: /conf}
- {name: sa-token, mountPath: /var/run/secrets/dex-exchange, readOnly: true}
volumes:
- name: tpl
configMap: {name: vector, defaultMode: 0555}
- name: conf
emptyDir: {medium: Memory}
- name: data
emptyDir: {}
- name: varlog
hostPath: {path: /var/log}
- name: sa-token
projected:
sources:
- serviceAccountToken:
path: token
audience: dex
expirationSeconds: 600
The rendered config lives on a memory-backed emptyDir, so the Dex token is never written to the node's disk. Check that the sidecar keeps working. Each pod logs one reload at startup and one about every 5 minutes:
kubectl -n logging logs ds/vector -c vector | grep -E 'Configuration file changed|ERROR'
kubectl -n logging logs ds/vector -c token-exchanger
7. Expire old indices
Every cluster writes a new daily index. An index state management (ISM) policy deletes each index once it is old enough. This one deletes after 7 days, and ism_template attaches it automatically to every new logs-*-* index:
- Dashboards
- API
- Terraform
-
In OpenSearch Dashboards, open Management > Index Management > State management policies.
-
Click Create policy and choose the JSON editor.
-
Set the policy ID to
logs-retentionand paste:{"policy": {"description": "Delete daily log indices after 7 days","default_state": "hot","states": [{"name": "hot", "actions": [], "transitions": [{"state_name": "delete", "conditions": {"min_index_age": "7d"}}]},{"name": "delete", "actions": [{"delete": {}}], "transitions": []}],"ism_template": [{"index_patterns": ["logs-*-*"], "priority": 100}]}} -
Click Create.
-
To attach the policy to indices that already exist, open Indices, select them, and choose Actions > Apply policy. The template only applies to indices created after the policy exists.
-
Open Managed indices to check that each index shows
logs-retention.
Change min_index_age to keep logs for longer or shorter.
curl -u admin:$ADMIN_PASSWORD -X PUT "$OS_URL/_plugins/_ism/policies/logs-retention" \
-H "Content-Type: application/json" -d '{
"policy": {
"description": "Delete daily log indices after 7 days",
"default_state": "hot",
"states": [
{"name": "hot", "actions": [], "transitions": [{"state_name": "delete", "conditions": {"min_index_age": "7d"}}]},
{"name": "delete", "actions": [{"delete": {}}], "transitions": []}
],
"ism_template": [{"index_patterns": ["logs-*-*"], "priority": 100}]
}
}'
The template only applies to indices created after the policy exists. Attach it to indices that already exist:
curl -u admin:$ADMIN_PASSWORD -X POST "$OS_URL/_plugins/_ism/add/logs-*-*" \
-H "Content-Type: application/json" -d '{"policy_id": "logs-retention"}'
Check that an index is managed:
curl -u admin:$ADMIN_PASSWORD "$OS_URL/_plugins/_ism/explain/logs-*-*"
Each managed index lists "policy_id": "logs-retention". Change min_index_age to keep logs for longer or shorter.
Add the policy to the configuration from the previous steps. The opensearch provider block is the one from step 4.
resource "opensearch_ism_policy" "logs_retention" {
policy_id = "logs-retention"
body = jsonencode({
policy = {
description = "Delete daily log indices after 7 days"
default_state = "hot"
states = [
{ name = "hot", actions = [], transitions = [{ state_name = "delete", conditions = { min_index_age = "7d" } }] },
{ name = "delete", actions = [{ delete = {} }], transitions = [] },
]
ism_template = [{ index_patterns = ["logs-*-*"], priority = 100 }]
}
})
}
The template only applies to indices created after the policy exists. Apply the policy to existing indices in Dashboards or with the API.
8. Verify delivery
Once a pod writes a log line, the document appears in the cluster's daily index. Each cluster's documents live only in its own pattern:
curl -u admin:$ADMIN_PASSWORD "$OS_URL/logs-k8s-prod-eu-*/_search?size=1&sort=timestamp:desc"
Each document carries the message and Kubernetes metadata (kubernetes.pod_namespace, kubernetes.pod_name, container and node details). In OpenSearch Dashboards, create one index pattern per cluster, for example logs-k8s-prod-eu-*, with timestamp as the time field.
9. Clean up
In every Kubernetes cluster, remove Vector and the Dex that runs there. Deleting the logging namespace removes the Dex Deployment, Service, ConfigMap and HTTPRoute together with the Vector DaemonSet, ConfigMap and ServiceAccount. The ClusterRole and ClusterRoleBinding are cluster-wide, so they go separately. Only delete the namespace if you created it for this guide:
kubectl delete clusterrolebinding/vector-logging clusterrole/vector-logging
kubectl delete namespace logging
Then delete the OpenSearch cluster. This also deletes its indices, roles and mappings, and the retention policy:
- Console
- API
- Terraform
- Open the cluster in the ClusterNest Console.
- Click Delete.
- Type the cluster name to confirm.
- Click Delete Cluster.
curl -X DELETE "https://api.clusternest.com/cluster/opensearch/$CLUSTER_ID" \
-H "Authorization: Bearer $ACCESS_TOKEN"
rm -f credentials.json
terraform destroy
Part 4: Let people sign in with SAML
The earlier parts let workloads write logs. This part lets people read them: users sign in to OpenSearch Dashboards through a SAML identity provider and see only the indices their group is allowed to read. It is complete on its own, so it works with any cluster that has logs in daily indices, whichever way they got there. The examples use per-cluster index patterns such as logs-k8s-prod-eu-* and Authentik as the identity provider.
A cluster accepts one SAML source, and it is always offered on the Dashboards login page as Login with followed by the source's name. It sits next to any other auth source, such as the OIDC source for Dex, and neither affects the other.
How it works
- A user opens Dashboards and clicks Login with authentik.
- The identity provider authenticates them and returns a signed SAML assertion that carries their username and groups.
- OpenSearch takes the username from
subject_keyand the groups fromroles_keyas backend roles. - A role mapping turns each group into a read-only role for one cluster's indices.
1. Register the application in your identity provider
You need the cluster's Dashboards address, https://<name>-<organization_id>-osd.c9t.io (the opensearch_dashboards_url field of the cluster). The settings OpenSearch requires of any identity provider:
| Setting | Value |
|---|---|
| ACS URL | https://<dashboards-hostname>/_opendistro/_security/saml/acs |
| Audience | the cluster source's sp_entity_id, here opensearch-dashboards |
| Binding | POST |
| Signing | sign both the assertion and the response |
| Attributes | the username, and the groups |
In Authentik, create a SAML Provider with those values: set Service Provider Binding to Post, pick a Signing Certificate and enable Sign assertions and Sign responses, and attach the username, email, name and groups property mappings. Then create an Application that uses the provider.
Unsigned assertions are rejected, and a provider left on the redirect binding returns the response as a GET, which Dashboards answers with a 401.
Next, bind the groups that may sign in to the application, under Policy / Group / User Bindings. An Authentik application with no bindings can refuse every user, and the sign-in then ends on Authentik's own Permission denied / Request has been denied page without ever reaching OpenSearch. Bind exactly the groups you map in step 3.
Other identity providers follow the same settings. See the SSO overview for provider-specific guides.
2. Add the SAML source to the cluster
idp_entity_id must match the entity ID in the identity provider's metadata exactly. For Authentik it is https://<authentik-host>/application/saml/<application-slug>/metadata/. subject_key is the attribute that holds the username, and roles_key is the attribute that holds the groups. For Authentik the defaults are:
subject_key:http://schemas.goauthentik.io/2021/02/saml/usernameroles_key:http://schemas.xmlsoap.org/claims/Group
- Console
- API
- Terraform
- Open the cluster in the ClusterNest Console and click Edit.
- Under Authentication, click Add SAML source and fill in:
- Name:
authentik(it becomes the text of the Dashboards button, "Login with" followed by the name) - IDP Metadata URL:
https://auth.example.com/api/v3/providers/saml/12/metadata/?download - IDP Entity ID:
https://auth.example.com/application/saml/opensearch/metadata/ - SP Entity ID:
opensearch-dashboards - Subject Key:
http://schemas.goauthentik.io/2021/02/saml/username - Roles Key:
http://schemas.xmlsoap.org/claims/Group
- Name:
- Save.
- Wait for the state to go from
updatingback toavailable. The change takes a few minutes.
Every request to the ClusterNest API carries an access token. Generate an app password from the MFA settings page of your account (your regular password does not work), then exchange it for a token:
export ACCESS_TOKEN=$(curl -sS -X POST https://api.clusternest.com/auth/token \
-H "Content-Type: application/json" \
-d '{"email": "[email protected]", "password": "<app password>"}' \
| jq -r .access_token)
The token expires, so run this again when requests start returning 401. See Get an access token for the full request and response.
Read the cluster, add the source, and write the result back. A PUT replaces the whole cluster, so send the merged result:
curl "https://api.clusternest.com/cluster/opensearch/$CLUSTER_ID" \
-H "Authorization: Bearer $ACCESS_TOKEN" \
| jq '.auth_sources += [{
"type": "saml",
"name": "authentik",
"idp_metadata_url": "https://auth.example.com/api/v3/providers/saml/12/metadata/?download",
"idp_entity_id": "https://auth.example.com/application/saml/opensearch/metadata/",
"sp_entity_id": "opensearch-dashboards",
"subject_key": "http://schemas.goauthentik.io/2021/02/saml/username",
"roles_key": "http://schemas.xmlsoap.org/claims/Group"
}]' > cluster.json
curl -X PUT "https://api.clusternest.com/cluster/opensearch/$CLUSTER_ID" \
-H "Authorization: Bearer $ACCESS_TOKEN" \
-H "Content-Type: application/json" \
--data-binary @cluster.json
The cluster is updating for a few minutes while the change applies. Poll GET /cluster/opensearch/$CLUSTER_ID until state is available.
Add the source to the auth_sources list of the cluster's clusternest_opensearch resource, next to any sources it already has:
resource "clusternest_opensearch" "logs" {
name = "logs"
tier = "basic"
organization_id = 123
auth_sources = [
{
name = "authentik"
type = "saml"
idp_metadata_url = "https://auth.example.com/api/v3/providers/saml/12/metadata/?download"
idp_entity_id = "https://auth.example.com/application/saml/opensearch/metadata/"
sp_entity_id = "opensearch-dashboards"
subject_key = "http://schemas.goauthentik.io/2021/02/saml/username"
roles_key = "http://schemas.xmlsoap.org/claims/Group"
},
]
}
terraform apply returns once the cluster is available again.
3. Map each group to a read-only role
A user who signs in has no permissions until you map their groups to roles. Create one reader role per cluster, limited to that cluster's index pattern, and map an identity provider group to it. This example uses a group logs-prod-eu:
- Dashboards
- API
- Terraform
Repeat steps 2 to 7 for each cluster's group.
- Sign in to OpenSearch Dashboards as
admin. - Open Security > Roles and click Create role. Name it
logs_reader_prod-eu. - Under Cluster permissions, add
cluster_composite_ops_ro. - Under Index permissions, set the index pattern to
logs-k8s-prod-eu-*and the permissions toread,indices:admin/mappings/getandindices:admin/resolve/index. Click Create. - Open the role's Mapped users tab and click Manage mapping.
- Under Backend roles, add
logs-prod-euand click Map. - Open Security > Roles, choose
kibana_user, open Mapped users > Manage mapping, addlogs-prod-euas a backend role and click Map.
The kibana_user mapping gives the group the basic Dashboards access that every user needs. When several groups share it, list all of them in that one mapping.
Fetch the admin credential from the ClusterNest API into a file, so the password is never printed:
(umask 077; curl -sS --fail-with-body \
-H "Authorization: Bearer $ACCESS_TOKEN" \
"https://api.clusternest.com/cluster/opensearch/$CLUSTER_ID/credentials" \
-o credentials.json)
export ADMIN_PASSWORD=$(jq -r .password credentials.json)
export OS_URL=https://logs-123-os.c9t.io
curl -u admin:$ADMIN_PASSWORD -X PUT "$OS_URL/_plugins/_security/api/roles/logs_reader_prod-eu" \
-H "Content-Type: application/json" -d '{
"cluster_permissions": ["cluster_composite_ops_ro"],
"index_permissions": [{
"index_patterns": ["logs-k8s-prod-eu-*"],
"allowed_actions": ["read", "indices:admin/mappings/get", "indices:admin/resolve/index"]
}]
}'
curl -u admin:$ADMIN_PASSWORD -X PUT "$OS_URL/_plugins/_security/api/rolesmapping/logs_reader_prod-eu" \
-H "Content-Type: application/json" \
-d '{"backend_roles": ["logs-prod-eu"]}'
curl -u admin:$ADMIN_PASSWORD -X PUT "$OS_URL/_plugins/_security/api/rolesmapping/kibana_user" \
-H "Content-Type: application/json" \
-d '{"backend_roles": ["logs-prod-eu"]}'
The kibana_user mapping gives the group the basic Dashboards access that every user needs. If you map several groups to it, send all of them in one request: PUT replaces the list, and PATCH appends to it.
terraform {
required_providers {
clusternest = {
source = "tf.clusternest.com/clusternest/clusternest"
version = ">=1.1.0"
}
opensearch = {
source = "opensearch-project/opensearch"
version = ">= 2.2.0"
}
}
}
provider "clusternest" {}
variable "cluster_id" {
type = number
}
variable "opensearch_url" {
type = string
}
locals {
clusters = ["prod-eu", "prod-us", "staging"]
}
data "clusternest_opensearch_credentials" "logs" {
cluster_id = var.cluster_id
}
provider "opensearch" {
url = var.opensearch_url
username = data.clusternest_opensearch_credentials.logs.username
password = data.clusternest_opensearch_credentials.logs.password
}
resource "opensearch_role" "logs_reader" {
for_each = toset(local.clusters)
role_name = "logs_reader_${each.key}"
cluster_permissions = ["cluster_composite_ops_ro"]
index_permissions {
index_patterns = ["logs-k8s-${each.key}-*"]
allowed_actions = ["read", "indices:admin/mappings/get", "indices:admin/resolve/index"]
}
}
resource "opensearch_roles_mapping" "logs_reader" {
for_each = toset(local.clusters)
role_name = opensearch_role.logs_reader[each.key].role_name
backend_roles = ["logs-${each.key}"]
}
resource "opensearch_roles_mapping" "kibana_user" {
role_name = "kibana_user"
backend_roles = [for cluster in local.clusters : "logs-${cluster}"]
}
The kibana_user mapping gives the groups the basic Dashboards access that every user needs, and it holds all of them in one list.
Repeat the role and its mapping with prod-us and staging for the other clusters, so a person in logs-prod-eu can read prod-eu logs and nothing else.
4. Sign in and check
Open the Dashboards address and click Login with authentik. After signing in, the user has the roles kibana_user and logs_reader_prod-eu (Dashboards also adds own_index, a built-in role for each user's own tenant).
Query the logs in Dev Tools (Management > Dev Tools). It sends each request as the signed-in user, so there is no cookie or password to handle. Read the newest documents of the cluster's own indices:
GET logs-k8s-prod-eu-*/_search
{
"size": 5,
"sort": [{"timestamp": "desc"}]
}
Everything outside the role's index pattern is denied with a security_exception, whether it is another cluster's logs, the audit log, or the security index:
GET logs-k8s-prod-us-*/_count
GET security-auditlog-*/_count
GET .opendistro_security/_count
To browse the logs, create an index pattern named exactly logs-k8s-prod-eu-*, with timestamp as the time field, and open Discover. An index pattern such as * fails, because it includes indices the role cannot see. Always create patterns that match the role's pattern.
Remove access
To remove someone's access, take them out of the group in the identity provider. To remove a group's access, unbind it from the application (the sign-in is then refused before it reaches OpenSearch) and delete its mapping:
- Dashboards
- API
- Terraform
- Sign in to OpenSearch Dashboards as
admin. - Open Security > Roles > logs_reader_prod-eu > Mapped users and click Manage mapping.
- Remove the group from Backend roles and click Map.
- Open Security > Roles > kibana_user > Mapped users, click Manage mapping, remove the group there too and click Map.
curl -u admin:$ADMIN_PASSWORD -X DELETE "$OS_URL/_plugins/_security/api/rolesmapping/logs_reader_prod-eu"
Remove the group from local.clusters and apply. Terraform deletes its role and mapping and takes the group out of the kibana_user list.