Operations
How to deploy, monitor, back up and upgrade KubeGlass in a cluster. The commands assume the Helm chart installed as release kubeglass in namespace kubeglass, so the Deployment, Service and ConfigMap are all called kubeglass. Adjust the names if yours differ.
The configuration reference lists every setting, and the chart’s values.yaml lists every chart option.
Deployment
Section titled “Deployment”Helm install
Section titled “Helm install”The chart is published as an OCI artifact at oci://ghcr.io/kubeglass/charts/kubeglass; there is no helm repo add step.
Behind an authenticating reverse proxy (the default impersonation mode), set the proxy’s CIDR. The install fails without it:
helm install kubeglass oci://ghcr.io/kubeglass/charts/kubeglass \ --namespace kubeglass \ --create-namespace \ --set config.trustedProxies=10.0.0.0/8With OIDC, the chart has no dedicated keys for the issuer and audience (the values schema rejects unknown keys), so pass them with extraEnv. Save this as values-oidc.yaml:
config: authMode: oidcextraEnv: - name: KUBEGLASS_OIDC_ISSUER_URL value: https://accounts.google.com - name: KUBEGLASS_OIDC_AUDIENCE value: <client-id> - name: KUBEGLASS_ENV value: productionhelm install kubeglass oci://ghcr.io/kubeglass/charts/kubeglass \ --namespace kubeglass \ --create-namespace \ -f values-oidc.yamlTo try it without a proxy or identity provider, use --set config.authMode=local and reach it with kubectl -n kubeglass port-forward svc/kubeglass 8090:8090. In local mode anyone who can reach the Service gets KubeGlass’s own access, Secrets included, so don’t expose it any other way.
deploy/helm/kubeglass/values-production.yaml is a starting point for a production install: ingress with TLS, larger resources, a scoped impersonation rule and a NetworkPolicy that admits only the ingress controller.
To add actions, list columns or palette aliases for everyone, put them under extensions: in your values; see Extensions.
RBAC for impersonation
Section titled “RBAC for impersonation”Outside local mode KubeGlass makes every Kubernetes call as the signed-in user, so its ServiceAccount needs the impersonate verb. The chart’s ClusterRole grants it on users, groups and serviceaccounts by default (rbac.impersonation.enabled), and leaves the rule out in local mode. Limit whom KubeGlass may impersonate with rbac.impersonation.resourceNames:
rbac: impersonation: resourceNames: - platform-teamrbac.readSecrets (default true) controls whether the ClusterRole can read Secrets.
If you manage RBAC yourself (rbac.create=false), the rule the chart creates is:
apiVersion: rbac.authorization.k8s.io/v1kind: ClusterRolemetadata: name: kubeglass-impersonatorrules: - apiGroups: [""] resources: ["users", "groups", "serviceaccounts"] verbs: ["impersonate"] # Add resourceNames to restrict which users and groups can be impersonated. - apiGroups: ["authentication.k8s.io"] resources: ["userextras/scopes", "userextras/uid"] verbs: ["impersonate"]Production checklist
Section titled “Production checklist”-
config.authModeisoidcorimpersonation, notlocal. - In
impersonationmode,config.trustedProxiesnames only your proxy’s addresses, and only the proxy can reach the pod. -
KUBEGLASS_ENV=productionis set throughextraEnv. This turns on stricter checks: the OIDC issuer must use HTTPS, and*is not allowed inKUBEGLASS_CORS_ALLOWED_ORIGINS. - TLS is terminated at the ingress or load balancer (or
KUBEGLASS_TLS_ENABLED=truewith a mounted certificate). -
KUBEGLASS_CORS_ALLOWED_ORIGINS, if set, lists specific origins. -
networkPolicy.enabledistrue(the default), withnetworkPolicy.ingressFromrestricted to your ingress controller. -
rbac.impersonation.resourceNamesis set where you can list who may be impersonated. -
persistence.existingClaimpoints to a PersistentVolumeClaim, so drift policies and results, inventory snapshots, the change history and terminal history survive restarts. Without it the data directory is anemptyDir. -
image.digestpins the image you verified (see Supply chain), so a moved tag can’t change what runs. - The namespace enforces Pod Security
restricted(label itpod-security.kubernetes.io/enforce: restricted). The chart’s pods meet it. -
resourcesare sized for your clusters. -
replicaCountstays at 1 andautoscaling.enabledstaysfalse(see High availability).
Health checks
Section titled “Health checks”| Endpoint | What it checks | Used for |
|---|---|---|
/healthz |
The process can answer | Liveness and startup probes |
/livez |
Same as /healthz |
Liveness probes in manual deployments |
/readyz |
The Kubernetes API server answers within 5 seconds; returns 503 if not | Readiness probe |
The chart sets up liveness, readiness and startup probes. For a manual deployment:
livenessProbe: httpGet: path: /healthz port: 8090 initialDelaySeconds: 10 periodSeconds: 15readinessProbe: httpGet: path: /readyz port: 8090 initialDelaySeconds: 5 periodSeconds: 10To check from your machine:
kubectl -n kubeglass port-forward svc/kubeglass 8090:8090 &curl -s http://localhost:8090/healthzcurl -s http://localhost:8090/readyzMonitoring
Section titled “Monitoring”Metrics are at /metrics on the service port. Set serviceMonitor.enabled=true to have the chart create a ServiceMonitor for the Prometheus Operator, and prometheusRule.enabled=true for its alert rules (KubeGlassHighErrorRate, KubeGlassPodRestarting, KubeGlassDown).
Metrics to watch
Section titled “Metrics to watch”The thresholds are starting points; tune them for your traffic.
| Metric | Type | Suggested alert |
|---|---|---|
kubeglass_api_requests_total |
counter, by method, path (the route) and status |
5xx above 5% of requests (the chart’s KubeGlassHighErrorRate) |
kubeglass_api_request_duration_seconds |
histogram, by method and path |
p99 above 2s |
kubeglass_ws_active_connections |
gauge | Near the cap for a sustained period: 100 by default (KUBEGLASS_MAX_WS_CONNECTIONS), 10 per user (KUBEGLASS_MAX_WS_CONNECTIONS_PER_USER) |
kubeglass_ws_messages_dropped_total |
counter | Rate above 0 for a sustained period |
kubeglass_ws_broadcast_duration_seconds |
histogram | p99 above 100ms |
Log queries
Section titled “Log queries”In a pod KubeGlass logs one JSON object per line. Request lines carry method, path, status, duration_ms and request_id; most subsystems add a component field.
# Errorskubectl -n kubeglass logs deploy/kubeglass | jq 'select(.level == "error")'
# Rejected sign-ins (bad tokens, proxy headers from untrusted addresses)kubectl -n kubeglass logs deploy/kubeglass | jq 'select((.message // "") | test("^Auth rejected|OIDC token validation failed"))'
# Slow requests (over 5 seconds)kubectl -n kubeglass logs deploy/kubeglass | jq 'select(.duration_ms > 5000)'
# Circuit breaker changeskubectl -n kubeglass logs deploy/kubeglass | jq 'select(.component == "circuit-breaker")'
# WebSocket eventskubectl -n kubeglass logs deploy/kubeglass | jq 'select(.component | strings | startswith("ws"))'Circuit breaker
Section titled “Circuit breaker”The circuit breaker opens when the cluster health check can’t reach the Kubernetes API server, so watches and log streams back off instead of hammering it.
Signs that it is open:
- The UI shows stale data and live updates stop.
- The log has
circuit opened - cluster unreachable, subsystems will back off, thencircuit still open - cluster unreachablewhile the outage lasts. /readyzreturns 503, so the pod drops out of the Service.
It closes by itself (circuit closed - cluster connectivity restored) when the health check succeeds again. To find the cause:
# Is the API server healthy?kubectl get --raw /healthz
# Can KubeGlass's ServiceAccount still list pods?kubectl auth can-i list pods --as=system:serviceaccount:kubeglass:kubeglassBackup and restore
Section titled “Backup and restore”What is stored
Section titled “What is stored”KubeGlass keeps its state in the data directory, /var/lib/kubeglass/data in the chart:
| File | Contents |
|---|---|
kubeglass.db |
Drift results and policies, inventory snapshots, change history |
terminal-snapshots.bolt |
Terminal session history |
config/settings.json |
Settings changed on the Settings page |
Everything else KubeGlass shows comes from the cluster. Without persistence.existingClaim the directory is an emptyDir and is lost when the pod is replaced, so there is nothing to back up.
Backup
Section titled “Backup”The image is distroless, with no shell, cp or tar, so kubectl exec and kubectl cp don’t work against the KubeGlass container. Stop KubeGlass, mount its PersistentVolumeClaim in a temporary pod, and copy from there. Replace kubeglass-data with the name of your claim.
# 1. Stop KubeGlass so nothing writes to the databasekubectl -n kubeglass scale deploy/kubeglass --replicas=0
# 2. Start a helper pod with the claim mounted at /datakubectl -n kubeglass run kubeglass-backup --image=busybox:1.36 --restart=Never \ --overrides='{ "spec": { "securityContext": {"runAsNonRoot": true, "runAsUser": 65532, "runAsGroup": 65532, "fsGroup": 65532, "seccompProfile": {"type": "RuntimeDefault"}}, "volumes": [{"name": "data", "persistentVolumeClaim": {"claimName": "kubeglass-data"}}], "containers": [{ "name": "kubeglass-backup", "image": "busybox:1.36", "command": ["sleep", "3600"], "securityContext": {"allowPrivilegeEscalation": false, "capabilities": {"drop": ["ALL"]}}, "volumeMounts": [{"name": "data", "mountPath": "/data"}] }] } }'kubectl -n kubeglass wait --for=condition=Ready pod/kubeglass-backup
# 3. Copy the data directory to your machinekubectl -n kubeglass cp kubeglass-backup:/data ./kubeglass-backup
# 4. Remove the helper and start KubeGlass againkubectl -n kubeglass delete pod kubeglass-backupkubectl -n kubeglass scale deploy/kubeglass --replicas=1Restore
Section titled “Restore”Follow steps 1 and 2 above, then copy the files back before removing the helper:
kubectl -n kubeglass cp ./kubeglass-backup/kubeglass.db kubeglass-backup:/data/kubeglass.dbkubectl -n kubeglass cp ./kubeglass-backup/terminal-snapshots.bolt kubeglass-backup:/data/terminal-snapshots.boltkubectl -n kubeglass exec kubeglass-backup -- mkdir -p /data/configkubectl -n kubeglass cp ./kubeglass-backup/config/settings.json kubeglass-backup:/data/config/settings.json
kubectl -n kubeglass delete pod kubeglass-backupkubectl -n kubeglass scale deploy/kubeglass --replicas=1High availability
Section titled “High availability”KubeGlass runs as a single replica. bbolt is a single-file embedded database, and only one process can open it for writing.
| Aspect | Constraint |
|---|---|
| Replicas | 1. There is no active/active or active/passive mode |
| Failover | Kubernetes restarts or reschedules the pod; there is no standby |
| Data replication | None built in. Rely on volume replication or backups |
| Horizontal scaling | Not supported. Keep replicaCount: 1 and autoscaling.enabled: false |
With a PersistentVolumeClaim the chart uses the Recreate update strategy, so the old pod releases the volume before the new one starts. The chart’s PodDisruptionBudget is off by default; with one replica, turning it on (podDisruptionBudget.enabled=true, minAvailable: 1) blocks voluntary evictions such as node drains until you move the pod yourself.
For production:
- Use a storage class whose CSI driver supports volume snapshots, and snapshot the claim on a schedule.
- Or run the backup above on a schedule and copy the files to object storage.
- Watch
/readyzand theKubeGlassDownalert.
Data retention
Section titled “Data retention”| Data | Kept for | Pruned |
|---|---|---|
| Drift results | 30 days (KUBEGLASS_DRIFT_RESULT_RETENTION) |
Every hour (KUBEGLASS_DRIFT_PRUNE_INTERVAL) |
| Change history | 30 days (same setting) | Every hour |
| Inventory snapshots | Not pruned by age | |
| Browser sessions (in memory) | Idle for session_max_age (default 4 hours) |
Every 5 minutes |
Incident response
Section titled “Incident response”High memory use
Section titled “High memory use”- Check
kubeglass_ws_active_connectionsfor many open WebSocket connections. - Look for a leak: connection count that grows and never falls.
- Check the size of
kubeglass.db(with the backup helper pod:ls -lh /data). bbolt doesn’t shrink the file when data is deleted. - Lists, watches and search come from informers that KubeGlass starts per user and type when someone uses them. Each stops after 10 minutes without use, and at most 256 run at once.
- The rate limiter tracks at most 100,000 client IPs and sweeps idle ones every 5 minutes.
Sign-in failures
Section titled “Sign-in failures”- Check the auth mode:
kubectl -n kubeglass get configmap kubeglass -o jsonpath='{.data.KUBEGLASS_AUTH_MODE}'. - In
impersonationmode, check that the proxy’s address is insideKUBEGLASS_TRUSTED_PROXIESand that it sendsX-Forwarded-User. Look forAuth rejected: impersonation headers from untrusted sourceandAuth rejected: missing X-Forwarded-User headerin the log. - In
oidcmode, check that the issuer URL is reachable from the pod and that the token’s audience matchesKUBEGLASS_OIDC_AUDIENCE. Look forOIDC token validation failedand for lines with"component":"oidc-validator". - Validated tokens are cached for one minute (up to 10,000), so a revoked token can keep working for up to a minute.
- If the pod doesn’t start, read the startup errors with
kubectl -n kubeglass logs deploy/kubeglass. Configuration errors (a missing OIDC issuer, a missingKUBEGLASS_TRUSTED_PROXIESin impersonation mode) stop the server before it listens.
WebSocket disconnections
Section titled “WebSocket disconnections”- Check the load balancer’s idle timeout. It should be at least 120 seconds, KubeGlass’s
IdleTimeout; the server pings every 30 seconds. - Check that every proxy in the path supports WebSockets and doesn’t buffer them. With ingress-nginx, raise
proxy-read-timeoutandproxy-send-timeout(values-production.yamlsets 3600). - In the browser, open DevTools, then Network, then the WS filter, and check the close codes.
- On the server, look for lines with a
componentofws,ws-huborws-handler.
Slow UI
Section titled “Slow UI”- Check the log for circuit breaker messages.
- Check API server latency:
kubectl get pods -A -v=6prints the time each request took. - Check PromQL queries: they are limited to 4,096 characters and a 30-day range.
- Check
kubeglass_ws_broadcast_duration_seconds.
Terminal problems
Section titled “Terminal problems”- Terminal sessions end after 30 minutes idle (with a warning at 24 minutes). Exec sessions in pods also end after 8 hours whatever their activity (with a 5-minute warning).
- Check that the user may exec:
kubectl auth can-i create pods/exec -n <namespace> --as=<user>. - A single WebSocket message may be at most 1 MB.
Secret and key rotation
Section titled “Secret and key rotation”TLS certificate
Section titled “TLS certificate”If TLS is terminated at the ingress (recommended), follow your ingress controller’s procedure; cert-manager renews certificates by itself.
If KubeGlass terminates TLS itself, you set KUBEGLASS_TLS_ENABLED=true, KUBEGLASS_TLS_CERT and KUBEGLASS_TLS_KEY with extraEnv and mounted a TLS Secret with extraVolumes and extraVolumeMounts. To rotate it:
# 1. Update the TLS Secret with the renewed certificate and keykubectl -n kubeglass create secret tls kubeglass-tls \ --cert=new-cert.pem \ --key=new-key.pem \ --dry-run=client -o yaml | kubectl apply -f -
# 2. Restart to load the new certificatekubectl -n kubeglass rollout restart deploy/kubeglass
# 3. Check the certificate being servedopenssl s_client -connect kubeglass.example.com:443 -servername kubeglass.example.com < /dev/null 2>/dev/null | \ openssl x509 -noout -dates -subjectUpgrade and rollback
Section titled “Upgrade and rollback”Upgrade
Section titled “Upgrade”# 1. Review what will change (needs the helm-diff plugin)helm diff upgrade kubeglass oci://ghcr.io/kubeglass/charts/kubeglass \ --namespace kubeglass \ -f custom-values.yaml
# 2. Back up the data directory (see Backup above)
# 3. Upgradehelm upgrade kubeglass oci://ghcr.io/kubeglass/charts/kubeglass \ --namespace kubeglass \ -f custom-values.yaml \ --wait --timeout 5m
# 4. Verifykubectl -n kubeglass rollout status deploy/kubeglasscurl -sf https://kubeglass.example.com/readyzAdd --version <chart-version> to helm diff upgrade and helm upgrade to pin a release.
Rollback
Section titled “Rollback”# List release historyhelm history kubeglass -n kubeglass
# Roll back to an earlier revisionhelm rollback kubeglass <revision> -n kubeglass --waitIf the newer version changed stored data in a way the older one can’t read, restore the backup you took before upgrading (see Restore).
Contacts
Section titled “Contacts”MAINTAINERS.md lists the maintainers. SECURITY.md describes how to report a vulnerability.