Skip to content

Operations

How to deploy, monitor, back up and upgrade KubeGlass in a cluster. The commands assume the Helm chart installed as release kubeglass in namespace kubeglass, so the Deployment, Service and ConfigMap are all called kubeglass. Adjust the names if yours differ.

The configuration reference lists every setting, and the chart’s values.yaml lists every chart option.

The chart is published as an OCI artifact at oci://ghcr.io/kubeglass/charts/kubeglass; there is no helm repo add step.

Behind an authenticating reverse proxy (the default impersonation mode), set the proxy’s CIDR. The install fails without it:

Terminal window
helm install kubeglass oci://ghcr.io/kubeglass/charts/kubeglass \
--namespace kubeglass \
--create-namespace \
--set config.trustedProxies=10.0.0.0/8

With OIDC, the chart has no dedicated keys for the issuer and audience (the values schema rejects unknown keys), so pass them with extraEnv. Save this as values-oidc.yaml:

config:
authMode: oidc
extraEnv:
- name: KUBEGLASS_OIDC_ISSUER_URL
value: https://accounts.google.com
- name: KUBEGLASS_OIDC_AUDIENCE
value: <client-id>
- name: KUBEGLASS_ENV
value: production
Terminal window
helm install kubeglass oci://ghcr.io/kubeglass/charts/kubeglass \
--namespace kubeglass \
--create-namespace \
-f values-oidc.yaml

To try it without a proxy or identity provider, use --set config.authMode=local and reach it with kubectl -n kubeglass port-forward svc/kubeglass 8090:8090. In local mode anyone who can reach the Service gets KubeGlass’s own access, Secrets included, so don’t expose it any other way.

deploy/helm/kubeglass/values-production.yaml is a starting point for a production install: ingress with TLS, larger resources, a scoped impersonation rule and a NetworkPolicy that admits only the ingress controller.

To add actions, list columns or palette aliases for everyone, put them under extensions: in your values; see Extensions.

Outside local mode KubeGlass makes every Kubernetes call as the signed-in user, so its ServiceAccount needs the impersonate verb. The chart’s ClusterRole grants it on users, groups and serviceaccounts by default (rbac.impersonation.enabled), and leaves the rule out in local mode. Limit whom KubeGlass may impersonate with rbac.impersonation.resourceNames:

rbac:
impersonation:
resourceNames:
- platform-team

rbac.readSecrets (default true) controls whether the ClusterRole can read Secrets.

If you manage RBAC yourself (rbac.create=false), the rule the chart creates is:

apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRole
metadata:
name: kubeglass-impersonator
rules:
- apiGroups: [""]
resources: ["users", "groups", "serviceaccounts"]
verbs: ["impersonate"]
# Add resourceNames to restrict which users and groups can be impersonated.
- apiGroups: ["authentication.k8s.io"]
resources: ["userextras/scopes", "userextras/uid"]
verbs: ["impersonate"]
  • config.authMode is oidc or impersonation, not local.
  • In impersonation mode, config.trustedProxies names only your proxy’s addresses, and only the proxy can reach the pod.
  • KUBEGLASS_ENV=production is set through extraEnv. This turns on stricter checks: the OIDC issuer must use HTTPS, and * is not allowed in KUBEGLASS_CORS_ALLOWED_ORIGINS.
  • TLS is terminated at the ingress or load balancer (or KUBEGLASS_TLS_ENABLED=true with a mounted certificate).
  • KUBEGLASS_CORS_ALLOWED_ORIGINS, if set, lists specific origins.
  • networkPolicy.enabled is true (the default), with networkPolicy.ingressFrom restricted to your ingress controller.
  • rbac.impersonation.resourceNames is set where you can list who may be impersonated.
  • persistence.existingClaim points to a PersistentVolumeClaim, so drift policies and results, inventory snapshots, the change history and terminal history survive restarts. Without it the data directory is an emptyDir.
  • image.digest pins the image you verified (see Supply chain), so a moved tag can’t change what runs.
  • The namespace enforces Pod Security restricted (label it pod-security.kubernetes.io/enforce: restricted). The chart’s pods meet it.
  • resources are sized for your clusters.
  • replicaCount stays at 1 and autoscaling.enabled stays false (see High availability).
Endpoint What it checks Used for
/healthz The process can answer Liveness and startup probes
/livez Same as /healthz Liveness probes in manual deployments
/readyz The Kubernetes API server answers within 5 seconds; returns 503 if not Readiness probe

The chart sets up liveness, readiness and startup probes. For a manual deployment:

livenessProbe:
httpGet:
path: /healthz
port: 8090
initialDelaySeconds: 10
periodSeconds: 15
readinessProbe:
httpGet:
path: /readyz
port: 8090
initialDelaySeconds: 5
periodSeconds: 10

To check from your machine:

Terminal window
kubectl -n kubeglass port-forward svc/kubeglass 8090:8090 &
curl -s http://localhost:8090/healthz
curl -s http://localhost:8090/readyz

Metrics are at /metrics on the service port. Set serviceMonitor.enabled=true to have the chart create a ServiceMonitor for the Prometheus Operator, and prometheusRule.enabled=true for its alert rules (KubeGlassHighErrorRate, KubeGlassPodRestarting, KubeGlassDown).

The thresholds are starting points; tune them for your traffic.

Metric Type Suggested alert
kubeglass_api_requests_total counter, by method, path (the route) and status 5xx above 5% of requests (the chart’s KubeGlassHighErrorRate)
kubeglass_api_request_duration_seconds histogram, by method and path p99 above 2s
kubeglass_ws_active_connections gauge Near the cap for a sustained period: 100 by default (KUBEGLASS_MAX_WS_CONNECTIONS), 10 per user (KUBEGLASS_MAX_WS_CONNECTIONS_PER_USER)
kubeglass_ws_messages_dropped_total counter Rate above 0 for a sustained period
kubeglass_ws_broadcast_duration_seconds histogram p99 above 100ms

In a pod KubeGlass logs one JSON object per line. Request lines carry method, path, status, duration_ms and request_id; most subsystems add a component field.

Terminal window
# Errors
kubectl -n kubeglass logs deploy/kubeglass | jq 'select(.level == "error")'
# Rejected sign-ins (bad tokens, proxy headers from untrusted addresses)
kubectl -n kubeglass logs deploy/kubeglass | jq 'select((.message // "") | test("^Auth rejected|OIDC token validation failed"))'
# Slow requests (over 5 seconds)
kubectl -n kubeglass logs deploy/kubeglass | jq 'select(.duration_ms > 5000)'
# Circuit breaker changes
kubectl -n kubeglass logs deploy/kubeglass | jq 'select(.component == "circuit-breaker")'
# WebSocket events
kubectl -n kubeglass logs deploy/kubeglass | jq 'select(.component | strings | startswith("ws"))'

The circuit breaker opens when the cluster health check can’t reach the Kubernetes API server, so watches and log streams back off instead of hammering it.

Signs that it is open:

  • The UI shows stale data and live updates stop.
  • The log has circuit opened - cluster unreachable, subsystems will back off, then circuit still open - cluster unreachable while the outage lasts.
  • /readyz returns 503, so the pod drops out of the Service.

It closes by itself (circuit closed - cluster connectivity restored) when the health check succeeds again. To find the cause:

Terminal window
# Is the API server healthy?
kubectl get --raw /healthz
# Can KubeGlass's ServiceAccount still list pods?
kubectl auth can-i list pods --as=system:serviceaccount:kubeglass:kubeglass

KubeGlass keeps its state in the data directory, /var/lib/kubeglass/data in the chart:

File Contents
kubeglass.db Drift results and policies, inventory snapshots, change history
terminal-snapshots.bolt Terminal session history
config/settings.json Settings changed on the Settings page

Everything else KubeGlass shows comes from the cluster. Without persistence.existingClaim the directory is an emptyDir and is lost when the pod is replaced, so there is nothing to back up.

The image is distroless, with no shell, cp or tar, so kubectl exec and kubectl cp don’t work against the KubeGlass container. Stop KubeGlass, mount its PersistentVolumeClaim in a temporary pod, and copy from there. Replace kubeglass-data with the name of your claim.

Terminal window
# 1. Stop KubeGlass so nothing writes to the database
kubectl -n kubeglass scale deploy/kubeglass --replicas=0
# 2. Start a helper pod with the claim mounted at /data
kubectl -n kubeglass run kubeglass-backup --image=busybox:1.36 --restart=Never \
--overrides='{
"spec": {
"securityContext": {"runAsNonRoot": true, "runAsUser": 65532, "runAsGroup": 65532, "fsGroup": 65532,
"seccompProfile": {"type": "RuntimeDefault"}},
"volumes": [{"name": "data", "persistentVolumeClaim": {"claimName": "kubeglass-data"}}],
"containers": [{
"name": "kubeglass-backup", "image": "busybox:1.36", "command": ["sleep", "3600"],
"securityContext": {"allowPrivilegeEscalation": false, "capabilities": {"drop": ["ALL"]}},
"volumeMounts": [{"name": "data", "mountPath": "/data"}]
}]
}
}'
kubectl -n kubeglass wait --for=condition=Ready pod/kubeglass-backup
# 3. Copy the data directory to your machine
kubectl -n kubeglass cp kubeglass-backup:/data ./kubeglass-backup
# 4. Remove the helper and start KubeGlass again
kubectl -n kubeglass delete pod kubeglass-backup
kubectl -n kubeglass scale deploy/kubeglass --replicas=1

Follow steps 1 and 2 above, then copy the files back before removing the helper:

Terminal window
kubectl -n kubeglass cp ./kubeglass-backup/kubeglass.db kubeglass-backup:/data/kubeglass.db
kubectl -n kubeglass cp ./kubeglass-backup/terminal-snapshots.bolt kubeglass-backup:/data/terminal-snapshots.bolt
kubectl -n kubeglass exec kubeglass-backup -- mkdir -p /data/config
kubectl -n kubeglass cp ./kubeglass-backup/config/settings.json kubeglass-backup:/data/config/settings.json
kubectl -n kubeglass delete pod kubeglass-backup
kubectl -n kubeglass scale deploy/kubeglass --replicas=1

KubeGlass runs as a single replica. bbolt is a single-file embedded database, and only one process can open it for writing.

Aspect Constraint
Replicas 1. There is no active/active or active/passive mode
Failover Kubernetes restarts or reschedules the pod; there is no standby
Data replication None built in. Rely on volume replication or backups
Horizontal scaling Not supported. Keep replicaCount: 1 and autoscaling.enabled: false

With a PersistentVolumeClaim the chart uses the Recreate update strategy, so the old pod releases the volume before the new one starts. The chart’s PodDisruptionBudget is off by default; with one replica, turning it on (podDisruptionBudget.enabled=true, minAvailable: 1) blocks voluntary evictions such as node drains until you move the pod yourself.

For production:

  • Use a storage class whose CSI driver supports volume snapshots, and snapshot the claim on a schedule.
  • Or run the backup above on a schedule and copy the files to object storage.
  • Watch /readyz and the KubeGlassDown alert.
Data Kept for Pruned
Drift results 30 days (KUBEGLASS_DRIFT_RESULT_RETENTION) Every hour (KUBEGLASS_DRIFT_PRUNE_INTERVAL)
Change history 30 days (same setting) Every hour
Inventory snapshots Not pruned by age
Browser sessions (in memory) Idle for session_max_age (default 4 hours) Every 5 minutes
  1. Check kubeglass_ws_active_connections for many open WebSocket connections.
  2. Look for a leak: connection count that grows and never falls.
  3. Check the size of kubeglass.db (with the backup helper pod: ls -lh /data). bbolt doesn’t shrink the file when data is deleted.
  4. Lists, watches and search come from informers that KubeGlass starts per user and type when someone uses them. Each stops after 10 minutes without use, and at most 256 run at once.
  5. The rate limiter tracks at most 100,000 client IPs and sweeps idle ones every 5 minutes.
  1. Check the auth mode: kubectl -n kubeglass get configmap kubeglass -o jsonpath='{.data.KUBEGLASS_AUTH_MODE}'.
  2. In impersonation mode, check that the proxy’s address is inside KUBEGLASS_TRUSTED_PROXIES and that it sends X-Forwarded-User. Look for Auth rejected: impersonation headers from untrusted source and Auth rejected: missing X-Forwarded-User header in the log.
  3. In oidc mode, check that the issuer URL is reachable from the pod and that the token’s audience matches KUBEGLASS_OIDC_AUDIENCE. Look for OIDC token validation failed and for lines with "component":"oidc-validator".
  4. Validated tokens are cached for one minute (up to 10,000), so a revoked token can keep working for up to a minute.
  5. If the pod doesn’t start, read the startup errors with kubectl -n kubeglass logs deploy/kubeglass. Configuration errors (a missing OIDC issuer, a missing KUBEGLASS_TRUSTED_PROXIES in impersonation mode) stop the server before it listens.
  1. Check the load balancer’s idle timeout. It should be at least 120 seconds, KubeGlass’s IdleTimeout; the server pings every 30 seconds.
  2. Check that every proxy in the path supports WebSockets and doesn’t buffer them. With ingress-nginx, raise proxy-read-timeout and proxy-send-timeout (values-production.yaml sets 3600).
  3. In the browser, open DevTools, then Network, then the WS filter, and check the close codes.
  4. On the server, look for lines with a component of ws, ws-hub or ws-handler.
  1. Check the log for circuit breaker messages.
  2. Check API server latency: kubectl get pods -A -v=6 prints the time each request took.
  3. Check PromQL queries: they are limited to 4,096 characters and a 30-day range.
  4. Check kubeglass_ws_broadcast_duration_seconds.
  1. Terminal sessions end after 30 minutes idle (with a warning at 24 minutes). Exec sessions in pods also end after 8 hours whatever their activity (with a 5-minute warning).
  2. Check that the user may exec: kubectl auth can-i create pods/exec -n <namespace> --as=<user>.
  3. A single WebSocket message may be at most 1 MB.

If TLS is terminated at the ingress (recommended), follow your ingress controller’s procedure; cert-manager renews certificates by itself.

If KubeGlass terminates TLS itself, you set KUBEGLASS_TLS_ENABLED=true, KUBEGLASS_TLS_CERT and KUBEGLASS_TLS_KEY with extraEnv and mounted a TLS Secret with extraVolumes and extraVolumeMounts. To rotate it:

Terminal window
# 1. Update the TLS Secret with the renewed certificate and key
kubectl -n kubeglass create secret tls kubeglass-tls \
--cert=new-cert.pem \
--key=new-key.pem \
--dry-run=client -o yaml | kubectl apply -f -
# 2. Restart to load the new certificate
kubectl -n kubeglass rollout restart deploy/kubeglass
# 3. Check the certificate being served
openssl s_client -connect kubeglass.example.com:443 -servername kubeglass.example.com < /dev/null 2>/dev/null | \
openssl x509 -noout -dates -subject
Terminal window
# 1. Review what will change (needs the helm-diff plugin)
helm diff upgrade kubeglass oci://ghcr.io/kubeglass/charts/kubeglass \
--namespace kubeglass \
-f custom-values.yaml
# 2. Back up the data directory (see Backup above)
# 3. Upgrade
helm upgrade kubeglass oci://ghcr.io/kubeglass/charts/kubeglass \
--namespace kubeglass \
-f custom-values.yaml \
--wait --timeout 5m
# 4. Verify
kubectl -n kubeglass rollout status deploy/kubeglass
curl -sf https://kubeglass.example.com/readyz

Add --version <chart-version> to helm diff upgrade and helm upgrade to pin a release.

Terminal window
# List release history
helm history kubeglass -n kubeglass
# Roll back to an earlier revision
helm rollback kubeglass <revision> -n kubeglass --wait

If the newer version changed stored data in a way the older one can’t read, restore the backup you took before upgrading (see Restore).

MAINTAINERS.md lists the maintainers. SECURITY.md describes how to report a vulnerability.