node
down when not ready, critical when not reporting, warn under memory, disk or PID pressure or when cordoned
- Attributes
hostnameipkubeletosarch
User guideWayseer 0.28.3Contents
The kubernetes module shows a cluster as nodes, namespaces, pods, services, deployments and replica sets, linked the way the cluster links them. It reads, and changes nothing unless you allow one of its actions and confirm it.
It connects the way kubectl does. It uses the same kubeconfig, contexts and credentials, including exec credential plugins (cloud CLIs, kubelogin) and OIDC. If kubectl get pods works for a context, this module works for that context too.
The Kubernetes module is built into Wayseer, so there is nothing to install:
modules:
- kind: kubernetes
name: k8s
options:
context: staging # leave out for the kubeconfig's current context
namespaces: [shop, payments] # leave out for every namespace
The module runs inside the app and sees the app's environment, as kubectl would in the same shell: HOME and KUBECONFIG to find the kubeconfig, PATH to find an exec credential plugin (aws, gke-gcloud-auth-plugin, kubelogin), and whatever that plugin reads, such as AWS_PROFILE. If a context works with kubectl but not here, check that Wayseer was started with the same variables. A desktop launcher may not set what your shell does.
Nothing connects to a cluster, and no kubeconfig is read, until a kind: kubernetes instance is configured.
| Option | Type | Default | Meaning |
|---|---|---|---|
kubeconfig | string | $KUBECONFIG, then ~/.kube/config | The kubeconfig file. |
context | string | the kubeconfig's current context | Which cluster and user to use. |
namespaces | list of string | every namespace | Only list objects in these namespaces. |
in_cluster | bool | false | Use the pod's service account, when running inside the cluster. |
timeout | duration | 10s | Longest wait for one request, 100ms to 5m. |
max_replicas | integer | 20 | The most replicas the scale action sets, 1 to 10000. |
With in_cluster, when Wayseer runs inside the cluster, leave out kubeconfig and context.
Credentials stay in the kubeconfig, or with the plugin it names. Nothing about them goes into Wayseer's config, and errors never show a token or the server's address. For example, a refused token shows as "the cluster refused the credentials (401 Unauthorized)".
Add one module instance for each cluster you want to see. Give each instance its own name and context.
node
down when not ready, critical when not reporting, warn under memory, disk or PID pressure or when cordoned
hostnameipkubeletosarchk8s/namespace
warn while being deleted
pod
the phase; critical for a container stuck in CrashLoopBackOff or an image pull error; warn while pending, when not every container is ready, or when a container restarted in the last 10 minutes
namespace, node, phase, ready (as 2/3), restarts, ip, hostname, ownerservice
critical when none of its pods is ready, warn when its selector matches no pods
namespacetypecluster_ipportsk8s/deployment
available replicas against those wanted: critical at none, warn at fewer; critical when a rollout has stalled
namespacereplicasreadyavailablestrategyk8s/replicaset
as for a deployment
namespacereplicasreadyavailableEach object's labels become its tags, as key=value. Old replica sets that a deployment keeps as history, scaled to nothing, are not shown.
The links are:
In Topology, each namespace is a round group of its deployments, replica sets, pods and services, and the nodes sit in the middle with the daemon set pods on them.
Nodes carry the hostname and ip attributes that the default identity rules match. So a node that another module also shows, such as a Prometheus node-exporter target, merges with it.
Some filters:
/source:k8s status:warn
/kind:pod namespace=shop
/kind:pod restarts>0
The module lists the cluster once, then watches it. A change shows within a fraction of a second: changes that arrive together wait 0.1 s and go in one update. Every 10 minutes the module sends the whole cluster again, which corrects anything a missed change left wrong.
If a watch stops, for example because the API server restarted, the module lists that kind of object again. Meanwhile its health shows a note such as "watch stopped; listing pods again". That is not an error: what shows is still the last known state. The module is only in error when listing fails, and then the error names what it could not list.
Kubernetes events show on the Timeline, on the object they are about, and in the detail panel's events. A repeated event, such as a container backing off again, shows each time it happens, with its count.
Normal
info
Warning
warn
Warning that means something broke: BackOff, Failed, FailedCreate, FailedMount, FailedAttachVolume, FailedScheduling, Evicted, OOMKilling, SystemOOM, NodeNotReady
error
Rollout events (ScalingReplicaSet, SuccessfulCreate, SuccessfulDelete) have the kind deploy; the rest have the kind k8s. An event about an object the module does not show, such as a Job, goes on that object's namespace, with the object named first: Job/backup: …. Each event carries reason and object fields.
When the module starts, the events the cluster still keeps (usually the last hour) show as well.
When metrics-server is installed, the module reads it every 15 seconds and offers:
cpu.utilisation
For a node, the share of its allocatable CPU in use; for a pod, the share of one core, so a pod using two cores shows 200%
nodepodmemory.utilisation
Working-set memory as a share of the node's allocatable memory
nodememory.used
Working-set memory
nodepodIt keeps the last hour. Without metrics-server there are no metrics, and nothing reports the module as unhealthy. For longer history, add a Prometheus module as well.
A Prometheus module with series adds metrics of your own to the pods and deployments this module finds, such as an application's queue depth, or the metric a HorizontalPodAutoscaler scales on, as kube-state-metrics reports it:
modules:
- kind: kubernetes
name: k8s
- kind: prometheus
name: prom
options:
url: http://localhost:9090
series:
- query: max by (namespace, pod, queue) (shop_queue_depth)
metric: queue.depth
unit: count
description: messages waiting in the pod's queues
entity: {labels: [namespace, pod], kind: pod}
- query: max by (namespace, horizontalpodautoscaler) (kube_horizontalpodautoscaler_status_target_metric{metric_name="queue_depth",metric_target_type="average"})
metric: hpa.queue_depth
description: the queue depth the deployment's autoscaler sees
entity: {labels: [namespace, horizontalpodautoscaler], kind: k8s/deployment}
A pod's namespace and pod labels name it as this module does, namespace/name. The second query names the deployment after its autoscaler, which works where they share a name, as they usually do. Then >grid.show metric=queue.depth kind=pod ranks every pod by its queue.
The module needs get, list and watch on the objects it shows and on events, and get and list on metrics-server's nodes and pods. It needs nothing else. This read-only ClusterRole covers every namespace:
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRole
metadata:
name: wayseer-read
rules:
- apiGroups: [""]
resources: [nodes, namespaces, pods, services, events]
verbs: [get, list, watch]
- apiGroups: [apps]
resources: [deployments, replicasets]
verbs: [get, list, watch]
- apiGroups: [metrics.k8s.io]
resources: [nodes, pods]
verbs: [get, list]
Bind it to the user or service account that Wayseer connects as, with a ClusterRoleBinding.
With access to only some namespaces, set namespaces to those namespaces and grant the same rules through a Role and RoleBinding in each of them. Nodes are cluster-wide, so a namespaced role cannot list them. The cluster then shows without nodes, and the module's health notes why. It is not an error. The same goes for events: without the right to list them, the cluster shows without its events, with a note.
Events about nodes are kept in the default namespace. With namespaces set, node events show only if default is one of them.
The module offers three actions. None runs unless the instance's actions lists it, and each waits for you to confirm it.
restart
Stamps kubectl.kubernetes.io/restartedAt on the pod template, as kubectl rollout restart does, so the pods are replaced in turn.
patch on deploymentsscale
Sets the replica count, from 0 to max_replicas (default 20).
patch on deployments/scaledelete
Deletes the pod with its default grace period. Its owner, if it has one, starts another.
delete on podsmodules:
- kind: kubernetes
name: k8s
actions: [restart, scale]
options:
max_replicas: 10
The read-only role above allows none of them. Grant only the rules for the actions you list, in a separate role, so reading and acting stay apart:
apiVersion: rbac.authorization.k8s.io/v1
kind: Role
metadata:
name: wayseer-act
namespace: shop
rules:
- apiGroups: [apps]
resources: [deployments]
verbs: [patch] # restart
- apiGroups: [apps]
resources: [deployments/scale]
verbs: [patch] # scale
- apiGroups: [""]
resources: [pods]
verbs: [delete] # delete
When the cluster refuses an action, the status bar names the verb, resource and namespace that a role would need, as in "the kubeconfig's user may not patch deployments/scale in shop". It never shows the server's address, your user name or a token. With in_cluster, it reads "the pod's service account" instead.