Wayseer

User guideWayseer 0.28.3Contents

Kubernetes

The kubernetes module shows a cluster as nodes, namespaces, pods, services, deployments and replica sets, linked the way the cluster links them. It reads, and changes nothing unless you allow one of its actions and confirm it.

It connects the way kubectl does. It uses the same kubeconfig, contexts and credentials, including exec credential plugins (cloud CLIs, kubelogin) and OIDC. If kubectl get pods works for a context, this module works for that context too.

Configuration

The Kubernetes module is built into Wayseer, so there is nothing to install:

modules:
  - kind: kubernetes
    name: k8s
    options:
      context: staging             # leave out for the kubeconfig's current context
      namespaces: [shop, payments] # leave out for every namespace

The module runs inside the app and sees the app's environment, as kubectl would in the same shell: HOME and KUBECONFIG to find the kubeconfig, PATH to find an exec credential plugin (aws, gke-gcloud-auth-plugin, kubelogin), and whatever that plugin reads, such as AWS_PROFILE. If a context works with kubectl but not here, check that Wayseer was started with the same variables. A desktop launcher may not set what your shell does.

Nothing connects to a cluster, and no kubeconfig is read, until a kind: kubernetes instance is configured.

OptionTypeDefaultMeaning
kubeconfigstring$KUBECONFIG, then ~/.kube/configThe kubeconfig file.
contextstringthe kubeconfig's current contextWhich cluster and user to use.
namespaceslist of stringevery namespaceOnly list objects in these namespaces.
in_clusterboolfalseUse the pod's service account, when running inside the cluster.
timeoutduration10sLongest wait for one request, 100ms to 5m.
max_replicasinteger20The most replicas the scale action sets, 1 to 10000.

With in_cluster, when Wayseer runs inside the cluster, leave out kubeconfig and context.

Credentials stay in the kubeconfig, or with the plugin it names. Nothing about them goes into Wayseer's config, and errors never show a token or the server's address. For example, a refused token shows as "the cluster refused the credentials (401 Unauthorized)".

Add one module instance for each cluster you want to see. Give each instance its own name and context.

What shows

node

down when not ready, critical when not reporting, warn under memory, disk or PID pressure or when cordoned

Attributes
hostname
ip
kubelet
os
arch

k8s/namespace

warn while being deleted

pod

the phase; critical for a container stuck in CrashLoopBackOff or an image pull error; warn while pending, when not every container is ready, or when a container restarted in the last 10 minutes

Attributes
namespace, node, phase, ready (as 2/3), restarts, ip, hostname, owner

service

critical when none of its pods is ready, warn when its selector matches no pods

Attributes
namespace
type
cluster_ip
ports

k8s/deployment

available replicas against those wanted: critical at none, warn at fewer; critical when a rollout has stalled

Attributes
namespace
replicas
ready
available
strategy

k8s/replicaset

as for a deployment

Attributes
namespace
replicas
ready
available

Each object's labels become its tags, as key=value. Old replica sets that a deployment keeps as history, scaled to nothing, are not shown.

The links are:

  • a deployment owns its replica sets, and a replica set owns its pods;
  • a service depends on the pods its selector picks;
  • a pod runs on its node;
  • every namespaced object is a member of its namespace.

In Topology, each namespace is a round group of its deployments, replica sets, pods and services, and the nodes sit in the middle with the daemon set pods on them.

Nodes carry the hostname and ip attributes that the default identity rules match. So a node that another module also shows, such as a Prometheus node-exporter target, merges with it.

Some filters:

/source:k8s status:warn
/kind:pod namespace=shop
/kind:pod restarts>0

Live updates

The module lists the cluster once, then watches it. A change shows within a fraction of a second: changes that arrive together wait 0.1 s and go in one update. Every 10 minutes the module sends the whole cluster again, which corrects anything a missed change left wrong.

If a watch stops, for example because the API server restarted, the module lists that kind of object again. Meanwhile its health shows a note such as "watch stopped; listing pods again". That is not an error: what shows is still the last known state. The module is only in error when listing fails, and then the error names what it could not list.

Events

Kubernetes events show on the Timeline, on the object they are about, and in the detail panel's events. A repeated event, such as a container backing off again, shows each time it happens, with its count.

Normal

info

Warning

warn

Warning that means something broke: BackOff, Failed, FailedCreate, FailedMount, FailedAttachVolume, FailedScheduling, Evicted, OOMKilling, SystemOOM, NodeNotReady

error

Rollout events (ScalingReplicaSet, SuccessfulCreate, SuccessfulDelete) have the kind deploy; the rest have the kind k8s. An event about an object the module does not show, such as a Job, goes on that object's namespace, with the object named first: Job/backup: …. Each event carries reason and object fields.

When the module starts, the events the cluster still keeps (usually the last hour) show as well.

Metrics

When metrics-server is installed, the module reads it every 15 seconds and offers:

cpu.utilisation

For a node, the share of its allocatable CPU in use; for a pod, the share of one core, so a pod using two cores shows 200%

Kinds
node
pod

memory.utilisation

Working-set memory as a share of the node's allocatable memory

Kinds
node

memory.used

Working-set memory

Kinds
node
pod

It keeps the last hour. Without metrics-server there are no metrics, and nothing reports the module as unhealthy. For longer history, add a Prometheus module as well.

Metrics from Prometheus

A Prometheus module with series adds metrics of your own to the pods and deployments this module finds, such as an application's queue depth, or the metric a HorizontalPodAutoscaler scales on, as kube-state-metrics reports it:

modules:
  - kind: kubernetes
    name: k8s
  - kind: prometheus
    name: prom
    options:
      url: http://localhost:9090
      series:
        - query: max by (namespace, pod, queue) (shop_queue_depth)
          metric: queue.depth
          unit: count
          description: messages waiting in the pod's queues
          entity: {labels: [namespace, pod], kind: pod}
        - query: max by (namespace, horizontalpodautoscaler) (kube_horizontalpodautoscaler_status_target_metric{metric_name="queue_depth",metric_target_type="average"})
          metric: hpa.queue_depth
          description: the queue depth the deployment's autoscaler sees
          entity: {labels: [namespace, horizontalpodautoscaler], kind: k8s/deployment}

A pod's namespace and pod labels name it as this module does, namespace/name. The second query names the deployment after its autoscaler, which works where they share a name, as they usually do. Then >grid.show metric=queue.depth kind=pod ranks every pod by its queue.

Permissions

The module needs get, list and watch on the objects it shows and on events, and get and list on metrics-server's nodes and pods. It needs nothing else. This read-only ClusterRole covers every namespace:

apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRole
metadata:
  name: wayseer-read
rules:
  - apiGroups: [""]
    resources: [nodes, namespaces, pods, services, events]
    verbs: [get, list, watch]
  - apiGroups: [apps]
    resources: [deployments, replicasets]
    verbs: [get, list, watch]
  - apiGroups: [metrics.k8s.io]
    resources: [nodes, pods]
    verbs: [get, list]

Bind it to the user or service account that Wayseer connects as, with a ClusterRoleBinding.

With access to only some namespaces, set namespaces to those namespaces and grant the same rules through a Role and RoleBinding in each of them. Nodes are cluster-wide, so a namespaced role cannot list them. The cluster then shows without nodes, and the module's health notes why. It is not an error. The same goes for events: without the right to list them, the cluster shows without its events, with a note.

Events about nodes are kept in the default namespace. With namespaces set, node events show only if default is one of them.

Actions

The module offers three actions. None runs unless the instance's actions lists it, and each waits for you to confirm it.

restart

Stamps kubectl.kubernetes.io/restartedAt on the pod template, as kubectl rollout restart does, so the pods are replaced in turn.

On
deployments
RBAC it needs
patch on deployments

scale

Sets the replica count, from 0 to max_replicas (default 20).

On
deployments
RBAC it needs
patch on deployments/scale

delete

Deletes the pod with its default grace period. Its owner, if it has one, starts another.

On
pods
RBAC it needs
delete on pods
modules:
  - kind: kubernetes
    name: k8s
    actions: [restart, scale]
    options:
      max_replicas: 10

The read-only role above allows none of them. Grant only the rules for the actions you list, in a separate role, so reading and acting stay apart:

apiVersion: rbac.authorization.k8s.io/v1
kind: Role
metadata:
  name: wayseer-act
  namespace: shop
rules:
  - apiGroups: [apps]
    resources: [deployments]
    verbs: [patch]         # restart
  - apiGroups: [apps]
    resources: [deployments/scale]
    verbs: [patch]         # scale
  - apiGroups: [""]
    resources: [pods]
    verbs: [delete]        # delete

When the cluster refuses an action, the status bar names the verb, resource and namespace that a role would need, as in "the kubeconfig's user may not patch deployments/scale in shop". It never shows the server's address, your user name or a token. With in_cluster, it reads "the pod's service account" instead.