Skip to main content
Version: Next

GPU Nodes Not Registering

Before HAMi can schedule anything, a GPU node has to complete registration. This page covers the failures that happen before scheduling: the node advertises no GPUs, the HAMi scheduler does not know the node exists, or a node that used to work silently drops out of the cluster's GPU capacity.

Typical symptoms:

  • kubectl describe node shows no nvidia.com/gpu under Capacity and Allocatable.
  • A Pod requesting nvidia.com/gpu stays Pending with 0/N nodes are available, and no FilteringFailed event from hami-scheduler.
  • nvidia-smi works on the host, but the node still contributes nothing to the cluster.
  • A node worked yesterday and stopped being selected today.

If your Pod does get a FilteringFailed event from hami-scheduler, registration already succeeded and the problem is scheduling instead. See Troubleshooting.

How registration works

Registration is not one action. It is three independent channels, and each can break on its own:

HAMi GPU node registration path
ChannelWritten byCarriesWhere to observe it
Device countDevice Plugin to kubeletAn integer count onlynvidia.com/gpu in node Allocatable
Device specificationDevice Plugin to the API serverUUID, memory, compute, model, NUMA, healthhami.io/node-nvidia-register node annotation
Liveness handshakeScheduler to the API serverA timestamphami.io/node-handshake node annotation

The count and the specification travel separately because the Device Plugin API can only report a single integer resource. A node can therefore advertise nvidia.com/gpu: 10 while the HAMi scheduler still refuses to use it, because the annotation the scheduler actually reads is missing.

The count kubelet advertises is the inflated count: physical GPUs multiplied by devicePlugin.deviceSplitCount (default 10). One physical card on a default install shows as nvidia.com/gpu: 10, not 1. See GPU Virtualization.

Step 1: Find the broken channel

Run all three checks against the affected node before changing anything:

NODE=<node-name>

# Channel 1: does kubelet advertise the resource?
kubectl get node $NODE -o jsonpath='{.status.allocatable.nvidia\.com/gpu}{"\n"}'

# Channel 2: did the Device Plugin write the device specification?
kubectl get node $NODE -o jsonpath='{.metadata.annotations.hami\.io/node-nvidia-register}{"\n"}'

# Channel 3: is a Device Plugin Pod actually running there?
kubectl get pods -n kube-system \
-l app.kubernetes.io/component=hami-device-plugin \
-o wide --field-selector spec.nodeName=$NODE

Match the result to the section to read next:

ResultBroken channelGo to
No Device Plugin Pod on the nodeNothing is registeringCase 1
Pod exists but is not RunningPlugin cannot startCase 2
Empty nvidia.com/gpu, Pod RunningPlugin to kubeletCase 3
nvidia.com/gpu set, annotation emptyPlugin to API serverCase 4
Both set, Pod still PendingScheduler does not see the nodeCase 5

Case 1: no Device Plugin Pod on the node

The HAMi NVIDIA Device Plugin DaemonSet carries a node selector. Its chart default is:

devicePlugin:
nvidiaNodeSelector:
gpu: "on"

A node without that label never receives a Device Plugin Pod, so none of the three channels start. This is the single most common cause of a GPU node contributing nothing.

Check

kubectl get node $NODE --show-labels | tr ',' '\n' | grep gpu
kubectl get daemonset -n kube-system hami-device-plugin \
-o jsonpath='{.spec.template.spec.nodeSelector}{"\n"}'

Fix

kubectl label node $NODE gpu=on --overwrite

Then wait for the DaemonSet to place a Pod:

kubectl rollout status daemonset/hami-device-plugin -n kube-system

If the label is already present and correct, check that the node is not tainted in a way the DaemonSet does not tolerate:

kubectl describe node $NODE | grep -A 3 Taints

Add matching entries under devicePlugin.tolerations if needed. Node labelling is also covered in Prerequisites.

Case 2: the Device Plugin Pod does not stay running

kubectl logs -n kube-system -l app.kubernetes.io/component=hami-device-plugin \
-c device-plugin --tail=100
kubectl describe pod -n kube-system -l app.kubernetes.io/component=hami-device-plugin

NVML initialization failure

nvml Init err: ERROR_LIBRARY_NOT_FOUND

The Device Plugin treats every NVML failure during a device scan as fatal and exits, so the container ends up in CrashLoopBackOff rather than running degraded. The same fatal path is taken for nvml get memory error, nvml get name error, and nvml new device by index error.

This almost always means the container did not receive the driver, which in turn means nvidia-container-runtime is not the default runtime on the node:

containerd config dump | grep default_runtime_name

The output must be nvidia. If it is not, follow Prerequisites, then restart the container runtime. On GPU Operator 25.10 and later the default runtime deliberately stays runc; that case needs devicePlugin.runtimeClassName=nvidia instead, and is covered in Troubleshooting.

Stuck in Init

Waiting for /run/nvidia/validations/toolkit-ready...

The toolkit-validation init container blocks until the NVIDIA Container Toolkit writes a toolkit-ready file under devicePlugin.gpuOperatorToolkitReady.hostPath (default /run/nvidia/validations). The gate is off by default and is meant for GPU Operator clusters. If it was enabled on a cluster without GPU Operator, that file is never created and the init container waits forever:

helm upgrade hami hami-charts/hami -n kube-system --reuse-values \
--set devicePlugin.gpuOperatorToolkitReady.enabled=false

Node name not resolved

The Device Plugin patches the node named by its NODE_NAME environment variable, which the chart populates from spec.nodeName. On charts older than v2.3.10 the variable was called NodeName, and a mismatched image and chart pair leaves the plugin unable to identify its own node. Upgrade rather than patching by hand:

helm upgrade hami hami-charts/hami -n kube-system --reuse-values

Case 3: the resource never appears in Allocatable

The Pod is Running and NVML works, but nvidia.com/gpu is absent. The Device Plugin registered with the API server but not with kubelet.

Check

kubectl logs -n kube-system -l app.kubernetes.io/component=hami-device-plugin \
-c device-plugin --tail=200 | grep -i -E "register|socket|kubelet"

ls -l /var/lib/kubelet/device-plugins/ # on the node itself

The plugin registers over kubelet.sock in that directory. The chart mounts it from devicePlugin.pluginPath, which defaults to /var/lib/kubelet/device-plugins. If your distribution relocates the kubelet root, the plugin writes its socket somewhere kubelet never reads and registration silently never completes.

Fix

Confirm the real path on the node, then point the chart at it:

# on the node
ps aux | grep kubelet | grep -o '\--root-dir=[^ ]*'

helm upgrade hami hami-charts/hami -n kube-system --reuse-values \
--set devicePlugin.pluginPath=<kubelet-root>/device-plugins

Restarting kubelet also forces every Device Plugin to re-register, which is a fast way to confirm the socket path is the problem.

Case 4: the register annotation is missing or stale

This is the case that most often looks like a scheduler bug. kubelet advertises the GPUs, so kubectl describe node looks healthy, but hami-scheduler never places a Pod there.

Read the annotation

kubectl get node $NODE \
-o jsonpath='{.metadata.annotations.hami\.io/node-nvidia-register}' | jq .

Expected output for one 24 GiB card on a default install:

[
{
"id": "GPU-fc28df76-54d2-c387-e52e-5f0a9495968c",
"count": 10,
"devmem": 24576,
"devcore": 100,
"type": "NVIDIA-NVIDIA L40S",
"mode": "hami-core",
"health": true
}
]
FieldMeaningSource
idGPU UUIDNVML
countLogical split countdevicePlugin.deviceSplitCount
devmemSchedulable memory in MiBPhysical memory times deviceMemoryScaling
devcoreSchedulable compute percentagedeviceCoreScaling times 100
typeModel, prefixed with NVIDIA-NVML
numaNUMA nodesysfs
modehami-core or migPlugin operating mode
healthDevice healthDevice Plugin health check
Zero-valued fields are omitted

The annotation is serialized with omitempty, so any field whose value is zero or false is absent rather than shown. An unhealthy GPU has no health key at all, it does not appear as "health": false. The same applies to "numa": 0 and "index": 0. Read a missing health key as unhealthy, not as healthy.

The annotation is only rewritten when it changes

The Device Plugin rescans devices every 30 seconds, but it compares the newly encoded list against the last one it wrote and skips the patch when they are identical:

Device info unchanged, skipping annotation update

That line appears at verbosity -v=3. At the default verbosity, a real update logs:

Updating node annotations with 1 device(s)

The consequence: an old timestamp on the annotation is normal and is not evidence of a stuck plugin. Conversely, deleting the annotation by hand does not get it rewritten within 30 seconds, because the plugin's in-memory cache still matches what it thinks it wrote. Restart the Pod instead. See Force a re-registration.

The patch is rejected

patch node error nodes "gpu-node-1" is forbidden: User "system:serviceaccount:kube-system:hami-device-plugin" cannot patch resource "nodes"

The ServiceAccount lost permission to patch nodes, usually after a partial upgrade or a hand-edited ClusterRole:

kubectl auth can-i patch nodes \
--as=system:serviceaccount:kube-system:hami-device-plugin

Reinstalling or upgrading the chart restores the RBAC objects.

Case 5: the scheduler does not see the node

Both the resource and the annotation are correct, but Pods still do not land. The node is missing from the scheduler's in-memory cache.

Check the scheduler log

kubectl logs -n kube-system deploy/hami-scheduler -c vgpu-scheduler-extender --tail=200

The registration loop runs every 15 seconds, and also on node events and on leader changes. Raise verbosity to -v=5 to see per-node decisions, which are otherwise silent:

Using label selector for list nodes
Listed nodes
Processing node
Failed to get node devices

Failed to get node devices for your node means the scheduler read the annotation and rejected it. There are three ways that happens: the annotation key is absent, the JSON does not decode, and the decoded list is empty. The second and third also log at the default verbosity:

failed to decode node devices
no nvidia gpu device found

A hand-edited annotation is the usual cause of a decode failure.

Cause A: a node label selector excludes the node

The scheduler can be restricted to a subset of nodes with scheduler.nodeLabelSelector. It is commented out in the chart by default, so an unmodified install lists every node. If it was set, the value is echoed on startup:

kubectl logs -n kube-system deploy/hami-scheduler -c vgpu-scheduler-extender \
| grep "label selector"

Any node missing those labels is never registered, no matter how healthy it is.

Cause B: the replica you are reading is not the leader

Only the leader performs registration. Every other replica logs:

Scheduler is not leader yet, skipping ...

With more than one hami-scheduler replica, a quiet log is expected on followers and proves nothing. Identify the leader before concluding the loop is dead:

kubectl get lease -n kube-system | grep hami

The handshake annotation

hami.io/node-handshake is a liveness marker maintained by the scheduler, not by the Device Plugin. Reading it backwards causes a lot of wasted debugging, so it is worth stating the actual behavior:

  • When the annotation is absent, or its value does not contain Requesting, the scheduler stamps it with Requesting_<timestamp> and treats the node as healthy.
  • While the timestamp is less than 60 seconds old, the node is healthy.
  • Once it is older than 60 seconds, the scheduler checks the node's allocatable nvidia.com/gpu. If that is still greater than zero, the node stays healthy and nothing is cleaned up.
  • Only when the handshake has expired and allocatable has dropped to zero does the scheduler run node cleanup: it drops the node's devices from its cache and deletes the handshake annotation.
kubectl get node $NODE -o jsonpath='{.metadata.annotations.hami\.io/node-handshake}{"\n"}'
# Requesting_2026-08-15 09:12:44

Two practical consequences:

  • An old Requesting_ timestamp is not a fault. It is the normal steady state, and on its own it never removes a node. Do not tune anything based on it.

  • The real trigger for a node leaving the scheduler's cache is allocatable falling to zero, which means the Device Plugin stopped reporting to kubelet. Debug that, not the handshake. The scheduler logs the removal:

    Device is unhealthy, cleaning up node
NVIDIA uses an unsuffixed key

The NVIDIA handshake key is hami.io/node-handshake. Other vendors use a suffixed form such as hami.io/node-handshake-dcu or hami.io/node-handshake-xpu. A kubectl get node -o yaml | grep node-handshake-nvidia returns nothing on an NVIDIA node, which is expected.

Force a re-registration

Restarting the Device Plugin Pod clears its in-memory device cache and forces a fresh annotation patch on the next scan. This is the correct recovery for a missing, truncated, or hand-edited register annotation:

kubectl delete pod -n kube-system \
-l app.kubernetes.io/component=hami-device-plugin \
--field-selector spec.nodeName=$NODE

Confirm the cycle completed:

kubectl logs -n kube-system -l app.kubernetes.io/component=hami-device-plugin \
-c device-plugin --tail=50 | grep "Updating node annotations"

kubectl get node $NODE \
-o jsonpath='{.metadata.annotations.hami\.io/node-nvidia-register}' | jq length

The scheduler picks the node back up within one registration cycle, so allow about 15 seconds before retesting with a Pod.

If the node still does not register, collect the following before opening an issue:

kubectl get node $NODE -o yaml > node.yaml
kubectl logs -n kube-system -l app.kubernetes.io/component=hami-device-plugin \
-c device-plugin --tail=500 > device-plugin.log
kubectl logs -n kube-system deploy/hami-scheduler \
-c vgpu-scheduler-extender --tail=500 > scheduler.log

Validation environment

The behavior described on this page was verified by reading the HAMi source at v2.9.0 and on master, specifically pkg/device-plugin/nvidiadevice/nvinternal/plugin/register.go, pkg/device/devices.go, pkg/device/nvidia/device.go, pkg/scheduler/scheduler.go, and charts/hami/values.yaml. Chart defaults quoted here are the NVIDIA values from charts/hami. Intervals, the 60-second handshake window, and log strings can change between releases; check the source for the version you run before relying on an exact number.

CNCFHAMi is a CNCF Incubating project