Node Drain, Upgrade & Recovery — Challenge
Take a node out of service without taking the application with it, and find out which workloads were never ready for it.
- Time
- 27 min
- Level
- Advanced
- Objectives
- 4 objectives
- Cost
- Free
Where this fits in the platform
Already built
This lab adds
- A node taken out of service with the application still up
Which lets you
—
Before you start
You will need
- kind (multi-node)
- kubectl 1.28+
You do not need these already — the lab environment below provides them.
You will be able to
- Cordon and drain a node safely
- Protect availability during voluntary disruption with a PDB
- Recognise workloads that cannot survive rescheduling
Cost — Free
— a multi-node kind cluster. See the setup note; a single node cannot demonstrate rescheduling.
The goal#
Achieve the same outcome as Node Drain, Upgrade & Recovery, from an empty starting point, without the steps.
The cluster needs a Kubernetes upgrade. That means taking each node out of service in turn, and the first one you try teaches you which of your workloads were only ever running by luck.
Hands-on environment
Run this lab in a real terminal, free and in your browser. The environment is temporary and yours alone — break it as much as you like.
Start the challengeOpens in Killercoda, in a new tab — keep this page open for the steps.
This is the guided lab's environment — the same machine, with its walkthrough on the left. Work from the task above and leave those steps alone until you are done, or you are reading the answers.
Run it on your own machine
Run this lab on your own machine. One command starts the environment, with everything the lab needs already installed:
You will need:
- docker
- kubectl
- kind
git clone https://github.com/EgyKode/EgyKode-lab.git
cd EgyKode-lab
./egykode start k8s
./egykode shellYou need Docker and Git installed. Everything else runs inside the environment. The first start downloads it and takes a few minutes; later starts are seconds.
Not sure what you already have? Run: npm run doctor — it checks and changes nothing.
Anything you tick here is your own record. EgyKode cannot see inside that terminal, so the success criteria stay self-assessed even when the environment checks your work for you.
What must be true when you are done
Step 1 of 3
What must be true when you are done#
- A node is drained with no failed requests to the application.
- A PodDisruptionBudget blocks a drain that would breach availability.
- You identified at least one workload that could not be rescheduled cleanly, and why.
- The node returns to service and receives Pods again.
Rules#
- Do not open the guided lab until you are finished, or until the same problem has held you up for 20 minutes.
- Documentation is allowed and encouraged.
- Verify every criterion with a command whose output you can read.
If you get stuck#
- What did you expect, exactly?
- What happened instead — the error text, not a paraphrase?
- Which layer is that error from?
- What is the smallest command that proves the layer below is fine?
You are done when
0 of 4
The concept behind it
Next up
Lab 58 of 59 on the project path