Patch hosts one at a time
This guide builds a rolling patch: you give it a list of Linux hosts, and it upgrades their packages one host at a time through the KloudMate agent. A host reboots when an update needs it, and the next host starts only after the previous one is back and has no failed services. If a host fails, the rollout stops there, so a bad update takes down one host instead of the whole fleet.
The rollout is two workflows. Patch one host does the work on a single host. Patch hosts one at a time reads your list and calls Patch one host for each host in turn:
Patch one host is also useful on its own, when you want to patch a single host by hand.
Before you start
Section titled “Before you start”- You need the Developer role in a KloudMate workspace whose plan includes workflows.
- Install the KloudMate agent on each host, and allow runbook scripts on each one.
- The hosts use
apt,dnf, oryumto manage packages, andsystemd. - Patch hosts that can go down one at a time without an outage, such as web servers behind a load balancer.
Step 1: Import and publish Patch one host
Section titled “Step 1: Import and publish Patch one host”Import this workflow first, because the rollout calls it and can only call a published workflow:
- Copy the YAML below.
- Open Workflows, click Import, paste the YAML, and click Import.
- Click Open workflow, then click Publish.
Step 2: Import Patch hosts one at a time
Section titled “Step 2: Import Patch hosts one at a time”Import the rollout the same way. The import finds Patch one host by its uid and fills in the Call Workflow step for you:
Publish it when you’re ready. Its trigger is Manual / API, so it only runs when someone clicks Run.
How each step works
Section titled “How each step works”Find the host’s agent
Section titled “Find the host’s agent”A Check a Host Agent step looks up the agent by Host name, {{ trigger.inputs.host }}. The name must match the host name the agent registered with, as shown under Settings → Agents.
A Branch then checks that steps.agent.output.status equals reachable, which means the agent checked in within the last 5 minutes. For any other status, such as unreachable or not_found, the Else block runs a Run Code step that throws an error:
The error fails the run with a message that names the host and the reason, and a failed call stops the rollout. Without the check, a command for a host that’s down would wait until the step timed out.
Upgrade packages
Section titled “Upgrade packages”A Run Command on Host step runs the upgrade on the host, with Host set to {{ steps.agent.output.agent_id }} from the lookup:
The lines run as one shell script. set -e makes the script stop, and the step fail, at the first command that fails, instead of carrying on after a failed apt-get update.
An upgrade can take several minutes, so the step allows for it in two places. Command timeout (seconds) is 720, which stops the upgrade on the host after 12 minutes. The step’s Timeout, under Settings, is 15m, which also covers the time the agent takes to pick up the command and report back.
Reboot if an update needs it
Section titled “Reboot if an update needs it”A second Run Command on Host step checks whether the host needs a reboot. On Debian and Ubuntu, an update that needs one creates /var/run/reboot-required. On Red Hat-based systems, needs-restarting -r exits with a non-zero code when a reboot is needed:
shutdown -r +1 schedules the reboot a minute later instead of rebooting at once. That gives the agent time to report the step’s result before the host goes down.
Wait for the reboot
Section titled “Wait for the reboot”A Branch checks whether the output, steps.reboot.output.stdout, contains rebooting. If it does, a Wait / Delay step waits 3 minutes for the reboot to happen.
A Wait until step then checks the agent with Check a Host Agent until steps.agent_again.output.seconds_since_checkin is less than 90. The agent checks in every 60 seconds, so a check-in less than 90 seconds old means the host is back up. The step’s Timeout is 15m. If the host isn’t back by then, the step fails, and so does the rollout.
Check that nothing failed to start
Section titled “Check that nothing failed to start”The last Run Command on Host step lists the systemd units that are in the failed state:
If any unit failed, the script prints them and exits with 1, which fails the step. The failed units appear in the step’s output in run history.
Read the host list
Section titled “Read the host list”In the rollout, the hosts input is Long text. A Run Code step turns it into a list, splitting it at line breaks, spaces, and commas, so you can paste one host name per line or a comma-separated list:
Its Input is {"text": "{{ trigger.inputs.hosts }}"}, and the list comes back as steps.split.output.output.hosts.
For each host
Section titled “For each host”A Loop runs a Call Workflow step for each host, with Items set to {{ steps.split.output.output.hosts }} and Run at once at 1. The call’s Inputs pass the host as host, which Patch one host reads as {{ trigger.inputs.host }}.
With Run at once at 1, the loop waits for each call to finish before it starts the next one, and a call that fails stops the loop. The rollout then fails, and run history shows which host failed and why.
A Wait until can’t go inside a loop, which is why the steps for one host are a workflow of their own. See Wait until.
Email the result
Section titled “Email the result”When every host is done, a Send Email step emails the address in the notify_email input, and lists the hosts it patched.
Step 3: Try it on one host
Section titled “Step 3: Try it on one host”- Open Patch one host, click Run, and enter a host that can go down for a few minutes.
- Open the run from Workflows → Runs, and follow the steps as they run. A host that reboots shows Waiting while the workflow waits for it.
- When that works, run Patch hosts one at a time with two hosts, one per line, and your email address.
A run patches at most 100 hosts, which is the most one loop handles.
Adapt the workflow
Section titled “Adapt the workflow”- Ask for approval first. Add an Approval step before For each host in the rollout, with a message that lists
{{ steps.split.output.output.hosts | join: ", " }}. - Check your application too. Add a command to Check that nothing failed to start that calls your service’s health endpoint on the host, such as
curl -fsS http://localhost:8080/health. - Install security updates only. On Red Hat-based systems, change the upgrade line to
dnf upgrade --security -y -q. - Patch on a schedule. Change the rollout’s trigger to a monthly Schedule, and keep the host names in a constant as a comma-separated list. Then set the Input of Read the host list to
{"text": "{{ vars.patch_hosts }}"}, and To on Email the result to your team’s address.
Related
Section titled “Related”- Run Command on Host for how commands and timeouts work.
- Control flow for Loop, Call Workflow, and Wait until.
- Run Code for the sandbox the host list runs in.