Booting every board, every night

Share
Booting every board, every night
Armbian automated testing · what it does today, and what it needs to do next

Armbian builds for hundreds of boards. A kernel bump that is clean on one SoC can leave another unbootable and no build log will tell you. So there is a rack, and the rack boots them.

Compiling is not evidence. A kernel that builds, packages that install, an image that flashes — none of it proves a board comes back after reboot. Armbian supports a very large number of boards across a dozen vendors and twenty SoC families, maintained largely by volunteers who each own one or two of them. The failure mode that hurts is not a broken build. It is a change that builds perfectly and quietly bricks the boot path on hardware nobody happened to have on their desk that week.

The automated test facility exists to close that gap. It is a rack of real boards, wired for remote power, that installs what Armbian published and then tries to use it: upgrade, switch kernel branches, reboot repeatedly, measure, report. What follows is what it does today, the hardware it takes to run, where AI fits into maintaining the stack around it, and the change we want to make next: moving hardware validation from nightly builds to pull requests.

What is in the rack

Fleet composition, as recorded in inventory

65 Boards under test 62 Distinct models 22 Vendors 20 SoC families

The boards are the point, and the spread is deliberate. Rockchip, Amlogic, Allwinner, NXP, Marvell, Broadcom, TI, Samsung, Qualcomm, SpacemiT and x86 are all represented, because that is where the differences live; a change to a shared kernel config lands differently on rockchip64 than on meson64, and only one of them will tell you so.

Vendors currently on the bench include Radxa, FriendlyELEC, Khadas, Orange Pi, Banana Pi, Odroid, Raspberry Pi, Pine64, SolidRun, Mekotronics, Kobol, Cubietech, Udoo, Inovato and Arduino. Of the 65, 56 are active and 9 are in a failed state at the moment article was done boards that need a hand in the rack. That number is itself a signal: a board that has stopped answering is a board that stopped being able to report, which is usually more interesting than a test that merely failed.

Every board is registered in NetBox, which is the source of truth for the whole system. Not just an asset list the orchestrator reads it to decide what a board can be tested with. Power is derived from cabling topology rather than a field someone typed: a board's power port is cabled to an outlet on a controller device, and the controller carries the driver. If the cable is not in the model, the board is treated as having no managed power, and the tests that need a hard power cut are skipped rather than faked.

The hardware it takes

Per bench, and the shared infrastructure behind it

The cost of adding a board is not the board. It is the wiring around it. Each bench needs some subset of the following, and which subset a board has determines which tests it is eligible for.

Remote power Required

A way to cut and restore power without a human. Three backends are in use: PoE switch ports for anything powered over Ethernet, a switched rack PDU for mains devices, and a 24-channel relay board driven over GPIO for DC-barrel and USB-powered boards. Without this there is no cold-boot test, only a soft reboot which is exactly the case that tends to pass while the real one fails.

Network Required

Wired Ethernet on a managed switch. This is the control channel (the harness drives boards over SSH), the measurement channel for throughput tests, and on PoE benches, the power channel too. A board that only has Wi-Fi is testable but much harder to recover.

Power metering Strongly recommended

PoE ports meter per-port draw, so the harness samples watts throughout a run. This is how you tell a board that genuinely rebooted from one whose SoC never stopped executing the power trace is flat across what was supposed to be a power cycle.

Serial console Recommended

USB-to-TTL adapters on a powered hub, exposed over the network as named consoles. Without one, a board that fails to boot tells you nothing at all: you get silence and have to guess. Several boards in the rack currently lack one, and every hang on those is an investigation that stalls at "it does not come back".

Clean-flash path Not yet wired

An SD-card switcher that presents the card to a flasher host or to the board, or USB access for vendor recovery modes. This is what makes a board recoverable from software that does not boot. The framework supports it; no bench in the rack is currently wired for it, which is the single biggest constraint on what comes next.

Behind the benches sits the shared infrastructure, the part people underestimate when they picture "a few boards on a shelf":

  • Switching. Three core switches and four access switches, several of them PoE, which do double duty as the fleet's power controllers. Between them they carry the control network, the test traffic and the power for a large fraction of the boards. This is no longer just a 1 GbE network: newer boards increasingly come with 2.5 GbE, while the uplinks and core need 10 GbE to aggregate traffic from many boards running tests in parallel. Without that headroom, the lab network itself becomes the bottleneck and network-performance results stop measuring the board under test.
  • Power delivery. A switched rack PDU, multi-channel DC supplies, multiple 16-way USB supply, switched power strips, and a UPS in front of all of it because a lab that loses power mid-write is a lab that corrupts SD cards.
  • Console and out-of-band. Dedicated console hosts, including a KVM device for the machines that need screen-level access.
  • Compute. Six servers in the same site, including two Ampere-class ARM machines, running the CI runners that execute the build and test jobs and the services the fleet depends on.

All of it is scripted. Each class of hardware has a small command-line tool with the same shape — status, on, off, powercycle so the harness does not care whether a board is switched by a PoE port, a PDU outlet or a relay channel. It resolves the path from the inventory model and calls whichever tool matches.

What a nightly run does

The cycle, end to end

The fleet does not stay powered. A full cycle brings it up, tests it, and puts it back to sleep. The last step runs even when everything before it failed.

  1. Power on The fleet is brought up from the PDU and the boards are given several minutes to boot.
  2. Scan and reconcile Every board is probed and the inventory updated: what version it is running, what kernel, when it was last seen. Drift between the model and reality is recorded rather than assumed away.
  3. Sync maintainer keys Board maintainers' SSH keys are pushed to the boards they own, so the person responsible for a board can log into the actual unit that failed.
  4. Run the board pipeline A matrix job per board, in parallel across the runners. This is the part below.
  5. Scan again Post-test state is captured; what the run left behind, not just what it reported.
  6. Power off Always, including after a failure. A rack left powered on a failed run is a rack that cooks.

Inside step four, each board runs the same pipeline. It upgrades to the nightly repository, reboots, and then walks every kernel branch that board is configured to test:

StepWhat it does
upgradePoint at the nightly repository and install what is currently published.
rebootVerify that the board comes back with what it already had installed.
For each branchRepeat the steps below for each kernel branch (current, edge, …).
↳ kernel-switchInstall that branch's kernel and verify that it is fully configured.
↳ rebootPerform warm reboots, followed by a cold power cycle where switched power is available.
↳ hw-performanceTest CPU, memory, disk and temperature.
↳ dvfsVerify that the governor actually reaches the frequencies it claims.
↳ network-iperfMeasure throughput on each cabled network interface.
↳ store-versionsRecord exactly what is installed and running.
restore-stablePut the board back the way it was found.

Two details in there carry more weight than they look. The reboot module does warm reboots followed by a cold power cycle, because those fail differently: a board can survive reboot indefinitely and still not come back from a real power cut. Testing both is the only way to catch both failure modes. And kernel-switch verifies that the package is genuinely configured rather than trusting an exit code, because the interesting failures can leave a kernel half-installed while every command reports success.

What the results look like

Current results are published publicly at docs.armbian.com/status/board-tests, refreshed as runs complete.

What it actually catches

Here are few examples from recent runs.

Reboots that hang on both kernels

One board fails every reboot attempt on both of its kernel branches. The power trace shows the draw holding steady through what should have been a restart — the SoC never stopped executing.

That points at firmware rather than the kernel, but the issue remains unresolved. It is also a good example of the console gap: that bench has no serial console, so the investigation is relying on power measurements and inference instead of a boot log.

A test that was wrong about x86

The frequency-scaling check assumed that a governor under load should reach at least 95% of the maximum advertised frequency. That works for ARM cpufreq, but not for Intel and AMD, where the driver manages turbo behaviour itself and the advertised maximum is not a promise.

Every x86 board was therefore failing a test that was itself wrong. The test was fixed by detecting the driver and applying the check only where it makes sense.

Telling a broken board from a broken runner

Self-hosted runners occasionally drop mid-job: the test finishes green, then the job dies later and the run is marked failed. Retrying everything would be wrong. A board pipeline power-cycles hardware and takes about half an hour, so retrying a genuine failure wastes rack time and keeps cycling a board that may already be unwell.

The retry logic therefore re-runs only jobs whose test step did not fail. A real hardware or test failure stays red rather than being hidden by an automatic retry.

Where AI fits

And where it explicitly does not

A large share of the tooling described here — the test modules, hardware control scripts, inventory reconciliation and documentation — is now written and maintained with AI assistance. That is worth being plain about, including the parts that go wrong.

What AI is genuinely good at is shortening the distance from symptom to candidate explanation. A board fails; there is a transcript, a package state, a power trace, a kernel version and forty thousand lines of shell across several repositories. Correlating those quickly, proposing a mechanism and drafting a patch is work that used to take an evening and can now take a few minutes.

What it is bad at is knowing when it is wrong. In the course of this work, an analysis confidently concluded that a particular board had never appeared in the test results. The conclusion came from a sample that covered about three-quarters of the archive and happened to miss both of that board's records. The reasoning was sound; the evidence was partial; nothing in the output said so.

The point

AI shortens the distance from symptom to patch. It does not shorten the distance from patch to proof. Those are different problems, and only one of them has been made easier.

That is precisely why the rack matters more now, not less. If proposing changes gets cheaper while validating them does not, the amount of unverified change grows faster than our ability to verify it. And the most dangerous bugs here are exactly the ones that compile cleanly, install without error, and fail only when a specific board is asked to come back from a power cut.

A machine that is confidently wrong is a fine collaborator as long as something downstream can check it. Sixty-five boards that will actually try to boot are that something.

So the working rule is simple: AI is free to propose anything, but nothing reaches a board on its say-so. Hardware evidence is the gate, and every conclusion that matters is expected to point at a run, a trace or a package state rather than an argument.

What is next: validation at pull-request time

The change we want to make

Today the facility validates nightly builds. Whatever was published overnight gets installed on real hardware and exercised. That is genuinely useful, and it is how the bugs above were found, but it is validation after the fact. By the time a board fails, the change is already in the nightly repository and, depending on timing, already on users' machines.

The goal is to move that gate earlier: a pull request that touches a board family gets that family's boards booted before it merges. Not the whole rack for every PR, just the boards the change can actually reach.

Several things have to be true first, and they are worth stating honestly rather than as a roadmap of solved problems.

Selection

A diff has to map to boards. A change to a board configuration file implicates that board; a change to a kernel family implicates every board in it; a change to the packaging code implicates everything. Select too broadly and every PR occupies the entire rack for half an hour. Select too narrowly and the one board that would have caught the regression never gets tested.

Artifacts

A PR build currently proves a compile. Hardware testing needs installable packages the rack can pull and install exactly as a user would, which means PR builds must publish to an isolated repository the fleet can reach. Our current build machinery is already stretched by existing CI and release workloads. Making hardware testing part of routine PR validation will require additional build capacity, storage and package-publishing infrastructure.

Time budget

A full board pipeline takes around thirty minutes. That is acceptable overnight and much too slow as a merge gate, particularly when a nightly cycle already has the rack. This needs a short profile for PRs: install, switch kernel, reboot warm and cold, and confirm it is running what was installed. Performance and throughput work can stay in the nightly run. It also needs real contention handling, so a PR does not simply queue behind a fleet cycle.

Recovery

This is the hard one, and it is a hardware problem rather than a software one. Testing unmerged code means occasionally installing a kernel that does not boot. Right now every bench in the rack is tested in place; there is no remote path to reflash a board whose boot is broken. A board that a bad PR bricks is a board that stays bricked until somebody walks to the rack.

So PR-stage testing starts on benches that can be recovered remotely: SD-card switchers for boards that boot from removable media, and vendor recovery modes over USB for those that support it. Wiring that up across the fleet is the prerequisite that gates everything else, and it is where the next round of hardware effort goes. Boards that cannot be recovered remotely stay on nightly testing, where a failure is an inconvenience instead of a dead unit.

Signal quality

A merge gate that fails for reasons unrelated to the change is worse than no gate: it trains people to ignore it. The distinction between a board failure and a broken bench, already important for nightly runs, becomes critical the moment a red result blocks a merge.

Helping

Infrastructure, capacity, partnerships

There are three practical ways to move this forward.

Bench infrastructure. Remote recovery is the immediate blocker for PR-stage testing, particularly SD-card switchers and the wiring around them. The network also needs to grow with the hardware being tested: managed PoE switches with 2.5 GbE access ports and 10 GbE uplinks, plus PoE splitters for boards that cannot be powered directly over Ethernet. These are not just infrastructure upgrades; they expand what can be tested, measured and recovered without somebody standing in front of the rack.

Build capacity. PR-stage hardware testing needs PR builds, and build queue time is already a constraint on how quickly a fix reaches hardware. More build capacity, storage and package-publishing infrastructure are needed to move validation from nightly runs into the pull-request workflow. If you can host a build server, the requirements are documented, and the current fleet is listed on the build machinery page.

Hardware partnerships. For vendors, the facility provides a way to keep hardware under continuous validation rather than testing it once around release. Devices covered by support and maintenance agreements can become part of the permanent test fleet, where kernel updates, upgrades, reboots, networking and other regressions are exercised on real hardware as Armbian evolves. The value is not the board itself; it is keeping that board working over its supported lifetime.

Read more