Picture a Nagios Core server that has been running since 2014: 1,900 service checks, four different admins’ fingerprints in the config over the years, and one custom plugin in /usr/local/nagios/libexec that is a shell script SSHing into a legacy database box and grepping a log file. Nobody remembers it exists until someone lists every command in use — and it turns out to be the only thing watching the nightly billing export. Rebuild the monitoring “from the templates” without that inventory step, and it vanishes silently, with an unhappy finance team as the first sign.
Migrations off Nagios fail in exactly that way: not with a dramatic outage, but with checks that quietly stop existing. This plan is built around making the old and new systems prove they cover the same ground before anything is switched off. Our examples target Icinga 2, which is the smoothest path for most Nagios shops, with notes on Checkmk and Zabbix.
Step 1: take a real inventory
Do not inventory from the config files by reading them; Nagios configs grow templates, use chains and hostgroup assignments that hide what actually runs. Use the object cache Nagios writes at startup, which contains every resolved object:
# Path varies: /usr/local/nagios/var/objects.cache or /var/cache/nagios4/objects.cache
CACHE=/usr/local/nagios/var/objects.cache
grep -c '^define host {' "$CACHE"
grep -c '^define service {' "$CACHE"
grep -E '^\s+check_command' "$CACHE" | awk '{print $2}' | cut -d'!' -f1 | sort | uniq -c | sort -rn
The last line gives every check command in use with a count. Build a spreadsheet with one row per command and these columns:
- Command name and plugin path.
- Source — standard Monitoring Plugins (
check_disk,check_http,check_ping), a community plugin, or something written in-house. - Execution model — runs on the Nagios host, via NRPE on the target, via SSH (
check_by_ssh), or passive via NSCA. - Who gets notified, from contacts and contact groups.
- Keep, replace or drop. Expect 10–20% of checks to be dead weight: hosts decommissioned years ago, services permanently acknowledged.
Also export event handlers, escalation objects, scheduled downtimes and anything that reads status.dat — dashboards and scripts often depend on it.
Step 2: choose the target by how much you want to change
- Icinga 2 keeps the plugin model, the soft/hard state logic and the notification concepts, and adds a proper config DSL, a REST API and clustered, TLS-secured distributed monitoring. The lowest-friction target. See Icinga vs Nagios for the details.
- Checkmk can run Nagios plugins through its “Integrate Nagios plugins” rule (active checks) or MRPE on the agent side, but its real value is auto-discovery with its own agent. Expect to replace many plugins with native checks.
- Zabbix can call plugins as external checks or agent
UserParameters, but it treats output as item values, not states, so each check has to be rebuilt as an item plus trigger. It is a redesign, not a port.
Step 3: reuse plugins — they mostly just work
The Nagios plugin contract is simple: exit code 0 (OK), 1 (WARNING), 2 (CRITICAL), 3 (UNKNOWN), one line of output, optional performance data after a |. Icinga 2, Naemon, Checkmk and Shinken-derived tools all honor it, so your custom scripts carry over unchanged.
For standard plugins, Icinga 2 ships the Icinga Template Library (ITL) with ready CheckCommand definitions — disk, http, ping4, load, procs, nrpe and many more. For in-house plugins, write a CheckCommand once:
object CheckCommand "billing_export" {
command = [ PluginDir + "/check_billing_export.sh" ]
arguments = {
"-H" = "$address$"
"-a" = "$billing_max_age$"
}
vars.billing_max_age = 26
}
Then replace hundreds of individual Nagios service definitions with apply rules keyed on host variables:
apply Service "disk" {
check_command = "disk"
vars.disk_wfree = "15%"
vars.disk_cfree = "8%"
command_endpoint = host.vars.agent_endpoint
assign where host.vars.os == "Linux"
}
NRPE keeps working through the ITL nrpe command while you roll out the Icinga agent, so you do not have to touch every target on day one.
Step 4: run both systems in parallel
This is the step people skip, and the one that catches the billing-export script.
- Point the new system at the same targets with the same intervals. Let it check everything, but send its notifications to a test channel only.
- Run for at least two weeks, ideally spanning a month-end and a patch window, so you see batch jobs and reboots.
- Reconcile daily. Compare the set of non-OK services in both systems. Any service present in Nagios but missing in Icinga is a coverage gap; any service with different states is a threshold or plugin argument difference.
- Compare counts. Hosts and services in the new system should equal the “keep” rows from your inventory. Script it: Nagios
objects.cacheon one side, the Icinga 2 REST API (/v1/objects/services) on the other.
Check load on targets during this period — you are running every NRPE check twice.
Step 5: cut over deliberately
- Freeze config changes in Nagios a few days before cutover.
- Move notifications: enable real contacts in the new system, then disable notifications globally in Nagios (
enable_notifications=0or via the web UI). Keep Nagios checking. - Move integrations: ticketing, chat webhooks, status pages, anything that read
status.dator the Nagios CGIs. - Keep Nagios running read-only for two to four weeks as a reference if someone asks “did this ever alert before?”
- Archive the config and
objects.cache, then shut it down.
Common mistakes
- Trusting templates instead of the object cache. The resolved config is the truth.
- Porting every check. Use the migration to drop what nobody acts on — and then tune the rest with our alert noise guide.
- Forgetting passive checks. NSCA-fed services show up only when something sends to them; list the senders.
- Changing thresholds during the migration. Port first, tune later, or you cannot tell a coverage gap from a deliberate change.
- No plan for dependencies. Nagios
parentsneed to become IcingaDependencyobjects, or the first switch reboot will page the whole team.
Related reading
Start with our Nagios Core review and Icinga review, then compare the wider field in the open-source monitoring category and server and service monitoring category.