← Back to Blog
isp monitoring maintenance

A Network Health Checklist for a Small ISP: Daily, Weekly and Monthly

A small ISP with a hundred or so subscribers usually has no network operations team. The person who built the network is also the one who answers the phone, and monitoring means noticing when customers complain. That works until the day two things break at once. What helps is not a bigger tool but a routine: a short list of things to look at at regular intervals, so problems show up as a trend before they show up as an outage.

1

Every day: is everything up and is anything unusual?

The daily check should take five minutes. Confirm every core router, access point and backhaul link is reachable, and look at the log for anything critical:

/log print where topics~"critical|error"
/system resource print

In /system resource print, look at uptime (an unexpected low number means an unplanned reboot), CPU load and free memory. A CPU that sits above 70-80% outside peak hours, or memory that shrinks every day, is worth investigating before it becomes an outage.

2

Every day: watch the upstream, not just the router

A router that responds does not mean the internet works. Keep a latency measurement running to something beyond your own network, such as your upstream gateway and a public address. Look at packet loss and response time together: a link that answers but drops 3% of packets already feels broken to subscribers, even though a simple "is it up" check says everything is fine.

3

Every week: errors, drops and traffic trends

Interface counters reveal problems that never trigger an alarm, such as a bad cable, a failing SFP or a duplex mismatch:

/interface print stats
/interface ethernet monitor ether1 once

Any interface where rx-error or tx-drop keeps growing needs attention. Compare peak traffic with the capacity of each link: when a backhaul regularly passes 70-80% of its capacity at peak, subscribers will start to feel congestion before you see a saturated line, and the upgrade conversation should start now.

4

Every week: sessions and power

If you run PPPoE, the number of active sessions is a quick health indicator. A sudden drop in a concentrator that nobody touched is a sign of a failed uplink or an access problem:

/ppp active print count-only
/system health print

The second command shows voltage and temperature on hardware that supports it. Equipment on a pole or in an unventilated cabinet often shows its trouble in the numbers days before it fails.

5

Every month: updates, backups and housekeeping

Once a month, do the maintenance that nobody remembers on a busy day. Check for new RouterOS versions and read the release notes before upgrading anything critical:

/system package update check-for-updates
/system routerboard print

Confirm that backups exist, that they are recent and that they are stored away from the router itself. Review user accounts for people who no longer work with you, and check that the firewall still blocks management access from the WAN.

6

Write down your baseline so you can recognize abnormal

None of the numbers above mean anything without a reference. Note what normal looks like for each key router: typical CPU, typical peak traffic, typical latency to the upstream, usual number of PPPoE sessions. A spreadsheet is enough. When something changes, you will know within seconds whether it is a real problem or just Friday evening.

Why do it this way

Most outages in a small network are preceded by a symptom that was visible days or weeks earlier: a CPU creeping up, a link filling at peak, an interface collecting errors, a router rebooting at night. A routine turns those into small tasks instead of emergencies. It also protects you from the opposite problem, depending on a single person's memory. Writing the checklist down means a colleague can cover for you on vacation, and that your network does not depend on what you happened to remember to look at.

How MoniTik helps

Most of this checklist is exactly what MoniTik automates, so you are not doing it by hand on every router. It shows device status and email alerts when a router goes down (and when it recovers), CPU and memory history, latency and interface traffic graphs per device, and a data cap meter for links like Starlink. A summary dashboard lists which devices are currently down. The routine remains yours, but the daily rounds become a glance at one screen instead of logging in to each router.

Start Free Trial
Santiago Rivas
Santiago Rivas Field Technician

Santiago spends his days on rooftops and in racks, keeping MikroTik links online for local ISPs.