Multi-vendor network monitoring and fault management

HyperNMS

HyperNMS is a carrier-grade network management system that discovers, monitors and maps every device in a mixed-vendor network — routers, switches, OLTs, servers, radios and power plant — and turns raw counters into alarms an operations team can act on.

Who it's for

ISPs, telcos, data centre operators and large enterprises running mixed-vendor infrastructure that no single vendor NMS can see all of.

SNMP v1/v2c/v3Streaming telemetryNetFlow / sFlowSyslog and traps
100k+
monitored objects per cluster
30s
default polling granularity
Any
vendor, via open protocols

Every vendor ships a console. None of them shows your network.

A typical operator runs optical line terminals from one vendor, core routing from another, switching from a third, plus servers, UPS, DC plant and radio. Each has its own element manager, its own alarm list and its own idea of what "critical" means. The network exists only in the heads of the people on shift.

One monitoring plane for equipment from every vendor in the rack.

What it costs you today

  • Alarms arrive in four consoles with no shared severity
  • A single fibre cut raises hundreds of unrelated-looking alarms
  • Capacity questions need a spreadsheet and a week
  • Nobody can say what changed on a device last Tuesday

How HyperNMS works

The decisions that shape the platform, and why they were made that way.

01

Discover, then keep discovering

Seed a subnet and the platform walks it — SNMP, LLDP, CDP, ARP and routing tables — to build an inventory of devices, interfaces, VLANs, links and neighbours. It re-walks continuously, so an interface added at 3am is monitored by 3:05.

02

Poll everything, agent optional

Network equipment is polled agentlessly over standard protocols. Where you own the host — servers, probes, virtual machines — an optional agent adds process, filesystem, service and log-level detail the SNMP MIB never exposes.

03

Alarms that understand topology

Because the platform knows the link graph, an upstream failure suppresses the hundreds of downstream alarms it caused and raises one root-cause event instead. The shift sees a fibre cut, not a wall of red.

04

Scale out, not up

Distributed pollers sit inside remote POPs, customer sites and isolated management networks, buffering locally through a WAN outage and forwarding to the central cluster when it returns.

What you get

  • Automatic discovery of devices, interfaces and dependencies
  • Agentless polling over SNMP, ICMP, SSH and vendor REST APIs
  • Optional agents for deep host, service and log metrics
  • Topology-aware alarm correlation that suppresses downstream noise
  • Flow analysis for traffic mix, top talkers and peering decisions
  • Configuration backup with scheduled diffs and change alerting
  • Threshold, baseline and trend-based alerting with escalation
  • Distributed pollers for remote sites and isolated management VRFs
  • Long-horizon capacity planning and forecast reporting
  • Multi-tenant dashboards with role-based access control

Capabilities in full

Everything the platform does, grouped by the team that uses it.

Collection

  • SNMP v1, v2c and v3 with bulk operations
  • ICMP reachability, latency, jitter and loss
  • SSH and CLI scraping for devices with thin MIBs
  • Vendor REST and gNMI streaming telemetry
  • NetFlow, sFlow, IPFIX and J-Flow collection
  • Syslog ingestion and SNMP trap handling
  • Synthetic checks: HTTP, DNS, TCP, TLS expiry
  • Optional host agent for OS and service metrics

Analysis and alerting

  • Static thresholds and dynamic baselines
  • Topology-aware root-cause correlation
  • Flapping suppression and maintenance windows
  • Escalation chains with on-call rotation
  • Notification by email, SMS, webhook and chat
  • Alarm acknowledgement, notes and audit trail
  • Ticketing integration in both directions
  • Per-tenant alarm scoping and severity policy

Visualisation and reporting

  • Auto-generated Layer 2 and Layer 3 topology maps
  • Geographic maps for POPs, towers and remote sites
  • Composable dashboards per team and per tenant
  • Interface utilisation and error-rate reporting
  • Availability and SLA reporting with export
  • Capacity forecasting on 6 and 12 month horizons
  • Scheduled PDF and CSV report delivery
  • Open API and data export for your own BI tools

Operations

  • Scheduled configuration backup and version history
  • Configuration diff alerting on unexpected change
  • Device inventory with lifecycle and warranty dates
  • Template-driven onboarding for new device classes
  • Role-based access control and tenant isolation
  • High availability with active and standby collectors
  • Retention tiering: raw, hourly and daily rollups
  • Full audit log of every operator action

Where it gets deployed

ISP core and access

Watch OLTs, BNGs, aggregation switches and upstream transit in one place, with per-POP dashboards for the field teams.

Data centre and hosting

Rack-level power, cooling, switching and host metrics correlated so a PDU fault does not read as forty server outages.

Enterprise WAN and campus

Branch reachability, circuit performance and per-application flow visibility across sites you do not physically visit.

Towers and remote plant

Distributed pollers monitor unmanned sites over cellular backhaul and buffer through the outages that come with them.

Common questions

Does it work with equipment from any vendor?

Yes. The platform collects over open protocols — SNMP, ICMP, SSH, syslog, flow and streaming telemetry — rather than vendor-specific SDKs, so a device that speaks any of those can be monitored. Device templates ship for common platforms and you can write your own.

Do we have to install agents on network devices?

No. All network equipment is monitored agentlessly. Agents are optional and only relevant for servers and virtual machines, where they add process, filesystem and log detail that SNMP does not expose.

How does it handle sites with unreliable backhaul?

Distributed pollers run at the remote site, collect locally and buffer to disk when the link to the central cluster drops. When connectivity returns they forward the backlog, so you get the history rather than a gap.

Can it live alongside our existing vendor NMS?

It usually does. Most operators keep the vendor element manager for deep device configuration and use HyperNMS as the cross-vendor fault and performance layer above it, with alarms forwarded on to the existing ticketing system.

Put HyperNMS in front of your own network

Bring your CPE models, your protocol mix and your integration constraints. A solutions engineer will show you what the platform does with them — no slides.