Your monitoring dashboard is green. Every link shows up. Tunnels are active. No alerts have fired. This is WAN link failure monitoring doing exactly what it was designed to do.Â
Now picture the BFSI front office it is failing tellers at their stations, core banking application open, customers waiting. The WAN link shows healthy in the console. What the console does not show is that latency has crept from 14ms to 110ms over the past forty minutes, and packet loss is sitting at 6%. Every transaction is timing out. The tellers are staring at frozen screens. The branch manager is on the phone.Â
This is not a misconfiguration. It is not a vendor failure. It is the gap between what monitoring reports and what the network is actually doing and missing the only thing that mattered.Â
The question traditional monitoring answers:Â Is the link up?Â
The question that actually matters: Is the link performing well enough for applications to function?Â
These are not the same question. The gap between them is where networks silently fail.Â
Â
What WAN Link Failure Monitoring Actually Measures (And What It Misses)Â
Most network monitoring platforms whether a basic SNMP Poller, a cloud-based NMS, or a legacy NOC tool, operate on the same fundamental mechanism: a reachability check.Â
A reachability check sends a probe to a destination and waits for a response. ICMP ping is the most common form. If a reply comes back, the link is declared healthy. If no reply comes back, an alert fires.Â
This is a binary test. Up or down. Reachable or unreachable.Â
The problem is that WAN degradation is rarely binary. Links do not go from perfectly healthy to completely dead. They degrade. Gradually. Intermittently. In ways that are invisible to a reachability check but catastrophically visible to applications.Â
- A link with 8% packet loss will pass a reachability check.
- A link with 180ms of jitter will pass a reachability check.
- A link whose latency has crept from 12ms to 95ms over the past two hours will pass a reachability check.
All three of those links are failing your users. None of them will fire an alert in a monitoring system that only measures reachability.Â
This is the core mechanism behind WAN monitoring gaps and it explains why IT teams regularly experience failures that their dashboards never predicted.Â
Â
The Tunnel Status Problem in SD-WAN EnvironmentsÂ
In SD-WAN and VPN environments, monitoring gaps take a specific form: tunnel status reporting.Â
SD-WAN controllers and management consoles expose tunnel status as a primary health indicator. A tunnel is either up or down. If it is up, the path is assumed to be healthy. If it goes down, an alert fires.Â
Tunnel status is a useful signal. It is not a sufficient one.Â
A tunnel can be technically up while the physical path underneath it is experiencing severe degradation. The control plane connection that keeps the tunnel alive requires very little bandwidth and is tolerant of moderate packet loss. The data plane traffic your applications depend on is far less forgiving.Â
Consider a common scenario: an LTE backup link in a branch office. The tunnel is up. The SD-WAN controller shows the link as active and healthy. But the LTE signal has degraded, the tower is congested, the signal quality has dropped, and the path is now experiencing 12% packet loss and 250ms of jitter. VoIP calls are unintelligible. Screen-sharing has frozen. The ERP is timing out on writes operations.Â
Tunnel up. Applications broken. Monitoring silent.Â
This is what a false healthy WAN looks like in practice. It is not a misconfiguration or a monitoring vendor failure. It is the structural limitation of using tunnel status as a proxy for path performance.Â
Â
Why SLA Thresholds Are Either Absent or Set WrongÂ
Proper WAN link failure monitoring should measure path quality against SLA thresholds, defined benchmarks for acceptable latency, jitter, and packet loss on each link.Â
Most traditional monitoring platforms fail here in one of two ways.Â
The first: they do not measure SLA thresholds at all. The monitoring system knows whether a link is up or down. It does not know whether latency on that link is 10ms or 110ms. It has no concept of acceptable versus unacceptable path quality. Every link that is up is healthy, by definition.Â
The second: they set SLA thresholds too loosely to catch real degradation. When thresholds are configured, they are often set conservatively, alerting only on severe degradation rather than on the subtle early stage decline that precedes failure. A threshold that fires at 30% packet loss will not help you when your VoIP calls are already broken at 5%.Â
The consequence in both cases is the same: WAN monitoring gaps that allow link degradation to persist undetected until it becomes a visible outage.Â
In Indian enterprise environments, this problem is compounded by ISP SLA reporting. An ISP may report 99.9% uptime and that figure may be accurate as measured at the ISP’s network boundary. It tells you nothing about path quality at the application layer. A link that meets its contractual SLA can still deliver poor performance to your users if the metrics being tracked are uptime and availability rather than latency, jitter, and packet loss.Â
Â
Where Indian Enterprise Networks Are Especially Exposed
The structural limitations of reachability-based monitoring affect every enterprise network. But the risk is meaningfully higher in Indian enterprise environments for several specific reasons.Â
Many Indian branch offices, particularly in Tier 2 and Tier 3 cities, rely on broadband or LTE as their primary WAN link, not MPLS. Broadband and LTE are inherently more susceptible to intermittent degradation: signal fluctuation, tower congestion, ISP peering issues, last-mile instability. These are grey failures by nature. They are exactly what reachability checks cannot detect and what path quality monitoring is designed to catch. A branch in a metro city and a branch 80km outside it may have radically different connectivity profiles and the semi-urban branch is far more likely to experience the kind of partial, intermittent degradation that basic monitoring misses entirely. When that branch is a manufacturing plant, a warehouse, or a government field office, the operational impact is significant.Â
Government and BFSI deployments face a sharper version of this risk. Regulatory and operational continuity requirements make undetected WAN degradation a compliance issue, not just a user experience problem. An undetected degradation event that disrupts a core banking transaction or a government service delivery system cannot be retroactively explained by pointing to a green dashboard.Â
Large enterprise and government deployments in India commonly mix MPLS, broadband, and LTE across dozens or hundreds of sites. Monitoring each link’s reachability tells you almost nothing about the actual performance profile of that heterogeneous WAN. What you need is genuine WAN visibility continuous, per-path measurement of quality metrics across every link which is exactly what managed SD-WAN delivers regardless of transport type.Â
Â
What Genuine WAN Visibility Actually Looks LikeÂ
The alternative to reachability-based monitoring is continuous path quality measurement across every WAN link: latency, jitter, packet loss, and path stability.Â
Latency and jitter are where most grey failures first appear. Latency creep, the slow increase in round-trip time that precedes many failures, is invisible to a reachability check and immediately obvious to path quality monitoring. Jitter is the variation in latency between successive packets; high jitter destroys real-time applications even when average latency looks acceptable. A link showing 40ms average latency but oscillating between 8ms and 180ms will pass every reachability check and fail every VoIP call.Â
Packet loss compounds both problems. Even at 1 to 2 percent, packet loss has disproportionate effects on TCP-based application performance due to retransmission overhead. At 5 to 10 percent, most applications begin to fail visibly, the same threshold range that traditional monitoring thresholds are typically set to ignore. Path stability, the consistency of these metrics over time – adds a further dimension that aggregate averages cannot capture. A path that oscillates between good and degraded states every few minutes will deliver a poor user experience even if its hourly average looks within range.Â
This level of WAN visibility requires proactive measurement, not reactive alerting. A network that appears healthy to reachability-based tools but is already failing at the application layer is a false healthy WAN and the only way to detect it is to measure what applications actually experience. The monitoring system must continuously probe each path and evaluate the results against meaningful SLA thresholds, not just check whether the tunnel is still up. For a broader look at how this fits into network design, see the SD-WAN overview. For more on how path quality affects application performance, see SD-WAN performance optimization.Â
Â
Why Choose Nirad NetworksÂ
The architectural difference between traditional monitoring and SD-WAN architecture is structural, not incremental. Traditional platforms rely on reachability checks and tunnel status. SD-WAN platforms continuously measure path quality on every active link using synthetic probes, measuring latency, jitter, and packet loss end-to-end, evaluating results against per-policy SLA thresholds in real time, and rerouting traffic automatically when a path degrades below threshold. The link has not failed in the binary sense. But it has been identified as unfit for latency-sensitive traffic and taken out of service before users are affected.Â
Nirad EdgeX implements this architecture as a core function of Nirad’s SD-WAN solutions for enterprise networks, not an add-on.Â
Every Nirad EdgeX deployment monitors each active WAN link for latency, jitter, and packet loss in real time. SLA thresholds are configurable per application policy, a VoIP policy can enforce tighter latency and jitter limits than a bulk data transfer policy. When a path degrades below threshold, traffic is automatically steered to a healthier path, whether that is an MPLS primary, a broadband secondary, or an LTE backup.Â
The centralised management console gives network administrators genuine WAN visibility across every site in the deployment, a live view of path quality, not just link status. In multi-site enterprise and government deployments across India, this means that the intermittent LTE degradation at a Tier 2 branch, the broadband congestion at a remote warehouse, and the MPLS latency spike affecting a BFSI front office are all visible in one place, in real time, before they become user-reported incidents. East Central Railway is one example of this architecture in production: a distributed deployment across multiple sites where centralised path quality visibility replaced the reactive, incident-driven monitoring model that had preceded it.Â
Nirad’s 4G LTE device range integrates directly with the EdgeX platform, enabling consistent path quality monitoring across both fixed and mobile WAN links. For deployments that rely on LTE as a primary or backup transport, which describes a significant portion of Indian enterprise branch networks, this integration closes the monitoring gap at the transport layer where it most commonly occurs. EdgeX is purpose-built for Indian enterprise network conditions, including the cost structure and transport heterogeneity that characterise domestic enterprise deployments.Â
For BFSI and government deployments where undetected degradation carries compliance and operational risk, Nirad Secure extends this architecture with integrated IPS/IDS capabilities, addressing secure branch connectivity alongside the performance monitoring layer.Â
The gap between a green dashboard and a broken application is a structural problem. Reachability checks and tunnel status tell you whether a path exists, not whether it is fit for the applications running over it. Until that distinction is built into the monitoring architecture, the dashboard will be green right up until your users stop working.Â
Ready to close the gap between your monitoring dashboard and your actual network performance? Contact Nirad Networks to see how Nirad EdgeX delivers genuine WAN visibility across every site and every link.Â
Frequently Asked Questions
Clear answers for IT teams dealing with WAN grey failures, link quality issues, and application performance across distributed branches.
01
Why does my monitoring show a WAN link as healthy when users are reporting problems?
Most monitoring platforms use a reachability check, typically an ICMP ping, to determine link health. A reachability check only confirms that the destination is reachable. It does not measure latency, jitter, or packet loss on the path. A link can pass a reachability check while experiencing 10% packet loss and 200ms of jitter, conditions that will break real-time applications and slow transactional systems but will never trigger a reachability-based alert. This is the core cause of the monitoring gap.
02
What is the difference between tunnel status and path quality monitoring?
Tunnel status indicates whether an SD-WAN or VPN tunnel is established between two endpoints. A tunnel can be technically up while the physical path underneath it is severely degraded. Path quality monitoring continuously measures the actual performance of the path, latency, jitter, and packet loss and evaluates it against defined SLA thresholds. Tunnel status tells you the connection exists. Path quality monitoring tells you whether it is performing well enough for applications to function.
03
What SLA thresholds should I set for WAN link monitoring?
The appropriate thresholds depend on the applications running over the link. As a general baseline: latency under 150ms for transactional applications, under 80ms for real-time applications such as VoIP and video; jitter under 30ms for VoIP; packet loss under 1% for real-time applications and under 3% for general enterprise traffic. The key principle is that thresholds should be set based on application requirements, not set loosely enough to avoid false alarms. A threshold that never fires is not conservative, it is broken.
04
Why are Indian enterprise networks more exposed to WAN monitoring gaps than networks in other markets?
Two factors compound the problem in India. First, many Indian branch deployments, particularly outside metro cities, rely on broadband and LTE as primary WAN transports rather than MPLS. These transports are inherently more susceptible to intermittent degradation: tower congestion, signal fluctuation, last-mile instability. These are grey failures that reachability checks cannot detect. Second, ISP SLA reporting in India typically measures contractual uptime at the network boundary, not application-layer performance. A link can meet its 99.9% uptime SLA while consistently delivering poor path quality to end users.
05
How does SD-WAN improve WAN link failure monitoring compared to traditional approaches?
SD-WAN platforms continuously probe every active WAN path using synthetic traffic, measuring latency, jitter, and packet loss at regular intervals. These measurements are evaluated against per-policy SLA thresholds in real time. When a path degrades below threshold, even if the tunnel remains up, the SD-WAN controller can automatically reroute traffic to a healthier path before users are affected. This converts WAN monitoring from a reactive, alert-based discipline into a proactive, policy-driven one. The centralised management console provides genuine WAN visibility across all sites and all link types, replacing the binary link-up/link-down view that traditional monitoring provides.


