Targeted Probing for Silent Failure Pinpointing in Large Networks

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Identifying silent failures in large-scale computer networks is challenging due to their unannounced nature and the resource-intensive process of active monitoring, which can be demanding in terms of time and resources.

Innovation Solution

A system and method that utilize a network monitoring agent on each host node to maintain route-data in a database, identifying a subject intermediary node for investigation by selecting target probe paths and testing them with targeted probes to determine the operational status of intermediary nodes.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If active monitoring is performed to identify silent failures, then fault detection capability is improved, but time and resource consumption increases

Engineering Contradiction:
Improvefault detection capabilityVSAvoidtime and resource consumption
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent segments the network into multiple zones and divides monitoring responsibilities among zone controllers and probing controllers. This segmentation allows distributed fault detection without requiring centralized monitoring of the entire network, reducing time and resource consumption while maintaining comprehensive fault detection capability.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent performs preliminary actions by pre-configuring route data in databases, pre-identifying potential faulty intermediary nodes, and pre-establishing probe paths. This preliminary preparation enables rapid fault detection when failures occur, improving response time without requiring extensive real-time resources.

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If comprehensive network monitoring is implemented, then fault detection accuracy is improved, but system complexity increases

Engineering Contradiction:
Improvefault detection accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent introduces intermediary components including zone controllers, probing controllers, and network monitoring agents that mediate between host nodes and the fault detection system. These intermediaries simplify the overall system architecture by distributing functions and providing standardized interfaces, reducing system complexity while maintaining accurate fault detection through coordinated monitoring.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Measurement precision

If targeted probing is performed on multiple probe paths, then fault identification precision is improved, but resource demands increase

Engineering Contradiction:
Improvefault identification precisionVSAvoidresource demands
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The patent applies local quality by focusing probing resources on specific target intermediary nodes identified as potential fault sources, rather than uniformly monitoring all network nodes. The system selectively directs probe paths through suspected faulty nodes based on route data analysis, improving fault identification precision while concentrating resources on high-probability areas.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent employs partial action by selecting a subset of probe paths for targeted probing rather than exhaustively testing all possible paths. The system identifies and probes only the most relevant paths through suspected faulty intermediary nodes, achieving sufficient fault identification precision without the resource expenditure required for complete path coverage.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS9712381B1Systems and methods for targeted probing to pinpoint failures in large scale networks
Publication Date: 2017.07.18 GOOGLE LLC
  • US9712381B1 patent drawing
  • US9712381B1 patent drawing
  • US9712381B1 patent drawing

AI summary

Systems and methods for locating network errors. The system includes a plurality of host nodes in a network of host nodes and intermediary nodes, and a database storing route data for each of a plurality of host node pairs. The system includes a controller configured to identify a subject intermediary node to investigate for network errors and select, using route data stored in the database, a set of target probe paths. Each target probe path includes a respective pair of host nodes separated by a network path including at least one target intermediary node, which is either the subject intermediary node or an intermediary node that is a next-hop neighbor of the subject intermediary node. The controller is configured to test each target probe path in the set of target probe paths and to determine, based on a result of the testing, an operational status of the subject intermediary node.