Bayesian Network Fault Location in Telecommunications
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current fault management systems in telecommunications networks face challenges in efficiently identifying the root cause of faults due to complexity in alarm correlation and the need for extensive knowledge of network apparatuses and their relationships, particularly in large-scale networks with multiple transmission protocol levels.
Innovation Solution
The method involves using Bayesian Networks at the apparatus level to build probabilistic models based on received alarms and conditioned probabilities, allowing for real-time fault location by aggregating partial networks constructed by autonomous agents to form a complete Bayesian network for a limited network region, thereby simplifying the learning and maintenance processes.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If traditional alarm correlation methods are used to identify root causes in large-scale networks, then comprehensive fault detection is achieved, but system complexity and difficulty of automated implementation increase significantly
Solution Approach 1:
The patent divides the large-scale network into multiple domains, each managed by an autonomous agent that constructs its own Bayesian network. This segmentation allows each agent to handle only local alarm correlation independently, reducing the overall system complexity while maintaining comprehensive fault detection capability across the entire network.
Solution Approach 2:
The patent introduces a hierarchical dimension by organizing Bayesian networks at different levels: local networks at the apparatus level, domain-level networks aggregating multiple apparatuses, and network-level networks combining multiple domains. This multi-dimensional structure enables automated fault identification without requiring centralized processing of all alarms simultaneously.
2Measurement precision
If comprehensive alarm aggregation and correlation is implemented across the entire network, then accurate root cause identification is achieved, but information processing time and computational resources increase
Solution Approach 1:
The patent segments the alarm processing task by creating independent Bayesian networks for each apparatus and domain. Each agent processes alarms locally without waiting for centralized aggregation, enabling parallel processing that reduces fault identification time while maintaining accuracy through the probabilistic reasoning capability of Bayesian networks.
Solution Approach 2:
The patent implements preliminary construction of Bayesian networks during normal operation, continuously learning and updating probabilistic relationships between alarms and faults. When a fault occurs, the pre-built models enable immediate inference without requiring time-consuming data collection and analysis from scratch.
3Measurement precision
If detailed knowledge of network apparatuses and topological relationships is maintained centrally, then accurate fault diagnosis is achieved, but maintenance and updating of information becomes extremely complex
Solution Approach 1:
The patent enables each autonomous agent to maintain its own Bayesian network and local knowledge base independently. Each agent automatically updates its model based on local observations and alarm patterns without requiring centralized coordination, making the system self-maintaining and adaptable to network changes without complex manual updates.
Solution Approach 2:
The patent implements dynamic Bayesian networks that automatically adapt to network changes by continuously learning from incoming alarms and fault patterns. The probabilistic models are updated in real-time to reflect changing network conditions, topology modifications, and new fault types, maintaining diagnostic accuracy without manual intervention.
Data Source
Figure 1~2
Figure 3~4c
Figure 5a~5c
AI summary
Disclosed herein is a method for locating a fault in a communication network, comprising receiving status information relating to alarms, events, polled statuses or test results in the communication network; and locating the fault based on the received status information, wherein locating the fault includes identifying a limited region of the communication network in which the fault has occurred based on the received status information and on topological and functional information relating to network apparatuses that have generated the status information; constructing a probabilistic model relating faults and status information in the identified limited region of the communication network; and locating the fault based on the constructed probabilistic model and on status information received from the identified limited region of the communication network.