A communication fault point sniffing and PHM management system based on communication network topology

CN122802352APending Publication Date: 2026-09-22ZHONGRUNJIAN COMMUNICATION TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202611193741.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-07
Publication Date
2026-09-22

AI Technical Summary

Technical Problem

[0005]本发明的目的在于提供一种基于通信网络拓扑的通信故障点嗅探及PHM管理系统,以解决上述背景技术中提出的多节点关联异常场景下,根因故障节点与故障传导关系难以精准识别和同一故障源对应的关联异常节点范围难以精准划定的问题

Benefits of technology

1.本发明构建故障传导梯度场,融合节点异常程度差值、异常发生时序差值、链路基础权重与场景约束系数量化单链路传导强度,经时序一致性与拓扑连通性双重校验修正传导路径,可精准区分根因故障节点与次生异常节点,明确多异常节点间的因果关联与传导次序,提升多节点关联异常场景下故障根因定位的精准度;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122802352A_ABST
    Figure CN122802352A_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of communication network fault location, in particular to a communication fault point sniffing and PHM management system based on communication network topology, comprising: a communication network topology sensing and data acquisition unit; a topology correlation fault point sniffing unit, which adopts a fault conduction gradient field construction method, calculates and processes node abnormal timing, link performance attenuation and topology attributes, and generates a fault conduction gradient field; a network equipment PHM state evaluation unit; a control instruction generation and interaction unit. The present application constructs a fault conduction gradient field, fuses node abnormal degree difference, timing difference, link weight, scene system quantitative conduction strength, and is checked by time sequence and topology, can distinguish root cause and secondary nodes, and clearly shows cause and effect correlation and conduction order; along the gradient attenuation path, the secondary nodes are matched and the connected influence domain is delimited, which can distinguish homologous correlation abnormality and independent abnormality event, and provides range support for fault operation and maintenance scheduling.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of communication network fault location technology, and more specifically, to a communication fault point detection and PHM management system based on communication network topology. Background Technology

[0002] As communication networks continue to expand, the number of network devices, connection layers, and service carrying relationships are becoming increasingly complex. The volume of equipment status, link performance, and fault alarm data generated by network operation is constantly growing, and fault propagation paths are becoming more concealed and their impact wider. Therefore, conducting operational status monitoring, precise fault location, comprehensive health assessment, and efficient operation and maintenance scheduling for communication equipment and networks as a whole has become a core supporting technology for ensuring the stable and reliable operation of communication networks, and a key research direction in the field of communication network operation and maintenance.

[0003] Chinese patent CN202210268238.7 discloses a cascading failure model and node vulnerability assessment method for a power-communication converged network. This method generates power layer and communication layer models for the power-communication converged network, constructs load-capacity models and load redistribution models for each layer, sets inter-network failure probabilities, and builds a cascading failure model for the power-communication converged network based on these models. Furthermore, it constructs node vulnerability assessment indicators considering cascading failures in the power-communication converged network, and ranks these indicators to obtain a set of key protected nodes. Chinese patent CN202310076514.4 discloses a health monitoring system and method for communication equipment based on PHM technology. This method collects operating performance data and obtains PHM modeling data through a data acquisition module, sets algorithm indicators and constructs an initial model through a feature extraction module, analyzes and processes the initial model through a data analysis module to output a target model, and then analyzes the target model through a health management module to determine the cause of communication equipment failure.

[0004] Although the aforementioned technical solutions study the operational status of communication networks from two dimensions—network cascading failure analysis and communication equipment health management—the following technical problems still urgently need to be addressed: First, in scenarios involving multiple nodes with associated anomalies, it is difficult to accurately identify the root cause fault node and the fault propagation relationship. Chinese patent CN202210268238.7 mainly screens key protected nodes based on cascading failure models and node vulnerability indicators, while Chinese patent CN202310076514.4 mainly relies on single-device operating data to construct a health model and locate equipment faults. Neither of these solutions combines network topology correlation and the timing of anomaly occurrence to quantitatively analyze the propagation strength and causal relationship between multiple anomalous nodes. When multiple nodes in the network experience anomalies sequentially, relying solely on node vulnerability ranking or single-device health analysis cannot distinguish between the root cause fault node and secondary anomalous nodes affected by propagation, nor can it quantify the strength of fault propagation between nodes, making it difficult to establish accurate fault correlation logic. Second, it is difficult to accurately define the range of associated anomalous nodes corresponding to the same fault source. Chinese patent CN202210268238.7 analyzes the inter-node relationships and completes vulnerability ranking only from the perspective of cascading impact, while Chinese patent CN202310076514.4 only performs health monitoring and fault diagnosis for individual devices. Neither of them combines topological connectivity constraints and anomaly temporal correlation to delineate the boundary of the associated abnormal nodes covered by a single fault source. When multiple abnormal events or multiple fault sources exist simultaneously in the network, it is impossible to effectively distinguish between related anomalies from the same source and independent abnormal events, which easily leads to fuzzy boundaries in determining the fault impact range, thereby interfering with subsequent fault location analysis and accurate scheduling of operation and maintenance resources. In view of this, we propose a communication fault point detection and PHM management system based on communication network topology. Summary of the Invention

[0005] The purpose of this invention is to provide a communication fault point detection and PHM management system based on communication network topology, so as to solve the problems mentioned in the background art, such as the difficulty in accurately identifying the root cause fault node and the fault propagation relationship, and the difficulty in accurately defining the range of associated abnormal nodes corresponding to the same fault source in multi-node associated anomaly scenarios.

[0006] To address the aforementioned technical problems, the present invention aims to provide a communication fault point detection and PHM management system based on communication network topology, comprising: A communication network topology sensing and data acquisition unit is used to acquire node link topology information, network hierarchical attributes and network-wide device operation status data of the communication network, and to standardize and integrate the topology information, hierarchical attributes and operation status data to obtain a structured topology dataset and a device operation dataset. The topology-related fault point detection unit receives a structured topology dataset and a device operation dataset. It uses a fault propagation gradient field construction method to calculate and process node anomaly timing, link performance degradation, and topology attributes to generate a fault propagation gradient field. Based on the fault propagation gradient field, it extracts root cause fault nodes, secondary abnormal nodes, and fault impact domain topology subsets. After time series and topology connectivity verification, it generates fault location results with propagation attributes. The network device PHM status assessment unit receives fault location results with conduction attributes and equipment operation datasets. It uses equipment health status modeling and trend inference to perform hierarchical health assessment and degradation trend analysis on root cause fault nodes, secondary abnormal nodes and nodes in the fault impact domain, and generates hierarchical PHM status assessment results and operation and maintenance priority ranking results. The control command generation and interaction unit receives the hierarchical PHM status assessment results and operation and maintenance priority ranking results, and uses network control command encapsulation and scheduling conversion methods to perform command-based processing on fault isolation strategies, traffic scheduling schemes and operation and maintenance tasks, generating device control commands and operation and maintenance scheduling commands, and outputting them to the corresponding network devices and operation and maintenance management terminals.

[0007] As a further improvement to this technical solution, the communication network topology sensing and data acquisition unit includes a topology and hierarchical attribute acquisition module, a device operating status acquisition module, and a standardized integration and processing module, wherein: The topology and hierarchical attribute acquisition module obtains node link topology information and network hierarchical attributes of the communication network through distributed network detection, and outputs the original topology and hierarchical data to the standardized integration and processing module. The device operation status acquisition module acquires the operation status data of all network devices through a multi-source performance acquisition protocol and outputs the raw operation status data to the standardized integration and processing module. The standardized integration processing module receives the original topology and hierarchical data and the original operating status data. After standardization and integration processing, including format unification, time sequence alignment and data cleaning, it generates a structured topology dataset and a device operating dataset.

[0008] As a further improvement to this technical solution, the topology-related fault point detection unit includes a fault propagation gradient field construction module, a fault attribute extraction module, and a time series and topology connectivity verification module, wherein: The fault propagation gradient field construction module receives the structured topology dataset and the device operation dataset, generates a fault propagation gradient field, and outputs it to the fault attribute extraction module. The fault attribute extraction module extracts the root cause fault nodes, secondary abnormal nodes and fault influence domain topology subsets based on the fault propagation gradient field, and outputs the extraction results to the time series and topology connectivity verification module. The timing and topology connectivity verification module performs timing and topology connectivity verification on the extracted results and generates fault location results with conduction attributes.

[0009] As a further improvement to this technical solution, the process of generating the fault propagation gradient field in the fault propagation gradient field construction module includes the following steps: S21.1 Receive the structured topology dataset and the device operation dataset, extract the anomaly degree, anomaly occurrence time and link basic weight of each node, and complete data normalization and time sequence alignment. S21.2 Calculate the difference in anomaly severity and the difference in anomaly occurrence time for each link, and calculate the propagation gradient value of a single link by combining the link's basic weight and the scenario constraint correction coefficient; the propagation gradient value of a single link is calculated using the following formula: ; In the formula, For link The propagation gradient value; upstream node The anomaly level value is taken from the device operation dataset; For downstream nodes The anomaly level value is taken from the device operation dataset; This is the time difference between the occurrence of anomalies in the downstream node and the upstream node, taken from the time stamp of the device operation dataset; For link The base weights are taken from the structured topology dataset; The scene constraint correction coefficient is used; when the time difference between the occurrence of anomalies in upstream and downstream nodes is zero, the propagation gradient value is the product of the link base weight and the scene constraint correction coefficient. S21.3. Summarize the propagation gradient values ​​of all links in the entire domain to form a fault propagation gradient field covering the entire network.

[0010] As a further improvement to this technical solution, the scene constraint correction coefficient is determined in step S21.2. The process includes the following steps: S21.21. Determine the link direction based on the network hierarchy attribute of the upstream and downstream nodes. A link from a higher network hierarchy node to a lower network hierarchy node is a downlink, and vice versa. Set the conduction attenuation coefficient according to the link direction. The conduction attenuation coefficient of the downlink is less than that of the uplink. S21.22. Adjust the conduction coefficient according to the relative relationship between the real-time load of the link and the preset load inflection point. When the real-time load of the link exceeds the preset load inflection point, the conduction coefficient increases by a step. S21.23. Multiply the conduction attenuation coefficient by the conduction coefficient to generate the corresponding scenario constraint correction coefficient for the link. .

[0011] As a further improvement to this technical solution, the process of extracting fault attributes by the fault attribute extraction module includes the following steps: S22.1 Calculate the node gradient value corresponding to each node, and take the maximum value of the propagation gradient value of all downstream links of the node as the node gradient value of the corresponding node. S22.2 Identify the extreme points in the gradient values ​​of all nodes and mark the corresponding nodes as root cause fault nodes; S22.3 Matching is performed along the downstream path with continuously decaying gradient, and abnormal nodes on the path are marked as secondary abnormal nodes. S22.4 Delineate the connected regions continuously covered by gradient values ​​and generate a topological subset of the fault-affected domain.

[0012] As a further improvement to this technical solution, the timing and topology connectivity verification module includes a timing verification submodule and a topology connectivity verification submodule, wherein: The timing verification submodule performs timing verification on the extracted results, and determines that the abnormal occurrence time of the root cause fault node is earlier than the abnormal occurrence time of all nodes with abnormal occurrence time in the fault influence domain topology subset. If the verification fails, the corresponding downstream path is removed. The topology connectivity verification submodule performs topology connectivity verification on the extraction results, determines that the topology subset of the fault influence domain is a connected region starting from the root cause fault node, and cuts off the corresponding non-continuous boundary if the verification fails. The timing and topology connectivity verification module generates fault location results with conduction attributes based on the extracted results after verification correction.

[0013] As a further improvement to this technical solution, the network device PHM status assessment unit includes a hierarchical health assessment module and an operation and maintenance priority ranking module, wherein: The graded health assessment module receives fault location results with transmission attributes and equipment operation datasets, performs graded health assessment and degradation trend analysis on root cause fault nodes, secondary abnormal nodes and nodes within the fault impact domain, generates graded PHM status assessment results and outputs them to the operation and maintenance priority ranking module. The maintenance priority ranking module receives the hierarchical PHM status assessment results, sorts them in conjunction with the fault impact domain range, and generates maintenance priority ranking results.

[0014] As a further improvement to this technical solution, the process of graded health assessment and degradation trend analysis in the graded health assessment module includes the following steps: S31.1 Perform a deep degradation assessment on the root cause failure node to generate data on the degree and development trend of root cause failure degradation; S31.2 Perform conduction damage assessment on secondary abnormal nodes and generate data on the degree of conduction damage and recovery trend of secondary nodes; S31.3 Perform a preliminary risk warning assessment on normal nodes within the fault impact domain and generate risk level and warning data; S31.4 Summarize the evaluation data of various nodes and generate hierarchical PHM status evaluation results.

[0015] As a further improvement to this technical solution, the control instruction generation and interaction unit includes a hierarchical instruction generation module and an instruction distribution and interaction module, wherein: The hierarchical instruction generation module receives the hierarchical PHM status assessment results and the operation and maintenance priority ranking results, performs instruction processing on the fault isolation strategy, traffic scheduling scheme and operation and maintenance tasks, generates equipment control instructions and operation and maintenance scheduling instructions and outputs them to the instruction distribution and interaction module. The instruction distribution and interaction module receives device control instructions and operation and maintenance scheduling instructions, and outputs the corresponding instructions to the corresponding network devices and operation and maintenance management terminals.

[0016] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. This invention constructs a fault propagation gradient field, integrates the difference in the degree of node anomaly, the difference in the timing of anomaly occurrence, the basic weight of the link, and the scenario constraint coefficient to quantify the propagation strength of a single link. After double verification of timing consistency and topological connectivity, the propagation path is corrected, which can accurately distinguish between the root cause fault node and the secondary abnormal node, clarify the causal relationship and propagation order between multiple abnormal nodes, and improve the accuracy of fault root cause location in multi-node associated abnormal scenarios. 2. This invention matches secondary abnormal nodes along the downstream path of continuously decaying gradient, delineates the connected regions continuously covered by gradient values ​​to generate a topological subset of the fault influence domain, and corrects the region boundary through topological connectivity verification. This can accurately delineate the range of associated abnormal nodes corresponding to a single fault source, effectively distinguishing between related abnormalities from the same source and independent abnormal events, and providing clear and reliable range support for fault operation and maintenance scheduling. Attached Figure Description

[0017] Figure 1 This is a schematic diagram of the overall system framework structure of the present invention; The meanings of the labels in the diagram are as follows: 1. Communication network topology sensing and data acquisition unit; 11. Topology and hierarchical attribute acquisition module; 12. Equipment operation status acquisition module; 13. Standardized integration and processing module; 2. Topology-related fault point detection unit; 21. Fault propagation gradient field construction module; 22. Fault attribute extraction module; 23. Temporal and topology connectivity verification module; 3. Network device PHM status assessment unit; 31. Hierarchical health assessment module; 32. Operation and maintenance priority ranking module; 4. Control instruction generation and interaction unit; 41. Hierarchical instruction generation module; 42. Instruction distribution and interaction module. Detailed Implementation

[0018] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.

[0019] like Figure 1 This embodiment provides a communication fault point detection and PHM management system based on communication network topology, suitable for operation and maintenance scenarios involving fault root cause localization and equipment health trend management in complex communication networks. The system can be deployed on a network operation and maintenance management platform, relying on the collection and integration of network-wide topology, hierarchical attributes, and equipment operation data to form a structured dataset. Its core mechanism is to construct a fault propagation gradient field, calculate the difference in anomaly degree, temporal difference, and topology weight between nodes, identify root cause fault nodes, secondary abnormal nodes, and their influence domains, and based on this, perform hierarchical health assessments and degradation trend analyses on different nodes. The system includes: The communication network topology sensing and data acquisition unit 1 is used to acquire node link topology information, network hierarchical attributes, and network-wide device operating status data of the communication network. It performs standardized integration processing on the topology information, hierarchical attributes, and operating status data to obtain a structured topology dataset and a device operating dataset. The communication network topology sensing and data acquisition unit 1 includes a topology and hierarchical attribute acquisition module 11, a device operating status acquisition module 12, and a standardized integration processing module 13, wherein: The topology and hierarchical attribute acquisition module 11 acquires node link topology information and network hierarchical attributes of the communication network through distributed network probing, and outputs the raw topology and hierarchical data to the standardization integration processing module 13. Distributed network probing can read link layer discovery information in the device management information database based on Simple Network Management Protocol (SNMP), obtain neighbor relationships based on Link Layer Discovery Protocol (LLDP), and parse network connection relationships through routing protocol information, thereby extracting the device identifier, interface identifier, and adjacency relationship of each node to form the raw topology data. At the same time, the topology and hierarchical attribute acquisition module 11 also obtains the network hierarchical attributes corresponding to the nodes from the network management system configuration database or the device's preset hierarchical identifiers. For example, it marks the nodes as core layer, aggregation layer, or access layer, or as specific management domain identifiers. The acquired hierarchical attributes, indexed by the node identifier, together with the raw topology data, constitute the raw topology and hierarchical data, which is then output to the standardization integration processing module 13.

[0020] The device operation status acquisition module 12 acquires network-wide device operation status data through a multi-source performance acquisition protocol and outputs the raw operation status data to the standardized integration and processing module 13. The multi-source performance acquisition protocol may include performance indicators such as interface traffic, packet loss rate, CPU utilization, and memory utilization obtained through SNMP polling, as well as high-speed data streams received in real-time via telemetry subscription, and device operation events obtained through methods such as system logs (Syslog). All types of operation status data carry timestamps and are recorded in time-series format to form raw operation status data, which is then output to the standardized integration and processing module 13.

[0021] The standardized integration processing module 13 receives the original topology and hierarchical data and the original operating status data. After standardization and integration processing, including format unification, time sequence alignment, and data cleaning, it generates a structured topology dataset and a device operating dataset. Specifically: Regarding format unification, the standardization and integration processing module 13 converts data from different sources and with different representations into a unified data model. For example, the standardization and integration processing module 13 maps device identifiers in topology information to unique node IDs across the entire network, and converts performance indicators from different acquisition protocols into unified numerical types and units, making the same type of indicators from different devices and different interfaces comparable.

[0022] Regarding time alignment, the standardization and integration processing module 13 uses the Network Time Protocol (NTP) time base to align the timestamps of the operational status data from different acquisition cycles. For data points with inconsistent sampling times, the standardization and integration processing module 13 uses linear interpolation or neighbor-value filling to form an equally spaced time series. For records whose timestamp deviations exceed the preset tolerance range due to acquisition delays, the standardization and integration processing module 13 marks them separately before participating in the alignment.

[0023] In terms of data cleaning, the standardization and integration processing module 13 identifies and removes obvious outliers in the original data. The rules for determining outliers can be based on the physical range of the indicator, the statistical distribution threshold, or the jump amplitude with adjacent sampling points. For the removed and originally missing data points, the standardization and integration processing module 13 performs interpolation to complete them based on the temporal relationship between the valid data points before and after them; and performs deduplication processing on duplicate records.

[0024] After the above standardization and integration process, the standardization and integration module 13 generates a structured topology dataset and a device operation dataset. The structured topology dataset includes a node table, a link table, and a node hierarchy attribute table. The node table records the node ID and device type, while the link table records the node IDs at both ends, interface information, and basic link weights. The basic link weights can be pre-configured based on inherent topology attributes such as link bandwidth and service carrying importance level, with a value range of (0,1]. Links with higher bandwidth and more important services carry have larger corresponding weight values. The node hierarchy attribute table records the hierarchy identifier corresponding to the node. The device operation dataset contains multi-dimensional time-series performance data indexed by node ID, recording indicators such as CPU utilization, memory utilization, interface throughput, packet loss rate, and device status of each node at different times in time sequence. The generated structured topology dataset and device operation dataset are output to the topology-related fault point sniffing unit 2, providing a data source for the construction of the fault propagation gradient field.

[0025] Topology-related fault detection unit 2 receives a structured topology dataset and a device operation dataset. It employs a fault propagation gradient field construction method to calculate and process node anomaly timing, link performance degradation, and topology attributes to generate a fault propagation gradient field. Based on this gradient field, it extracts the root cause fault node, secondary anomaly nodes, and a subset of the fault-affected topology. After timing and topology connectivity verification, it generates fault location results with propagation attributes. Specifically, traditional fault location methods typically rely on alarm thresholds for single nodes or simple alarm association rules. When facing cascading faults, they are prone to misjudging secondary alarms as the root cause and cannot quantify the intensity and range of fault propagation along the topology. Unlike these methods, this unit introduces a fault propagation gradient field construction and utilization mechanism, unifying the anomaly degree, anomaly timing, and link attributes of each node in the network topology into the quantitative calculation framework of the gradient field. This allows the propagation relationship of the fault in the topology structure to be characterized by scalarized propagation gradient values. Based on this, the source, propagation path and impact range of the fault are extracted from the gradient field by extreme point identification and gradient continuous decay path tracing. Then, the causal consistency of the location results is ensured by dual verification of temporal sequence and topological connectivity. Finally, the fault location result includes not only the fault node identifier, but also attribute information such as the propagation path and propagation intensity.

[0026] In this embodiment, the topology-related fault detection unit 2 includes a fault propagation gradient field construction module 21, a fault attribute extraction module 22, and a time series and topology connectivity verification module 23, wherein: The fault propagation gradient field construction module 21 receives the structured topology dataset and the equipment operation dataset, generates the fault propagation gradient field, and outputs it to the fault attribute extraction module 22. The process of generating the fault propagation gradient field in the fault propagation gradient field construction module 21 includes the following steps: S21.1 Receive the structured topology dataset and the device operation dataset, extract the anomaly degree, anomaly occurrence time and link basic weight of each node, and complete data normalization and time sequence alignment. In S21.1, the fault propagation gradient field construction module 21 extracts the time series data of the performance indicators of each node during the monitoring period from the equipment operation dataset, determines the abnormality degree value and the time of abnormality occurrence of each node, and reads the basic weight of each link from the link table of the structured topology dataset to complete the normalization processing of the abnormality degree value and the time series alignment of the time of abnormality occurrence of each node.

[0027] The equipment operation dataset centrally stores multi-dimensional time-series performance data indexed by node ID. The fault propagation gradient field construction module 21 selects data related to fault propagation from this dataset. Performance metrics are used to calculate the degree of anomalies; for example, CPU utilization, memory utilization, and interface packet loss rate can be selected. For ease of reference in subsequent formulas, let's denote... This represents the total number of performance metrics selected.

[0028] The anomaly score of a node measures the overall degree to which the node deviates from its normal operating state during the current monitoring period. Let node... During the monitoring period, each sampling point was at the [number]th [number]. The statistical representative value of the measured value of the indicator (the maximum value of the indicator within the monitoring period can be selected) is: Calculate the normalized outlier score for this indicator using the following formula: ; In the formula, Represents a node In the The normalized outlier score on the performance index is dimensionless and ranges from [0,1]. Indicates the first The normal baseline value of each performance indicator is obtained and pre-configured during the system initialization phase by offline statistical analysis of historical stable data from long-term operation of the equipment. Indicates the first The abnormal saturation reference value of a performance indicator is taken as the historical maximum value or the upper limit of the physical range of that indicator.

[0029] Fault propagation gradient field construction module 21 obtains node all After calculating the normalized outlier scores for each indicator, the maximum value among the normalized outlier scores is taken as the degree of outlier for that node. ; In the formula, Represents a node The anomaly degree value is dimensionless and ranges from [0,1]. The larger the value, the more significant the anomaly degree of the node. The more serious the deviation of the operating state from the normal state.

[0030] The time of an anomaly occurrence for a node is determined by the degree of the anomaly, starting from a value below a preset anomaly threshold. Transform into equal to or higher than The first sampling time. Among them, To preset the anomaly threshold, a dimensionless value is set based on network operation and maintenance experience, for example, a value between 0.3 and 0.5. Assume a node... The time series of abnormality values ​​is as follows ,in Indicates the first Each sampling time, This indicates the total number of sampling times within the monitoring period, and the time of anomaly occurrence. Determine using the following formula: ; In the formula, Represents a node The time of an anomaly occurrence is measured in seconds. If the anomaly severity value of node i is lower than a certain value at all sampling times within the monitoring period, then the anomaly severity value is considered normal. If no abnormality occurs during the monitoring period, the node will not be included in the determination of upstream or downstream abnormal nodes in the subsequent gradient calculation.

[0031] S21.2 Calculate the difference in anomaly severity and the difference in anomaly occurrence time for each link, and calculate the propagation gradient value of a single link by combining the link's basic weight and the scenario constraint correction coefficient; the propagation gradient value of a single link is calculated using the following formula: ; In the formula, For link The propagation gradient value; upstream node The anomaly level value is taken from the device operation dataset; For downstream nodes The anomaly level value is taken from the device operation dataset; This is the time difference between the occurrence of anomalies in the downstream node and the upstream node, taken from the time stamp of the device operation dataset; For link The base weights are taken from the structured topology dataset; The scene constraint correction coefficient is used; when the time difference between the occurrence of anomalies in upstream and downstream nodes is zero, the propagation gradient value is the product of the link base weight and the scene constraint correction coefficient. In S21.2, the fault propagation gradient field construction module 21 calculates the difference in anomaly severity and the difference in anomaly occurrence time between upstream and downstream nodes on a link-by-link basis, and calculates the propagation gradient value of a single link by combining the link's basic weight and scenario constraint correction coefficient. This applies to nodes with interconnected relationships in the network. and nodes , defined by nodes Pointing to node A directed link is a link Downstream nodes With upstream nodes The time difference between the occurrence of the anomaly is calculated using the following formula: ; In the formula, Indicates downstream node With upstream nodes The time difference between the occurrence of the anomaly, in seconds.

[0032] when At that time, the link The conduction gradient value is calculated using the following formula: ; In the formula, Indicates link The propagation gradient value, with the dimension [time]. -1 .

[0033] when When downstream and upstream nodes exhibit anomalies at the same sampling time, it becomes difficult to distinguish the propagation direction based solely on the difference in anomaly severity and time difference. The propagation strength is primarily determined by topological weights and scene constraints. The propagation gradient value is the product of the link's base weights and the scene constraint correction coefficients. ; when If the link does not satisfy the temporal causal relationship of fault propagation from upstream to downstream, it is determined to be an invalid propagation link, the propagation gradient value is set to 0, and it is not included in the fault propagation gradient field.

[0034] Among them, the scene constraint correction coefficient is determined. The process includes the following steps: S21.21. Determine the link direction based on the network hierarchy attribute of the upstream and downstream nodes. A link from a higher network hierarchy node to a lower network hierarchy node is a downlink, and vice versa. Set the conduction attenuation coefficient according to the link direction. The conduction attenuation coefficient of the downlink is less than that of the uplink. In S21.21, the fault propagation gradient field construction module 21 determines the link direction based on the network hierarchy attributes of the nodes at both ends of the link. The network hierarchy attributes are provided by the node hierarchy attribute table in the structured topology dataset. Let the node... The network hierarchy is ,node The network hierarchy is The values ​​for each layer decrease sequentially from the core layer to the aggregation layer and then to the access layer. At that time, the link For downlink; when At that time, link This is for the uplink. The conducted attenuation coefficient is set according to the link direction. : ; In the formula, Indicates link The conduction attenuation coefficient is dimensionless; This represents the conduction attenuation coefficient of the downlink. This represents the uplink conduction attenuation coefficient. The conduction attenuation coefficients of links at the same level are all preset constants, dimensionless, and satisfy the following conditions: For example. Take 0.5, Take 1.0, We set it to 1.5. This setting is based on empirical patterns in network operation and maintenance practices: when a fault propagates from the core to the edge in the downlink direction, its impact is usually more significant and rapid, while when it propagates in the uplink direction, it is often weakened by the redundancy and protection mechanisms of the aggregation layer or the core layer. Therefore, the conduction attenuation coefficient of the downlink is smaller than that of the uplink.

[0035] S21.22. Adjust the conduction coefficient according to the relative relationship between the real-time load of the link and the preset load inflection point. When the real-time load of the link exceeds the preset load inflection point, the conduction coefficient increases by a step. In S21.22, the fault propagation gradient field construction module 21 determines the propagation coefficient based on the relative relationship between the real-time load of the link and the preset load inflection point. Assume the link... The real-time load is The throughput data of the corresponding interfaces at both ends of the link is extracted from the device operation dataset and calculated. For example, the average link utilization rate during the monitoring period is taken, which is dimensionless; the preset load inflection point is... Dimensionless, for example, taking 70% of the link bandwidth. Transmission coefficient. Determine using the following formula: ; In the formula, Indicates link The conductivity coefficient is dimensionless. This represents the conduction coefficient under normal load conditions. The conduction coefficients under congestion conditions are all preset constants, dimensionless, and satisfy the following conditions: ,For example Take 1.0, The value is set to 1.8. This step increase setting reflects the "amplifier" effect of high-load links in fault propagation, meaning that when a link approaches or reaches a congested state, a small performance degradation can trigger a significant chain reaction.

[0036] S21.23. Multiply the conduction attenuation coefficient by the conduction coefficient to generate the corresponding scenario constraint correction coefficient for the link. .

[0037] In S21.23, the fault propagation gradient field construction module 21 uses the propagation attenuation coefficient obtained in step S21.21. The conduction coefficient obtained in step S21.22 Multiply to obtain the link Scene constraint correction coefficient : .

[0038] S21.3. Summarize the propagation gradient values ​​of all links in the entire domain to form a fault propagation gradient field covering the entire network.

[0039] In S21.3, the fault propagation gradient field construction module 21 summarizes the propagation gradient values ​​of all links in the entire network that satisfy the temporal causality condition and have a propagation gradient value greater than 0, forming a fault propagation gradient field covering the entire network. The fault propagation gradient field is represented in the form of a set of directed links: ; In the formula, The fault propagation gradient field is represented and stored in the form of a link list or adjacency matrix. Each record contains the upstream node ID, downstream node ID, and the propagation gradient value of the link, which can be called by the fault attribute extraction module 22.

[0040] The fault attribute extraction module 22 extracts root cause fault nodes, secondary abnormal nodes, and fault influence domain topology subsets based on the fault propagation gradient field, and outputs the extraction results to the time series and topology connectivity verification module 23. The process of extracting fault attributes by the fault attribute extraction module 22 includes the following steps: S22.1 Calculate the node gradient value corresponding to each node, and take the maximum value of the propagation gradient value of all downstream links of the node as the node gradient value of the corresponding node. In S22.1, the fault attribute extraction module 22 calculates the node gradient value corresponding to each node. For any node... Its nodal gradient value Take the maximum value of the propagation gradient values ​​of all downstream links of this node: ; In the formula, Represents a node The nodal gradient values, measured in units of [time]. -1 The node gradient value characterizes the maximum intensity of fault propagation from that node to the downstream of the network. For network edge nodes or nodes that do not exhibit anomalous propagation, the fault propagation gradient field... There is no upstream link to this node, so its node gradient value is 0.

[0041] S22.2 Identify the extreme points in the gradient values ​​of all nodes and mark the corresponding nodes as root cause fault nodes; In S22.2, the fault attribute extraction module 22 identifies extreme points in the gradient values ​​of all nodes and marks the corresponding nodes as root cause fault nodes. The condition for being identified as an extreme point is: for the fault propagation gradient field All nodes For links to downstream nodes, the propagation gradient values ​​are all smaller than those of the node. The nodal gradient values, i.e.: ; If the fault propagation gradient field There is no node in For the link of the downstream node, then the node It automatically meets the criteria for determining extreme points.

[0042] The implication of this judgment logic is that the intensity of the fault propagated outward from the root cause node is greater than the intensity of the fault propagated to it from any other node, which conforms to the propagation characteristic of faults "radiating outward" from that node. Nodes that satisfy the above conditions... Nodes marked as root cause failures. All root cause failure nodes constitute the root cause failure node set. .

[0043] S22.3 Matching along the downstream path with continuously decaying gradient, and marking abnormal nodes on the path as secondary abnormal nodes; In S22.3, the fault attribute extraction module 22 performs matching along the downstream path of continuously decaying gradients, marking abnormal nodes on the path as secondary abnormal nodes. The matching condition for continuously decaying gradients is: for nodes already marked... Departure Link If node It is an abnormal node, and the link The propagation gradient value and the node The ratio of the node gradient values ​​is not lower than the preset gradient decay ratio threshold. That is, satisfying: ; In the formula, This represents the preset gradient decay ratio threshold, which is dimensionless. For example, a value of 0.3 can be used, which can be adjusted based on network size and tolerance for propagation chain length. If the above conditions are met, then the node... Marked as a secondary anomalous node, and with node Continue recursively tracing downstream from the new starting point; if the above conditions or nodes are not met... If a node is not anomalous, tracing along that path ceases. All secondary anomalous nodes constitute a set of secondary anomalous nodes. .

[0044] S22.4 Delineate the connected regions continuously covered by gradient values ​​and generate a topological subset of the fault-affected domain.

[0045] In S22.4, the fault attribute extraction module 22 delineates the connected regions continuously covered by gradient values, generating a fault influence domain topological subset. It consists of the union of the set of root cause failure nodes and the set of secondary anomalous nodes: ; In the formula, This represents a subset of the topological domain affected by the fault. Preserve the gradient field between nodes in the fault propagation process The existing link relationships form a connected topology structure corresponding to the topological subset of the fault-affected domain. In this subgraph, any node has at least one path originating from the root cause fault node and consisting entirely of the set. The path composed of internal nodes. When there are multiple root cause fault nodes, for each root cause fault node... From the set Extract the root cause node and all secondary abnormal nodes traced from the root cause node through step S22.3, forming an independent fault influence domain topology subset corresponding to the root cause node. Different root cause nodes corresponding to Secondary nodes that are partially shared are allowed between them.

[0046] The timing and topology connectivity verification module 23 performs timing and topology connectivity verification on the extracted results, generating fault location results with conduction attributes. The timing and topology connectivity verification module 23 includes a timing verification submodule and a topology connectivity verification submodule, wherein: The timing verification submodule performs timing verification on the extracted results, determining whether the anomaly occurrence time of the root cause fault node is earlier than the anomaly occurrence time of all nodes with anomaly occurrence times within the fault-affected domain topology subset. If the verification fails, the corresponding downstream path is removed. The verification condition is: for any root cause fault node... and the topological subset of the fault impact domain corresponding to the root cause fault node. Any node within that has the time of an anomaly occurrence All satisfy If a node exists Make This indicates that the anomaly of this node cannot be caused by the root cause failure node, the timing check fails, and the node and its slave nodes will be removed. Starting along the fault propagation gradient field In the mid-link direction, All downstream nodes and corresponding links traced within the range are from the set of secondary abnormal nodes. and corresponding Remove from the list.

[0047] The topology connectivity verification submodule performs topology connectivity verification on the extracted results, determining whether the topological subset of the fault impact domain is a connected region originating from the root cause fault node. If the verification fails, the corresponding discontinuous boundaries are truncated. The verification condition is: the topological subset of the fault impact domain after time-series verification correction. Topologically, it is composed of root cause failure nodes. The connected region originating from [a certain point]. If [the region is] connected to [a certain point]... There are root cause failure nodes in it. Disconnected isolated nodes or isolated subgraphs indicate that these isolated parts cannot be propagated from the root cause failure node through continuous topology links. The topology connectivity check fails, and the nodes or isolated subgraphs are truncated at discontinuous boundaries. Remove from the list.

[0048] The timing and topology connectivity verification module 23 generates a fault location result with conduction attributes based on the extracted results after verification and correction. If no abnormal nodes are identified in the entire network during the monitoring period, the fault conduction gradient field is empty, or no root cause fault nodes are identified after verification, it is determined that there are currently no effective conduction faults in the entire network. A fault-free location result is generated and output to the network device PHM status assessment unit 3, which performs routine health monitoring.

[0049] Specifically, the timing and topology connectivity verification module 23 generates a fault location result with conduction attributes based on the extracted result after verification correction. The result includes a set of root cause fault nodes. Secondary abnormal node set The propagation paths between secondary anomaly nodes and root cause failure nodes, and the topological subsets of the failure impact domains corresponding to each root cause failure node. The system generates a set of nodes and links, along with transmission attribute information such as the transmission gradient value of each link. The generated fault location results with transmission attributes are output to the network device PHM status assessment unit 3.

[0050] Network device PHM status assessment unit 3 receives fault location results with transmission attributes and device operation datasets. Using device health status modeling and trend extrapolation, it performs hierarchical health assessments and degradation trend analyses on root cause fault nodes, secondary abnormal nodes, and nodes within the fault impact domain, generating hierarchical PHM status assessment results and maintenance priority ranking results. Specifically, traditional device health assessments typically use a uniform assessment model for all alarm devices, failing to differentiate the roles of devices in fault transmission. This leads to mixed assessment results for root cause devices and affected devices, making it difficult to accurately allocate maintenance resources. Unlike such methods, this unit, based on the fault location results with transmission attributes output by topology-related fault point detection unit 2, performs different health assessment and degradation trend analyses according to the different roles of nodes in the fault event—root cause fault nodes, secondary abnormal nodes, and normal nodes within the fault impact domain—ensuring that the assessment results match the actual location and damage mechanism of the nodes in the fault transmission. Based on this, the maintenance priority ranking of the assessment results is combined with the fault impact domain range, providing an orderly basis for subsequent management and control instructions.

[0051] In this embodiment, the network device PHM status assessment unit 3 includes a hierarchical health assessment module 31 and an operation and maintenance priority ranking module 32, wherein: The graded health assessment module 31 receives fault location results with transmission attributes and equipment operation datasets, performs graded health assessments and degradation trend analysis on root cause fault nodes, secondary abnormal nodes, and nodes within the fault influence domain, generates graded PHM status assessment results, and outputs them to the maintenance priority ranking module 32. The graded health assessment and degradation trend analysis process in the graded health assessment module 31 includes the following steps: S31.1 Perform a deep degradation assessment on the root cause failure node to generate data on the degree and development trend of root cause failure degradation; In S31.1, the graded health assessment module 31 extracts the root cause fault node set from the fault location results with transmission attributes. A deep degradation assessment is performed on each root cause failure node within the set. The deep degradation assessment extracts time-series performance metrics data of the root cause failure node from the equipment operation dataset, covering the monitoring period and a historical time window preceding the monitoring period. From this data, it selects those closely related to equipment health degradation. Key performance indicators, such as CPU utilization, memory utilization, and interface packet loss rate.

[0052] For the For each key performance indicator, the graded health assessment module 31 calculates the deviation of the statistical representative value of this indicator from the historical normal baseline during the current monitoring period. Let the root cause failure node be... In the The statistical representative value for the current monitoring period for this indicator is The historical normal baseline value is The maximum allowable value for this indicator is The individual degradation score for this indicator is calculated using the following formula: ; In the formula, Indicates the root cause failure node In the The individual degradation score for each key performance indicator is dimensionless and ranges from [0,1]. Indicates the first The historical normal baseline value of each indicator is determined statistically from historical data of long-term stable operation of the equipment; Indicates the first The maximum allowable value for each indicator is set based on the equipment specifications or historical extreme values.

[0053] Furthermore, the graded health assessment module 31 obtains the root cause failure node. all After calculating the individual degradation scores for each indicator, the maximum value among these individual degradation scores is taken as the degradation degree value for that root cause failure node. ; In the formula, Indicates the root cause failure node The degradation degree value is dimensionless and ranges from [0,1]. The larger the value, the more severe the performance degradation of the node.

[0054] In addition, the graded health assessment module 31 also considers root cause failure nodes. The slope of the degradation scores for each key performance indicator during the monitoring period is used to generate trend data. This trend data can be qualitatively described; for example, if the degradation score continuously rises during the monitoring period and the rate of increase exceeds a preset acceleration threshold, it is considered an accelerating degradation trend; if the degradation score remains relatively stable, it is considered a maintaining trend; and if the degradation score declines, it is considered a slowing trend. The generated root cause node failure degradation degree value... The development trend data is written into the assessment record of the corresponding node in the hierarchical PHM status assessment results.

[0055] S31.2 Perform conduction damage assessment on secondary abnormal nodes and generate data on the degree of conduction damage and recovery trend of secondary nodes; In S31.2, the graded health assessment module 31 extracts the set of secondary abnormal nodes from the fault location results with conduction attributes. For each secondary anomalous node in the set, a conduction damage assessment is performed. Unlike the deep degradation assessment of the root cause failure node, the conduction damage assessment focuses on evaluating the degree of conduction damage suffered by the secondary anomalous node during the failure propagation process.

[0056] For secondary abnormal nodes The graded health assessment module 31 first extracts the time-series performance index data of the node from the equipment operation dataset, and calculates the preliminary degradation score of the node based on its own performance index using the same method as in S31.1, denoted as... Simultaneously, the graded health assessment module 31 obtains the transmission gradient values ​​of each link on the transmission path between the secondary abnormal node and the corresponding root cause fault node from the fault location results with transmission attributes, and takes the maximum value of the transmission gradient values ​​of each link on the transmission path as the maximum transmission gradient received by the node. The degree of conductive damage experienced by this node is calculated using the following formula: ; In the formula, Indicates secondary abnormal nodes The degree of conduction damage is a dimensionless value, ranging from [0,1] (when the calculated value exceeds 1, it is taken as 1). This represents the initial degradation score calculated based on the node's own performance indicators; This represents the maximum propagated gradient value received by this node, measured in units of [time]. -1 The link propagation gradient value is taken from the fault location result with propagation properties on the propagation path between the node and the corresponding root cause fault node. This is the reference value for the propagation gradient, measured in units of [time]. -1 , is a preset constant used to normalize the propagation gradient value to a scale comparable to the degradation score. This value can be set according to the network size and typical fault propagation intensity.

[0057] The graded health assessment module 31 also performs a recovery trend analysis on the time-series data of the performance indicators of the secondary abnormal node s. The recovery trend is determined by: if the initial degradation score of the node... If, during the later stages of the monitoring period, a sustained decline relative to the peak value occurs, and the decline exceeds a preset recovery threshold, it is considered to have a recovery trend; otherwise, it is considered to have an insignificant recovery trend or to be continuously deteriorating. The generated secondary node conduction damage degree value... The recovery trend data is written into the assessment record of the corresponding node in the hierarchical PHM status assessment results.

[0058] S31.3 Perform a preliminary risk warning assessment on normal nodes within the fault impact domain and generate risk level and warning data; In S31.3, the graded health assessment module 31 extracts the topological subsets of each fault influence domain from the fault location results with conduction properties. The system identifies nodes that are not marked as root cause failure nodes or secondary abnormal nodes; these nodes are considered normal nodes within the failure impact domain. For each normal node within the failure impact domain, the hierarchical health assessment module 31 performs a preliminary risk warning assessment.

[0059] The risk assessment for a normal node is determined by considering the following factors: the node's location within the topological subset of the fault-affected domain, the topological distance (in hops) between the normal node and the most recent secondary anomaly node or root cause fault node, and whether the node's own performance indicators show a deterioration trend even though they have not yet reached the anomaly threshold during the current monitoring period. Risk levels are categorized into high, medium, and low. For example, a normal node with a topological distance of 1 hop from a secondary anomaly node and a deteriorating performance indicator is classified as high-risk; a normal node with a topological distance of 2 hops and stable performance indicators is classified as medium-risk; and a normal node with a topological distance of 3 hops or more and no deteriorating performance indicator is classified as low-risk.

[0060] Furthermore, the tiered health assessment module 31 generates corresponding early warning data based on the risk level. For example, it generates an early warning message "Immediate attention recommended" for high-risk nodes and an early warning message "Continuous monitoring recommended" for medium-risk nodes. The generated risk level and early warning data are written into the assessment record of the corresponding node in the tiered PHM status assessment results.

[0061] S31.4 Summarize the evaluation data of various nodes and generate hierarchical PHM status evaluation results.

[0062] In S31.4, the graded health assessment module 31 summarizes the root cause failure node assessment data, secondary abnormal node assessment data, and normal node assessment data within the fault impact domain generated in steps S31.1 to S31.3 to generate a graded PHM status assessment result. The graded PHM status assessment result includes the node identifier, node role, corresponding health assessment data, and the fault impact domain identifier associated with each node. Node roles include three categories: root cause failure nodes, secondary abnormal nodes, and normal nodes within the impact domain; the health assessment data corresponds to the root cause node's failure degradation degree value according to the node role. Compared with development trend data and secondary node transmission damage degree values Combine recovery trend data or normal node risk level and early warning data. The generated graded PHM status assessment results are output to the operation and maintenance priority sorting module 32.

[0063] The maintenance priority ranking module 32 receives the hierarchical PHM status assessment results, sorts them in conjunction with the fault impact domain range, and generates maintenance priority ranking results.

[0064] Specifically, the maintenance priority ranking module 32 receives the hierarchical PHM status assessment results, sorts them based on the fault impact domain, and generates a maintenance priority ranking result. The maintenance priority ranking module 32 uses multi-dimensional ranking rules when ranking: The first ranking dimension is the node role. The operation and maintenance priority of the root cause failure node is higher than that of the secondary failure node, and the operation and maintenance priority of the secondary failure node is higher than that of the normal nodes in the failure impact domain. The second ranking dimension is the size of the topological subset of the fault-affected domain where the node is located. The more nodes the affected domain contains, the higher the priority. The third ranking dimension is the node's health assessment data; for root cause failure nodes, the degradation level value. The larger the value, the higher the priority. For secondary anomalous nodes, this indicates the degree of damage propagation. The larger the value, the higher the priority. For normal nodes within the fault impact domain, the higher the risk level, the higher the priority.

[0065] Finally, after multi-dimensional sorting, an operation and maintenance priority ranking result is generated, which includes a sequence of nodes arranged from high to low priority, along with information such as the node role, health assessment summary, and the size of the fault impact domain for each node. The generated operation and maintenance priority ranking result is output to the control instruction generation and interaction unit 4.

[0066] The control command generation and interaction unit 4 receives the hierarchical PHM status assessment results and maintenance priority ranking results. It then uses network control command encapsulation and scheduling conversion methods to process fault isolation strategies, traffic scheduling schemes, and maintenance tasks into commands, generating device control commands and maintenance scheduling commands, and outputting them to the corresponding network devices and maintenance management terminals. The control command generation and interaction unit 4 includes a hierarchical command generation module 41 and a command distribution and interaction module 42, wherein: The hierarchical instruction generation module 41 receives the hierarchical PHM status assessment results and the operation and maintenance priority ranking results, performs instruction-based processing on the fault isolation strategy, traffic scheduling scheme, and operation and maintenance tasks, generates device control instructions and operation and maintenance scheduling instructions, and outputs them to the instruction distribution and interaction module 42. Specifically, when performing instruction-based processing, the hierarchical instruction generation module 41 generates corresponding instructions for each node in descending order of priority according to the node sequence in the operation and maintenance priority ranking results. For each node in the operation and maintenance priority ranking results, the hierarchical instruction generation module 41 obtains the node role and health assessment data of that node from the hierarchical PHM status assessment results, and matches the corresponding instruction generation strategy according to the node role.

[0067] For the root cause failure node, the hierarchical instruction generation module 41 determines the degree of degradation of the node. Based on the fault-affected domain topology subset range, generate device management instructions corresponding to the fault isolation strategy, such as executing a shutdown instruction on the faulty port or a traffic bypass instruction on the faulty node. At the same time, generate operation and maintenance scheduling instructions, such as generating an emergency repair work order for the root cause device.

[0068] For secondary abnormal nodes, the hierarchical instruction generation module 41 determines the degree of conduction damage based on the node's value. Based on the recovery trend data, generate device management instructions corresponding to the traffic scheduling scheme, such as traffic rerouting instructions to schedule part or all of the traffic flowing through the node to the backup path, and generate operation and maintenance scheduling instructions, such as generating secondary device inspection work orders.

[0069] For nodes within the fault-affected domain that are assessed as high-risk or medium-risk, the graded instruction generation module 41 generates operation and maintenance scheduling instructions based on the node's risk level and early warning data. For example, it generates preventive inspection work orders for high-risk equipment or enhanced monitoring work orders for medium-risk equipment.

[0070] Finally, the hierarchical instruction generation module 41 outputs all generated device control instructions and operation and maintenance scheduling instructions to the instruction distribution and interaction module 42.

[0071] The instruction distribution and interaction module 42 receives device control instructions and operation and maintenance scheduling instructions, and outputs the corresponding instructions to the corresponding network devices and operation and maintenance management terminals. Specifically, when distributing instructions, the instruction distribution and interaction module 42 selects different distribution channels and interface protocols according to the instruction type. For device control instructions, the instruction distribution and interaction module 42 sends the instructions to the corresponding network devices through the southbound device management interface, such as based on the NETCONF protocol, RESTCONF protocol, or command line interface, and receives the response receipt after the device executes the instructions. For operation and maintenance scheduling instructions, the instruction distribution and interaction module 42 pushes work orders and early warning information to the operation and maintenance management terminal through the northbound interface with the operation and maintenance management terminal, such as based on message queues or RESTful APIs, for operation and maintenance personnel to process according to priority.

[0072] Those skilled in the art will understand that the process of implementing all or part of the steps of the above embodiments can be carried out by hardware or by a program instructing the relevant hardware.

[0073] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely preferred examples and are not intended to limit the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the claimed invention.

Claims

1. A communication fault point detection and PHM management system based on communication network topology, characterized in that, include: The communication network topology sensing and data acquisition unit (1) is used to acquire node link topology information, network hierarchical attributes and network device operation status data of the communication network, and to standardize and integrate the topology information, hierarchical attributes and operation status data to obtain structured topology dataset and device operation dataset. The topology-related fault point sniffing unit (2) receives a structured topology dataset and a device operation dataset. It uses a fault propagation gradient field construction method to calculate and process the node anomaly timing, link performance degradation and topology attributes to generate a fault propagation gradient field. Based on the fault propagation gradient field, it extracts the root cause fault node, secondary abnormal node and fault influence domain topology subset. After timing and topology connectivity verification, it generates a fault location result with propagation attributes. The network device PHM status assessment unit (3) receives the fault location results with conduction attributes and the equipment operation dataset. It adopts the equipment health status modeling and trend inference method to perform hierarchical health assessment and degradation trend analysis on the root cause fault nodes, secondary abnormal nodes and nodes in the fault influence domain, and generates hierarchical PHM status assessment results and operation and maintenance priority ranking results. The control instruction generation and interaction unit (4) receives the hierarchical PHM status assessment results and operation and maintenance priority ranking results, adopts the network control instruction encapsulation and scheduling conversion method, performs instruction processing on the fault isolation strategy, traffic scheduling scheme and operation and maintenance tasks, generates equipment control instructions and operation and maintenance scheduling instructions, and outputs them to the corresponding network devices and operation and maintenance management terminals.

2. The communication fault point detection and PHM management system based on communication network topology according to claim 1, characterized in that, The communication network topology sensing and data acquisition unit (1) includes a topology and hierarchical attribute acquisition module (11), a device operation status acquisition module (12), and a standardized integration processing module (13), wherein: The topology and hierarchical attribute acquisition module (11) acquires node link topology information and network hierarchical attributes of the communication network through distributed network detection, and outputs the original topology and hierarchical data to the standardized integration processing module (13). The device operation status acquisition module (12) acquires the operation status data of all network devices through the multi-source performance acquisition protocol and outputs the raw operation status data to the standardized integration processing module (13); The standardized integration processing module (13) receives the original topology and hierarchical data and the original operating status data, and completes the standardized integration processing through format unification, time sequence alignment and data cleaning to generate a structured topology dataset and a device operating dataset.

3. The communication fault point detection and PHM management system based on communication network topology according to claim 2, characterized in that, The topology-related fault detection unit (2) includes a fault propagation gradient field construction module (21), a fault attribute extraction module (22), and a time series and topology connectivity verification module (23), wherein: The fault propagation gradient field construction module (21) receives the structured topology dataset and the device operation dataset, generates a fault propagation gradient field and outputs it to the fault attribute extraction module (22). The fault attribute extraction module (22) extracts the root cause fault node, secondary abnormal node and fault influence domain topology subset based on the fault propagation gradient field, and outputs the extraction results to the time series and topology connectivity verification module (23). The timing and topology connectivity verification module (23) performs timing and topology connectivity verification on the extracted results and generates fault location results with conduction attributes.

4. The communication fault point detection and PHM management system based on communication network topology according to claim 3, characterized in that, The process of generating the fault propagation gradient field in the fault propagation gradient field construction module (21) includes the following steps: S21.1 Receive the structured topology dataset and the device operation dataset, extract the anomaly degree, anomaly occurrence time and link basic weight of each node, and complete data normalization and time sequence alignment. S21.2 Calculate the difference in anomaly severity and the difference in anomaly occurrence time for each link, and calculate the propagation gradient value of a single link by combining the link's basic weight and the scenario constraint correction coefficient; the propagation gradient value of a single link is calculated using the following formula: ; In the formula, For link The propagation gradient value; upstream node The anomaly level value is taken from the device operation dataset; For downstream nodes The anomaly level value is taken from the device operation dataset; This is the time difference between the occurrence of anomalies in the downstream node and the upstream node, taken from the time stamp of the device operation dataset; For link The base weights are taken from the structured topology dataset; The scene constraint correction coefficient is used; when the time difference between the occurrence of anomalies in upstream and downstream nodes is zero, the propagation gradient value is the product of the link base weight and the scene constraint correction coefficient. S21.

3. Summarize the propagation gradient values ​​of all links in the entire domain to form a fault propagation gradient field covering the entire network.

5. The communication fault point detection and PHM management system based on communication network topology according to claim 4, characterized in that, In step S21.2, the scene constraint correction coefficient is determined. The process includes the following steps: S21.

21. Determine the link direction based on the network hierarchy attribute of the upstream and downstream nodes. A link from a higher network hierarchy node to a lower network hierarchy node is a downlink, and vice versa. Set the conduction attenuation coefficient according to the link direction. The conduction attenuation coefficient of the downlink is less than that of the uplink. S21.

22. Adjust the conduction coefficient according to the relative relationship between the real-time load of the link and the preset load inflection point. When the real-time load of the link exceeds the preset load inflection point, the conduction coefficient increases by a step. S21.

23. Multiply the conduction attenuation coefficient by the conduction coefficient to generate the corresponding scenario constraint correction coefficient for the link. .

6. The communication fault point detection and PHM management system based on communication network topology according to claim 5, characterized in that, The process of extracting fault attributes by the fault attribute extraction module (22) includes the following steps: S22.1 Calculate the node gradient value corresponding to each node, and take the maximum value of the propagation gradient value of all downstream links of the node as the node gradient value of the corresponding node. S22.2 Identify the extreme points in the gradient values ​​of all nodes and mark the corresponding nodes as root cause fault nodes; S22.3 Matching along the downstream path with continuously decaying gradient, and marking abnormal nodes on the path as secondary abnormal nodes; S22.4 Delineate the connected regions continuously covered by gradient values ​​and generate a topological subset of the fault-affected domain.

7. The communication fault point detection and PHM management system based on communication network topology according to claim 6, characterized in that, The timing and topology connectivity verification module (23) includes a timing verification submodule and a topology connectivity verification submodule, wherein: The timing verification submodule performs timing verification on the extracted results, and determines that the abnormal occurrence time of the root cause fault node is earlier than the abnormal occurrence time of all nodes with abnormal occurrence time in the fault influence domain topology subset. If the verification fails, the corresponding downstream path is removed. The topology connectivity verification submodule performs topology connectivity verification on the extraction results, determines that the topology subset of the fault influence domain is a connected region starting from the root cause fault node, and cuts off the corresponding non-continuous boundary if the verification fails. The timing and topology connectivity verification module (23) generates fault location results with conduction properties based on the extracted results after verification correction.

8. The communication fault point detection and PHM management system based on communication network topology according to claim 7, characterized in that, The network device PHM status assessment unit (3) includes a hierarchical health assessment module (31) and an operation and maintenance priority ranking module (32), wherein: The graded health assessment module (31) receives the fault location results with transmission attributes and the equipment operation dataset, performs graded health assessment and degradation trend analysis on the root cause fault nodes, secondary abnormal nodes and nodes in the fault influence domain, generates graded PHM status assessment results and outputs them to the operation and maintenance priority ranking module (32). The operation and maintenance priority ranking module (32) receives the hierarchical PHM status assessment results, sorts them in combination with the fault impact domain range, and generates operation and maintenance priority ranking results.

9. The communication fault point detection and PHM management system based on communication network topology according to claim 8, characterized in that, The process of graded health assessment and degradation trend analysis in the graded health assessment module (31) includes the following steps: S31.1 Perform a deep degradation assessment on the root cause failure node to generate data on the degree and development trend of root cause failure degradation; S31.2 Perform conduction damage assessment on secondary abnormal nodes and generate data on the degree of conduction damage and recovery trend of secondary nodes; S31.3 Perform a preliminary risk warning assessment on normal nodes within the fault impact domain and generate risk level and warning data; S31.4 Summarize the evaluation data of various nodes and generate hierarchical PHM status evaluation results.

10. The communication fault point detection and PHM management system based on communication network topology according to claim 9, characterized in that, The control instruction generation and interaction unit (4) includes a hierarchical instruction generation module (41) and an instruction distribution and interaction module (42), wherein: The hierarchical instruction generation module (41) receives the hierarchical PHM status assessment results and operation and maintenance priority ranking results, performs instruction processing on the fault isolation strategy, traffic scheduling scheme and operation and maintenance tasks, generates equipment control instructions and operation and maintenance scheduling instructions and outputs them to the instruction distribution and interaction module (42). The instruction distribution and interaction module (42) receives the device control instruction and the operation and maintenance scheduling instruction, and outputs the corresponding instruction to the corresponding network device and operation and maintenance management terminal.

Citation Information

Patent Citations

  • Cascaded Failure Model and Node Vulnerability Assessment Method for Power Communication Convergence Networks

    CN114598612B

  • Communication equipment health detection system and method based on PHM technology

    CN116050156A