Network equipment fault intelligent detection method and system based on data transmission

By constructing a topology and link change rule base, network device faults can be monitored in real time, and the root node of the fault can be accurately located. This solves the problem of lack of topology structure in existing technologies and improves the accuracy and adaptability of fault detection.

CN120934995APending Publication Date: 2025-11-11SMIC WANYE TECHNOLOGY CO LTD
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
CN202511117843.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-11
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Existing technologies for network device fault detection lack the ability to construct a topology structure, making it impossible to clearly define the dependencies between devices. This leads to a high probability of misjudgment and the inability to dynamically adjust the topology hierarchy, reducing the timeliness and effectiveness of fault monitoring.

Method used

By constructing a rule base for topology structure and link changes, we can monitor topology hierarchical changes in real time, mark abnormal topology hierarchies and associated parent-child links, and accurately locate the root node of the fault using the link change rule base.

Benefits of technology

It improves fault location efficiency, reduces misjudgments, adapts to dynamic network changes, reduces the risk of fault propagation, and enhances fault detection effectiveness and adaptability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120934995A_ABST
    Figure CN120934995A_ABST
Patent Text Reader

Abstract

The invention discloses a network equipment fault intelligent detection method and system based on data transmission, and relates to the technical field of network equipment fault detection.A topological structure and a link change rule base are constructed, node addition, offline and hierarchical attribute change in the topological structure are captured in real time, and topological hierarchical change is sensed in real time, so that the network equipment fault intelligent detection method and system based on data transmission are realized. Meanwhile, monitoring each topology level, marking an abnormal topology level and associated father and son links, utilizing a link change rule base, accurately positioning a fault root node, determining a dependency relationship between equipment, accurately dividing a fault influence range, performing hierarchical troubleshooting, improving fault positioning efficiency, responding to dynamic network change in real time, and improving fault positioning accuracy. The fault diffusion risk is reduced, the timeliness and the detection effect of fault detection are improved, in addition, the fault detection requirements of different scenes are met, and the adaptability of fault detection is greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of network equipment fault detection technology, and specifically to an intelligent fault detection method and system for network equipment based on data transmission. Background Technology

[0002] Network devices, including routers, switches, firewalls, and servers, are the core carriers of information systems. Their stable operation directly affects the stability of business operations and public services. However, with the expansion of network scale, the increase in device types, and the improvement of business complexity, intelligent detection of network device faults can achieve real-time fault perception, accurate fault location, and early warning.

[0003] Existing technologies, such as the invention patent application CN120429149A, disclose a network fault management system and method based on multi-source data, relating to the field of server hard drive fault monitoring technology. Specifically, it includes the following steps: collecting hardware health data, service performance indicators, network traffic data, and historical fault records, and determining the collection frequency; acquiring key features of server hard drive faults and performing data quantification processing on the key features; using the pre-processed data as a benchmark, selecting several key feature data values ​​of server hard drives, and constructing a server hard drive fault probability prediction model through logistic regression; acquiring real-time key feature data, calculating the probability of server hard drive fault occurrence through the server hard drive fault probability prediction model, and comparing it with a preset classification range to obtain the fault level; and triggering automated response actions based on the fault level. This invention effectively reduces fault diagnosis time and service interruption risk, achieving more intelligent and efficient network fault management.

[0004] Existing technologies, such as the invention patent application CN120263628A, which discloses a method and system for rapid fault location of network devices based on historical data, relate to the field of network device fault detection technology. The method includes the following steps: collecting historical operating data of network devices; classifying network faults into three application scenarios and locating them according to different logics; cleaning and preprocessing the collected network device data, performing recursive feature elimination and filtering based on the preprocessing results, and matching real-time features with the historical database to calculate the priority of application scenarios; evaluating the degree of matching between current fault features and the three application scenarios, selecting the most matching scenario as the fault location strategy, and recording the execution priority order in the historical scenario database; for newly identified fault scenarios, expanding them as sub-scenarios, and statistically analyzing their occurrence frequency based on the number of occurrences to determine whether they need to be upgraded to the main scenario. This invention achieves rapid fault location of network devices.

[0005] The existing technology has at least the following shortcomings: 1. In network device applications, multiple devices are connected. When a device fails, it may affect other devices. Building a topology structure can provide underlying support for fault location, impact assessment and processing efficiency by clarifying the connection relationship and hierarchical dependency of devices in the network. However, the above-mentioned solutions lack the construction of a topology structure, cannot clarify the dependency relationship between devices, increase the probability of misjudging device faults, and cannot accurately divide the scope of fault impact and determine processing priorities. In addition, for complex networks, without a topology hierarchy, hierarchical investigation cannot be carried out, reducing the efficiency of fault location.

[0006] 2. Networks are not static structures but are in a state of dynamic change. If the topology information is outdated, fault monitoring based on the old topology will result in misjudgments. However, the above solutions lack dynamic perception of topology hierarchical changes and cannot dynamically adjust the logical structure in the topology, which reduces the timeliness of fault monitoring, increases the probability of misjudgments, and reduces the effectiveness of fault monitoring.

[0007] 3. A link refers to a physical or logical channel connecting two network nodes. It determines the direct transmission path of data from one node to another. The link change rule base in the topology can reflect the correlation between link changes and faults. However, the above scheme lacks the construction of a link change rule base, and cannot establish a deterministic correlation between "link change" and "fault type". Therefore, it cannot distinguish between normal link changes and faulty link changes, resulting in an increased false alarm rate. Furthermore, link changes often have a chain relationship. Without a rule base, the monitoring system can only perceive faults in isolation and cannot determine the relationship between faults, thus reducing the effectiveness of fault detection. Summary of the Invention

[0008] To address the aforementioned technical shortcomings, the present invention aims to provide a method and system for intelligent fault detection of network devices based on data transmission.

[0009] To solve the above-mentioned technical problems, the present invention adopts the following technical solution: In the first aspect, the present invention provides a network device fault intelligent detection method based on data transmission, including the following steps: S1, acquiring historical network device monitoring records, processing and analyzing the historical network device monitoring records, and constructing a topology structure and link change rule library.

[0010] S2. Capture the addition, disconnection, and hierarchical attribute changes of nodes in the topology structure, monitor the status of each parent-child link in the topology structure, perform real-time perception of topology hierarchical changes, and monitor each topology hierarchy to determine the status of data transmission. If the status is abnormal, execute S3.

[0011] S3. Mark the abnormal topology level and associated parent and child links, and use the link change rule base to accurately locate the fault root node.

[0012] Secondly, the present invention provides a network device fault intelligent detection system based on data transmission, comprising: a rule building module, used to acquire historical network device monitoring records, process and analyze the historical network device monitoring records, and build a topology and link change rule library.

[0013] The anomaly monitoring module is used to capture the addition and disconnection of nodes and changes in hierarchical attributes in the topology, monitor the status of each parent-child link in the topology, perceive changes in the topology hierarchy in real time, monitor each topology hierarchy, determine the status of data transmission, and execute the fault location module if the status is abnormal.

[0014] The fault location module is used to mark the abnormal topology level and the associated parent and child links, and to accurately locate the root node of the fault by using the link change rule base.

[0015] The beneficial effects of this invention are as follows: This invention provides a network device fault intelligent detection method and system based on data transmission. By constructing a topology and link change rule base, and capturing in real time the addition, disconnection, and hierarchical attribute changes of nodes in the topology, as well as real-time perception of topology hierarchical changes, it monitors each topology hierarchical level, marks abnormal topology hierarchical levels and associated parent-child links, accurately locates the fault root node using the link change rule base, accurately divides the fault impact range by clarifying the dependencies between devices, performs hierarchical investigation, improves fault location efficiency, responds to dynamic network changes in real time, reduces the risk of fault propagation, improves the timeliness of fault detection, and enhances fault correlation analysis capabilities, avoiding viewing anomalies in isolation and improving fault detection effectiveness. In addition, it adapts to the fault detection needs of different scenarios, significantly improving the adaptability of fault detection. Attached Figure Description

[0016] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0017] Figure 1 This is a schematic diagram of the implementation steps of the method of the present invention.

[0018] Figure 2 This is a schematic diagram of the system structure connection of the present invention. Detailed Implementation

[0019] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0020] See Figure 1 As shown, a network device fault intelligent detection method based on data transmission includes the following steps: S1, acquiring historical network device monitoring records, processing and analyzing the historical network device monitoring records, and constructing a topology and link change rule base.

[0021] It should be noted that historical network device monitoring records are extracted from the data center.

[0022] In a specific embodiment, the specific process of S1 is as follows: S1.1, collect network device operation logs, topology change records and link performance indicators within a preset period from historical network device monitoring records, remove transient interference data through data cleaning, and then standardize the data format.

[0023] Among them, data cleaning to remove transient interference data and standardization of data formats are existing technologies that can be found on the Internet, and will not be elaborated on here.

[0024] It should be noted that link performance metrics are key parameters for measuring the data transmission capability and quality of a link, including bandwidth, throughput, and latency.

[0025] S1.2 Based on device performance, the topology is divided into core layer, aggregation layer, access layer and terminal layer. Each node in each topology layer is marked with a layer attribute. At the same time, the parent and child links between the topology layers are defined. Fault records within a preset period are obtained from historical network device monitoring records to establish a layer influence mapping table. The layer influence mapping table is used to set weights for each parent and child link. The topology structure consists of each topology layer and each parent and child link.

[0026] It should be noted that, based on their expertise and testing needs, professionals set performance index ranges for each topology level. When the performance index of a device in a node falls within the performance index range of a certain topology level, that node is considered a node in that topology level, thus completing the layering. The performance index of the device in a node is a key indicator reflecting the operating status, resource utilization efficiency, and processing capacity of the device within the node, including CPU utilization, memory utilization, and network interface traffic.

[0027] In the above, the hierarchical attribute is marked as follows: for example, if node A is in the core layer, it is marked as core layer-A node; the example is only for illustration and is not the only limitation.

[0028] The labeling of hierarchical attributes is the process of assigning clear identifiers to different nodes according to their position in the hierarchical structure. It can intuitively distinguish the roles and functions of different nodes in the architecture, and realize the orderly management, efficient operation and maintenance and accurate analysis of network resources.

[0029] Parent-child link: A parent-child link specifically refers to a directional connection link between an upstream node (parent node) and a downstream node (child node). In a four-level network topology, a directional connection link between adjacent upstream nodes (parent nodes) and downstream nodes (child nodes) carries unidirectional dominant traffic, transmits hierarchical responsibilities, and strictly follows a single-level span. It is the core channel for maintaining hierarchical subordinate relationships and functional transmission.

[0030] Parent-child links only allow connections between "adjacent levels," meaning the parent node and child node must be one level apart. Traffic in parent-child links primarily flows from upstream to downstream (e.g., core layer → aggregation layer → access layer → terminal layer). Return traffic from downstream nodes must be transmitted back through the same link, but the control authority of the link (e.g., bandwidth allocation, policy configuration) is dominated by the upstream parent node. A downstream child node can connect to multiple upstream parent nodes, but all connections belong to the same parent-child link pair, and the core configuration of the child node must depend on the parent node and cannot operate independently of the parent node.

[0031] For example, the link between a node in the core layer and a node in the aggregation layer can be called a parent-child link, where the node in the core layer is the parent node in the parent-child link, and the node in the aggregation layer is the child link in the parent-child link.

[0032] Preferably, the weight setting for each parent-child link is specifically as follows: S1.2-1, obtain the impact range ratio of each node failure in each topology level from the fault records within a preset period, calculate the reference impact factor corresponding to each topology level, and construct a hierarchical impact mapping table.

[0033] In the above, the impact range ratio refers to the ratio of the number of nodes directly or indirectly affected after a fault occurs at a certain node in a certain layer of the topology to the total number of nodes; where the number of nodes directly or indirectly affected is the total number of nodes in the abnormal links marked with anomalies in the topology structure.

[0034] The weighted average of the impact range ratios corresponding to the faults of each node in each topology level is calculated, and the result is used as the impact range ratio of each topology level. The impact range ratio of each topology level is divided by the sum of the impact range ratios of all topology levels to obtain the reference impact factor of each topology level. The reference impact factors of each topology level constitute the hierarchical impact mapping table.

[0035] S1.2-2. Obtain the number of faults in each parent-child link from the fault records within the preset period, the impact range ratio, link switching rate and recovery time corresponding to each fault, and obtain the topology level to which the node of each fault in each parent-child link belongs. Calculate the fault severity feature value of each parent-child link using the hierarchical impact mapping table, and set the weight of each parent-child link using the fault severity feature value of each parent-child link.

[0036] The calculation process of the fault severity characteristic value of each parent-child link mentioned above is as follows: the number of faults in each parent-child link, the impact range ratio, link switching rate, recovery time and the topology level to which the node of each fault in each parent-child link belongs are preprocessed to obtain the number of node faults, average impact range ratio, average link switching rate, average recovery time and reference impact factor in each topology level of each parent-child link, and then the data format is standardized.

[0037] It should be noted that the topology levels of the parent and child nodes in each parent-child link and the topology levels of the nodes to which each fault occurred are obtained from the fault records. The number of node faults in each topology level of each parent-child link, as well as the impact range ratio, link switching rate, and recovery time corresponding to each fault, are obtained. The average of the impact range ratio, link switching rate, and recovery time corresponding to each fault is calculated to obtain the number of node faults, average impact range ratio, average link switching rate, and average recovery time in each topology level of each parent-child link. The reference impact factor for each topology level of each parent-child link is obtained from the level impact mapping table.

[0038] Substitute the standardized node failure count, average impact range ratio, average link switching rate, average recovery time, and reference impact factor of each parent-child link into the fault severity assessment model to output the fault severity characteristic value of each parent-child link.

[0039] Among them, the number of node failures, average impact range ratio, average link switching rate, average recovery time, and reference impact factor in each topology level of each parent-child link are denoted as a1. jk a2 jk a3 jk a4 jk and a5 jk Substitute the expression from the fault severity assessment model:

[0040] In the formula, j represents the number of each parent-child link, j is a positive integer, k represents the number of each topology layer, k = 1, 2, 3, 4, where 1, 2, 3, 4 represent the core layer, aggregation layer, access layer, and terminal layer, respectively, K represents the number of topology layers, and γ1, γ2, γ3, γ4 represent the weighting factors of the number of node failures, the average impact range ratio, the average link switching rate, and the average recovery time, respectively.

[0041] Among them, γ1, γ2, γ3, and γ4 are set and modified by professionals based on their expertise and testing needs, and no specific numerical restrictions are imposed here.

[0042] The weight of each parent-child link is obtained by dividing the fault severity characteristic value of each parent-child link by the sum of the fault severity characteristic values ​​of each parent-child link.

[0043] The core purpose of assigning weights to parent-child links is to accurately characterize the differentiated features of the links. Without weights, the topology becomes an abstract model with only connections, failing to reflect these differences and causing topology-based analysis to become detached from reality. Weights quantify these differences numerically, upgrading the topology model from a qualitative description to a quantitative analysis tool. Furthermore, when tracing equipment faults at multiple levels, priorities can be set based on weights, improving the efficiency of the tracing process.

[0044] S1.3 Extract the link change rules and hierarchical tracing rules corresponding to the events of each node in each topology level from the historical network device monitoring records, and build a link change rule library.

[0045] The above-mentioned link change rule base construction process is as follows: S1.3-1, extract the original parent-child link and the subsequent parent-child link from the link change rules corresponding to the events of each node in each topology layer from the historical network device monitoring records, and obtain the performance indicators and weights of the original parent-child link and the subsequent parent-child link. At the same time, based on the hierarchical tracing rules, score the link change rules corresponding to the events of each node in each topology layer.

[0046] It should be noted that hierarchical tracing rules include forward hierarchical tracing rules and reverse hierarchical tracing rules. Forward hierarchical tracing rules trace from upstream to downstream. For example, when an event occurs at a core layer node, forward tracing prioritizes its impact on the downstream aggregation layer. For instance, if a port on a core switch fails, according to the rules, the status of the directly connected aggregation switch links is first checked to determine if the port failure caused data transmission interruption or link switching. Furthermore, based on the link change rule base, it is determined whether the core-to-aggregation link weights need adjustment to maintain overall network performance.

[0047] The reverse hierarchical tracing rule follows a downstream-to-upstream approach, based on the hierarchical relationship of parent-child links in the topology. Reverse tracing proceeds in the order of "terminal layer → access layer → aggregation layer → core layer." For example, if a terminal layer device experiences a network connectivity problem, this rule first checks the link between the terminal layer and access layer devices, then checks the uplink links of the access layer devices and the status of the aggregation switch, gradually tracing back to the core layer devices to avoid missing fault points due to skipping levels of investigation. The example hierarchical tracing rule is for illustrative purposes only and is not the only valid approach.

[0048] In a parent-child link, when a child node fails, the tracing process is performed back to the parent node according to the reverse hierarchical tracing rules. For example, if packet loss is detected in the parent-child link from the terminal layer (child node) to the access layer (parent node), the link of the access layer parent node is checked according to the reverse hierarchical tracing rules to determine whether the failure of the parent node caused the child node's abnormality. At the same time, the optimization of the data transmission path is evaluated.

[0049] Optimize data transmission paths: The network structure has redundant links. When a link fails, the redundant link is activated and the failed link is shut down to reduce the impact on network operation.

[0050] In the above, the link change rules corresponding to the events of each node in each topology level are scored. The specific scoring process is as follows: the performance indicators of the original parent-child link and the subsequent parent-child link are compared, and the number of parameters in the performance indicators of the original parent-child link that are less than the parameters in the performance indicators of the subsequent parent-child link is counted as the performance improvement quantity. Then, the performance improvement quantity is standardized, and the processed value is recorded as the performance improvement value. Whether the subsequent parent-child link satisfies the adjacent level connection and does not cross level, if it satisfies it, the link matching value is recorded as 1; if it does not satisfy it, the link matching value is recorded as 0. If the weight of the original parent-child link is less than the weight of the subsequent parent-child link, the link security value is recorded as 1; otherwise, it is recorded as 0. The score = performance improvement value + link matching value + link security value.

[0051] S1.3-2. Cluster and merge the events of each node in each topology level, obtain the comprehensive score of each link change rule corresponding to each node event class in each topology level, and then construct the link change rule library.

[0052] It should be noted that clustering events according to type is an existing technology and will not be elaborated here; for example, bandwidth congestion of aggregation layer switches and CPU overload of firewalls can be clustered as aggregation layer resource overload events.

[0053] The average score of each link change rule in each node event class in each topology level is calculated, and the result is used as the comprehensive score of each link change rule in each node event class in each topology level.

[0054] S2. Capture the addition, disconnection, and hierarchical attribute changes of nodes in the topology structure, monitor the status of each parent-child link in the topology structure, perform real-time perception of topology hierarchical changes, and monitor each topology hierarchy to determine the status of data transmission. If the status is abnormal, execute S3.

[0055] In a specific embodiment, the specific process of S2 is as follows: S2.1, real-time acquisition of the addition, disconnection and hierarchical attribute changes of nodes in each topology level through a real-time monitoring protocol, updating the addition and disconnection of nodes in each topology level in the topology structure, and reclassifying each node in each topology level in the topology structure according to the attribute changes of the nodes, correcting the weight of the parent-child links associated with the nodes, and then updating the topology structure.

[0056] In the above, the process of changing the hierarchical attribute is as follows: obtain the performance indicators of each node device in each topology level, and obtain the performance indicator range corresponding to each topology level. When the performance indicator of a node device in a certain topology level is not within the performance indicator range corresponding to that topology level, the hierarchical attribute of the device needs to be changed. Then, the performance indicator of the device is compared with the performance indicator ranges of other topology levels, and the topology level corresponding to the performance indicator within the performance indicator range is taken as the changed hierarchical attribute of the device.

[0057] S2.2 Monitor the performance indicators of each topology level in the topology structure, obtain the performance indicators of each topology level, and compare them with their performance indicator thresholds. If the performance indicator of at least one topology level is greater than its performance indicator threshold, the data transmission status is judged to be abnormal; otherwise, it is normal.

[0058] It should be noted that a combination of distributed probes and device native counters is used to detect the performance indicators of each node device in each topology layer and the performance indicators of each topology layer in the topology structure.

[0059] The performance indicator thresholds for each topology level are reference values ​​for judging whether the level's performance is normal. They are set by professionals according to network service requirements, and no specific numerical limits are imposed here.

[0060] S3. Mark the abnormal topology level and associated parent and child links, and use the link change rule base to accurately locate the fault root node.

[0061] In a specific embodiment, the specific process of S3 is as follows: S3.1, topology layers with performance indicators greater than their performance indicator thresholds are regarded as abnormal topology layers. If the parent node in a parent-child link belongs to an abnormal topology layer, all its downstream child links are marked as abnormal links. If the child node belongs to an abnormal topology layer and is confirmed to be related to the parent node through tracing rules, all links between the child node and the parent node are marked as abnormal links. This is how abnormal marking is performed.

[0062] Among them, for link interruptions confirmed to be caused by node offline (such as equipment shutdown), even if the performance indicators exceed the standard, if the node status record of S2.1 determines that it is "normal offline", it will not be marked as an abnormal parent-child link to avoid interference.

[0063] S3.2 Obtain the indicator data and link change rules for each abnormal link. Compare the link change rules of each abnormal link with the link change rules of each node event class in the abnormal topology level of the link change rule library. If the link change rule of an abnormal link is the same as the link change rule of a node event class in the abnormal topology level of the link change rule library, then the node is regarded as a fault node, and the event class of the node is regarded as the fault class of the fault node. At the same time, obtain the score of the link change rule of the fault node event class in the abnormal topology level from the link change rule library, thereby obtaining the score of the link change rule of each fault node event class.

[0064] S3.3. Based on the scoring of the link change rules of each fault node event class, determine whether each fault node is a fault root node.

[0065] Preferably, the specific process for determining whether each fault node is a fault root node is as follows: when the score of the link change rule of a certain fault node event class is greater than or equal to a preset score threshold, and the tracing rule indicates that the fault node is the upstream node of the corresponding abnormal link, then the fault node is determined to be a fault root node.

[0066] It should be noted that the upstream node is the parent node.

[0067] If the score of the link change rule for a certain fault node event is less than the preset score threshold, the correlation between the fault node and the abnormal link is calculated. If the correlation is greater than the preset correlation threshold, the fault node is determined to be the root node of the fault. Otherwise, the hardware indicators and configuration indicators of the fault node are obtained. When the hardware indicators are not within the normal range, or the configuration change time in the configuration indicators is less than the preset time, the fault node is determined to be the root node of the fault. This method is used to determine whether each fault node is the root node of the fault.

[0068] It should be noted that the performance indicators of the faulty node and the abnormal link are obtained, and then the correlation between the faulty node and the abnormal link is calculated using the Pearson correlation coefficient. The calculation method of the Pearson correlation coefficient is existing technology and can be found on the Internet, so it will not be elaborated here.

[0069] In the above, the preset score threshold, correlation threshold, normal range of hardware indicators, and preset duration are the critical values ​​for whether a segment is faulty. These are all set by professionals according to the testing requirements, and no specific numerical restrictions are imposed here.

[0070] See Figure 2 As shown, a network device fault intelligent detection system based on data transmission includes: a rule construction module, an anomaly monitoring module, and a fault location module.

[0071] The rule building module is used to obtain historical network device monitoring records, process and analyze these records, and build a rule base for topology and link changes.

[0072] The anomaly monitoring module is used to capture the addition and disconnection of nodes and changes in hierarchical attributes in the topology, monitor the status of each parent-child link in the topology, perceive changes in the topology hierarchy in real time, monitor each topology hierarchy, determine the status of data transmission, and execute the fault location module if the status is abnormal.

[0073] The fault location module is used to mark the abnormal topology level and the associated parent and child links, and to accurately locate the root node of the fault by using the link change rule base.

[0074] The above description is merely an example and illustration of the concept of the present invention. Those skilled in the art can make various modifications or additions to the specific embodiments described or use similar methods to replace them, as long as they do not deviate from the concept of the invention or exceed the scope defined in this specification, they should all fall within the protection scope of the present invention.

Claims

1. A method for intelligent fault detection of network devices based on data transmission, characterized in that, Includes the following steps: S1. Obtain historical network device monitoring records, process and analyze the historical network device monitoring records, and construct a topology and link change rule base; S2. Capture the addition, disconnection, and hierarchical attribute changes of nodes in the topology, monitor the status of each parent-child link in the topology, perceive topology hierarchical changes in real time, monitor each topology hierarchical level, determine the status of data transmission, and if the status is abnormal, execute S3. S3. Mark the abnormal topology level and associated parent and child links, and use the link change rule base to accurately locate the fault root node.

2. The intelligent fault detection method for network devices based on data transmission according to claim 1, characterized in that, The specific process of S1 is as follows: S1.1 Collect network device operation logs, topology change records, and link performance indicators within a preset period from historical network device monitoring records. Remove transient interference data through data cleaning and then standardize the data format. S1.2 Based on device performance, the topology is divided into core layer, aggregation layer, access layer and terminal layer. Each node in each topology layer is marked with a layer attribute. At the same time, the parent and child links between the topology layers are defined. Fault records within a preset period are obtained from historical network device monitoring records to establish a layer influence mapping table. Using the layer influence mapping table, weights are set for each parent and child link. The topology structure consists of each topology layer and each parent and child link. S1.3 Extract the link change rules and hierarchical tracing rules corresponding to the events of each node in each topology level from the historical network device monitoring records, and build a link change rule library.

3. The intelligent fault detection method for network devices based on data transmission according to claim 2, characterized in that, The process of setting weights for each parent-child link is as follows: S1.2-1. Obtain the impact range ratio of each node failure in each topology level from the fault records within the preset period, calculate the reference impact factor corresponding to each topology level, and construct the hierarchical impact mapping table. S1.2-2. Obtain the number of faults in each parent-child link from the fault records within the preset period, the impact range ratio, link switching rate and recovery time corresponding to each fault, and obtain the topology level to which the node of each fault in each parent-child link belongs. Calculate the fault severity feature value of each parent-child link using the hierarchical impact mapping table, and set the weight of each parent-child link using the fault severity feature value of each parent-child link.

4. The intelligent fault detection method for network devices based on data transmission according to claim 3, characterized in that, The calculation process for the fault severity characteristic values ​​of each parent-child link is as follows: The number of failures in each parent-child link, the impact range ratio of each failure, the link switching rate, the recovery time, and the topology level of each node to which each failure occurs in each parent-child link are preprocessed to obtain the number of node failures, average impact range ratio, average link switching rate, average recovery time, and reference impact factor in each topology level of each parent-child link. Then, the data format is standardized. Substitute the standardized node failure count, average impact range ratio, average link switching rate, average recovery time, and reference impact factor of each parent-child link into the fault severity assessment model to output the fault severity characteristic value of each parent-child link.

5. The intelligent fault detection method for network devices based on data transmission according to claim 2, characterized in that, The construction process of the link change rule base is as follows: S1.3-1 Extract the original parent-child link and the subsequent parent-child link from the trigger link change rules corresponding to the events of each node in each topology level from the historical network device monitoring records, and obtain the performance indicators and weights of the original parent-child link and the subsequent parent-child link. At the same time, based on the hierarchical tracing rules, score the trigger link change rules corresponding to the events of each node in each topology level. S1.3-2. Cluster and merge the events of each node in each topology level, obtain the comprehensive score of each link change rule corresponding to each node event class in each topology level, and then construct the link change rule library.

6. The intelligent fault detection method for network devices based on data transmission according to claim 1, characterized in that, The specific process of S2 is as follows: S2.

1. Real-time monitoring protocol is used to obtain the addition, disconnection and hierarchical attribute changes of nodes in each topology level in real time, and the addition and disconnection of nodes in each topology level are updated in the topology structure. At the same time, according to the attribute changes of the nodes, the nodes in each topology level in the topology structure are reclassified, and the weights of the parent and child links associated with the nodes are corrected, and then updated in the topology structure. S2.2 Monitor the performance indicators of each topology level in the topology structure, obtain the performance indicators of each topology level, and compare them with their performance indicator thresholds. If the performance indicator of at least one topology level is greater than its performance indicator threshold, the data transmission status is judged to be abnormal; otherwise, it is normal.

7. The intelligent fault detection method for network devices based on data transmission according to claim 6, characterized in that, The process of changing hierarchical attributes is as follows: Obtain the performance metrics of each node device in each topology layer, and obtain the performance metric range corresponding to each topology layer. When the performance metric of a node device in a certain topology layer is not within the performance metric range corresponding to that topology layer, the layer attribute of the device needs to be changed. Then, compare the performance metric of the device with the performance metric ranges of other topology layers, and take the topology layer corresponding to the performance metric within the performance metric range as the changed layer attribute of the device.

8. The intelligent fault detection method for network devices based on data transmission according to claim 1, characterized in that, The specific process of S3 is as follows: S3.

1. Topology layers with performance indicators exceeding their performance indicator thresholds are considered abnormal topology layers. If the parent node in a parent-child link belongs to an abnormal topology layer, all its downstream child links are marked as abnormal links. If a child node belongs to an abnormal topology layer and is confirmed to be related to its parent node through tracing rules, all links between the child node and the parent node are marked as abnormal links. This is how abnormal marking is performed. S3.2 Obtain the indicator data and link change rules of each abnormal link. Compare the link change rules of each abnormal link with the link change rules of each node event class in the abnormal topology level in the link change rule library. If the link change rule of an abnormal link is the same as the link change rule of a node event class in the abnormal topology level in the link change rule library, then the node is regarded as a fault node, and the event class of the node is regarded as the fault class of the fault node. At the same time, obtain the score of the link change rule of the fault node event class in the abnormal topology level from the link change rule library, thereby obtaining the score of the link change rule of each fault node event class. S3.

3. Based on the scoring of the link change rules of each fault node event class, determine whether each fault node is a fault root node.

9. The intelligent fault detection method for network devices based on data transmission according to claim 8, characterized in that, The specific process for determining whether each faulty node is a root fault node is as follows: If the score of the link change rule of a certain fault node event class is greater than or equal to the preset score threshold, and the tracing rule indicates that the fault node is the upstream node of the corresponding abnormal link, then the fault node is determined to be the fault root node. If the score of the link change rule for a certain fault node event is less than the preset score threshold, the correlation between the fault node and the abnormal link is calculated. If the correlation is greater than the preset correlation threshold, the fault node is determined to be the root node of the fault. Otherwise, the hardware indicators and configuration indicators of the fault node are obtained. When the hardware indicators are not within the normal range, or the configuration change time in the configuration indicators is less than the preset time, the fault node is determined to be the root node of the fault. This method is used to determine whether each fault node is the root node of the fault.

10. A network device fault intelligent detection system that implements the network device fault intelligent detection method based on data transmission according to any one of claims 1-9, characterized in that, include: The rule building module is used to obtain historical network device monitoring records, process and analyze these records, and build a rule base for topology and link changes. The anomaly monitoring module is used to capture the addition and disconnection of nodes and changes in hierarchical attributes in the topology, monitor the status of each parent and child link in the topology, perceive changes in the topology hierarchy in real time, monitor each topology hierarchy, determine the status of data transmission, and execute the fault location module if the status is abnormal. The fault location module is used to mark the abnormal topology level and the associated parent and child links, and to accurately locate the root node of the fault by using the link change rule base.

Citation Information

Patent Citations

  • Network equipment fault rapid positioning method and system based on historical data

    CN120263628A

  • Network fault management system and method based on multi-source data

    CN120429149A

  • Method and device for diagnosing root of network faults

    CN104796273A

  • Abnormal root cause analysis method and device and storage medium

    CN112882796A

  • Abnormal node positioning method and device, equipment and storage medium

    CN115437871A