Multi-source network equipment fault mode identification and positioning method and system
By collecting and fusion of multi-source data from network equipment, and generating fault feature vectors and abnormal probability distribution maps, the problems of low fault recognition accuracy and inaccurate positioning in the existing technology are solved, and accurate identification and rapid positioning of network equipment faults are achieved, and the stability and reliability of the network system are improved.
Patent Information
- Application Number
- CN202510828719.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-20
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-06-20
AI Technical Summary
The existing network equipment fault identification and positioning technology lacks the comprehensive analysis ability of multi-source data at the hardware and software levels, resulting in a single-sided fault judgment and the inability to fully grasp the operating status of the equipment. Especially in complex fault modes, the identification accuracy is low, and the failure correlation analysis between multiple devices is not effectively handled, and it is difficult to accurately locate the root cause of the fault, and the ability to adapt to dynamic changes in the network environment, resulting in frequent false alarms and missed alarms.
Collect multi-source data of network equipment, generate fault feature matrix of hardware and software layers, obtain fault feature vectors through fusion calculation, build device communication relationship diagram and link abnormality probability distribution diagram, perform correlation analysis to determine the root cause of the fault equipment, and generate fault processing instructions, record recovery process data and update baseline data.
It realizes accurate identification and positioning of network equipment faults, improves the accuracy and efficiency of fault detection, reduces network maintenance complexity, reduces diagnosis time, and improves the stability and reliability of the network system.
Smart Images

Figure CN120342902A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer fault diagnosis, and particularly to a method and system for identifying and locating multi-source network device fault modes. Background Art
[0002] With the continuous expansion of the network scale and the improvement of complexity, the identification and location of network device faults have become a key link to ensure the stable operation of the network. In the current network environment, there are various types of network devices, and different devices are interconnected to form a complex network topology structure; During daily operation, due to reasons such as hardware aging, software errors, network attacks, or improper configurations, network devices may encounter various faults, such as central processor overload, memory exhaustion, network interface abnormalities, system service interruptions, etc., resulting in a decline in network performance, service interruptions, or even paralysis of the entire network, which has a serious impact on the normal operation of enterprises. Therefore, quickly and accurately identifying and locating network device faults is of great significance for network management and maintenance; However, the existing network device fault identification and location technologies still lack the comprehensive analysis ability of multi-source data at the hardware layer and software layer, resulting in one-sided fault judgments and an inability to comprehensively grasp the device operation status. Especially in complex fault modes, the recognition accuracy is relatively low, and the fault correlation analysis between multiple devices cannot be effectively processed, resulting in low accuracy in identifying network-level faults. Especially when the fault link involves multiple devices, it is difficult to accurately locate the root cause of the fault, and there is also a lack of adaptability to the dynamic changes of the network environment. It is unable to automatically update the judgment baseline based on historical fault handling experience, resulting in a large number of false alarms or missed alarms when the network load fluctuates or the business model changes, reducing the reliability of fault identification and location. Therefore, there is an urgent need for a solution to solve the problems existing in the prior art. Summary of the Invention
[0003] Embodiments of the present invention provide a method and system for identifying and locating multi-source network device fault modes, which can at least solve some of the problems existing in the prior art.
[0004] In the first aspect of the embodiments of the present invention, a method for identifying and locating multi-source network device fault modes is provided, including: Collect the central processor utilization rate, memory utilization rate, network interface traffic data, system log data, and network protocol data of the target network device; Calculate the deviation values of the central processor utilization rate, memory utilization rate, and network interface traffic data from the preset baseline respectively, generate a hardware layer fault feature matrix, perform semantic segmentation on the system log data, extract the log event sequence and process call chain information, combine the protocol interaction records of the network protocol data, generate a software layer fault feature matrix, and perform fusion calculation on the hardware layer fault feature matrix and the software layer fault feature matrix to obtain the fault feature vector of the target network device; Collect the adjacent device identification information and link performance index data that communicate directly with the target network device; Construct a device communication relationship graph based on the adjacent device identification information, calculate the abnormal level of the link performance index data, map the abnormal level to the device communication relationship graph to obtain a link abnormal probability distribution graph, perform correlation analysis on the fault feature vector and the link abnormal probability distribution graph, and determine the influence range of the faulty device through recursive calculation, and output the faulty root device identification information and the fault occurrence probability; Generate a fault handling instruction according to the faulty root device identification information and the fault occurrence probability and execute it, record the fault recovery process data and update the baseline data.
[0005] In an alternative embodiment, Calculating the deviation values of the central processor utilization rate, memory utilization rate, and network interface traffic data from the preset baseline respectively, generating a hardware layer fault feature matrix, performing semantic segmentation on the system log data, extracting the log event sequence and process call chain information, and combining the protocol interaction records of the network protocol data to generate a software layer fault feature matrix includes: Perform time series segmentation on the central processor utilization rate, memory utilization rate, and network interface traffic data of the target network device, use the exponentially weighted moving average to calculate the baseline values of the central processor utilization rate, memory utilization rate, and network interface traffic data, and calculate the standardized deviation value through Z-score standardization; Perform weighted calculation on the standardized deviation value and the time decay factor to obtain the cumulative abnormal score, construct a three-dimensional feature vector based on the cumulative abnormal score, and construct a hardware layer fault feature matrix through a sliding time window; Use hierarchical clustering to classify the system log data into log templates, perform semantic segmentation on the system log data based on the dynamic programming algorithm and calculate the event sequence entropy value, extract the process identifier, process creation time, process end time, and parent-child process identifier from the system log data to construct the process call chain information, calculate the delay weighted value of the critical path in the process call chain information, analyze the protocol interaction records of the network protocol data, and calculate the protocol interaction abnormal score in combination with the violation weight and violation frequency; Align the event sequence entropy value, time delay weighted value, and protocol interaction anomaly score according to the time window and splice them to construct a software layer feature vector, and perform time series accumulation on the software layer feature vector to obtain a software layer fault feature matrix.
[0006] In an alternative embodiment, Fuse and calculate the hardware layer fault feature matrix and the software layer fault feature matrix to obtain a fault feature vector of the target network device, including: Determine the features in the hardware layer fault feature matrix that are greater than a preset hardware bottleneck threshold as hardware bottleneck features, calculate the influence weight of the hardware bottleneck features on the features in the software layer fault feature matrix, and calculate the change rate of adjacent moments of the features in the software layer fault feature matrix based on the influence weight to obtain the hardware bottleneck influence degree; Determine the features in the software layer fault feature matrix that deviate from the feature mean by more than a preset anomaly determination coefficient and feature standard deviation as software anomaly features, calculate the influence coefficient of the software anomaly features on the features in the hardware layer fault feature matrix, and calculate the change rate of the relative maximum value at the future moment of the features in the hardware layer fault feature matrix based on the influence coefficient to obtain the hardware resource exhaustion degree; Calculate the weighted sum of the hardware bottleneck influence degree and the hardware resource exhaustion degree to obtain the performance bottleneck propagation intensity, accumulate and calculate the performance bottleneck propagation intensity within a preset time window to obtain the node importance, multiply the propagation intensities between each node on the propagation path to obtain the propagation path importance, and calculate the weighted sum of the node importance and the propagation path importance to obtain the feature dynamic importance; Exponentially normalize the feature dynamic importance to obtain the propagation perception weight, obtain the fusion feature based on the product of the propagation perception weight and the corresponding features in the hardware layer fault feature matrix and the software layer fault feature matrix, and input the fusion feature into a preset activation function to obtain the fault feature vector of the target network device.
[0007] In an alternative embodiment, Collect the adjacent device identification information and link performance index data that communicate directly with the target network device, including: Obtain the adjacent device list of the target network device, and extract the corresponding adjacent device identification information of each adjacent device in the adjacent device list; Establish a communication connection with each adjacent device in the adjacent device list, and collect the link performance index data between the target network device and each adjacent device; Associate the adjacent device identification information with the link performance index data to construct the link communication feature matrix of the target network device.
[0008] In an alternative embodiment, Construct a device communication relationship graph based on the adjacent device identification information, calculate the anomaly level of the link performance index data, map the anomaly level to the device communication relationship graph to obtain a link anomaly probability distribution graph, perform correlation analysis on the fault feature vector and the link anomaly probability distribution graph, and determine the influence range of the faulty device through recursive calculation, and output the faulty root device identification information and the fault occurrence probability, including: Construct a device communication relationship graph based on the adjacent device identification information, calculate the link weights of different communication types between devices, perform segmented sliding analysis on the link performance index data, calculate the product of the mutation amplitude and the duration of the performance index, and use the deviation degree of the product from the historical baseline corresponding to the current device as the link anomaly level. Calculate the link anomaly probability based on the link anomaly level, and map the link anomaly probability to the corresponding link in the device communication relationship graph to obtain a link anomaly probability distribution graph; Extract the frequency-domain features and time-domain features of the fault feature vector through multi-scale decomposition, combine the extracted features with the link anomaly probability with weights to obtain the node local correlation degree, calculate the node importance based on the shortest path between nodes and the link weights, and obtain the global correlation degree through non-linear mapping of the product of the node local correlation degree and the node importance. Calculate the product of the cumulative attenuation effect of the link anomaly probability in the time series dimension and the node local correlation degree to obtain the propagation intensity matrix; Construct a node influence propagation chain according to the propagation intensity matrix, calculate the node state transition probability through hierarchical recursion until convergence to obtain the final node state, and perform probability fusion on the final node state and the global correlation degree to obtain the faulty root device identification information and the fault occurrence probability.
[0009] In an alternative embodiment, Construct a node influence propagation chain according to the propagation intensity matrix, calculate the node state transition probability through hierarchical recursion until convergence to obtain the final node state, and perform probability fusion on the final node state and the global correlation degree to obtain the faulty root device identification information and the fault occurrence probability, including: Construct a node influence propagation chain according to the propagation intensity matrix. The node influence propagation chain includes a node set, an edge set, and a propagation intensity weight matrix. Sample the historical state sequence of each node within a preset time series window to obtain a node state sequence. Calculate the time series dependence relationship between node state sequences based on mutual information measurement to obtain a time series dependence matrix. Perform Granger causality test on the node state sequence to obtain the direct causality intensity. Obtain all paths between nodes and calculate the cumulative product of the direct causality intensity on each path to obtain the indirect causality intensity; Calculate the weighted sum of the abnormal state duration and the decay coefficient for the node state sequence to obtain the state duration feature, calculate the average value of the state differences at adjacent times to obtain the change trend feature, calculate the periodic feature based on the autocorrelation coefficient of the state sequence, and combine the state duration feature, the change trend feature, and the periodic feature to construct a time series feature vector; Construct a state transition probability matrix according to the propagation intensity weight matrix, the indirect causality intensity, and the time series feature vector, and fuse the state transition probability matrix and the time series dependence matrix through hierarchical recursive calculation to update the node state until convergence to obtain the final node state; Perform probability fusion on the final node state and the global correlation degree to obtain the faulty root cause device identification information and the fault occurrence probability.
[0010] In an alternative implementation manner, Generate and execute a fault handling instruction according to the faulty root cause device identification information and the fault occurrence probability, record the fault recovery process data, and update the baseline data, including: According to the faulty root cause device identification information and the fault occurrence probability, match the fault handling strategy corresponding to the faulty root cause device identification information based on a preset fault type library, and perform priority sorting on the fault handling strategy according to the fault occurrence probability to generate a fault handling instruction; Execute the fault handling instruction, collect the fault recovery process data, and record the execution effect of the fault handling instruction; Update the baseline data based on the fault recovery process data, and feedback the execution effect of the fault handling instruction to the fault type library to complete the dynamic optimization of the fault handling strategy.
[0011] In the second aspect of the embodiments of the present invention, there is provided a multi-source network device fault mode recognition and localization system, including: A first unit, configured to collect the central processor usage rate, memory usage rate, network interface traffic data, system log data, and network protocol data of a target network device; A second unit, configured to calculate the deviation values of the central processor usage rate, memory usage rate, and network interface traffic data from a preset baseline respectively, generate a hardware layer fault feature matrix, perform semantic segmentation on the system log data, extract log event sequences and process call chain information, and combine the protocol interaction records of the network protocol data to generate a software layer fault feature matrix, and perform fusion calculation on the hardware layer fault feature matrix and the software layer fault feature matrix to obtain a fault feature vector of the target network device; A third unit, configured to collect the adjacent device identification information and link performance index data that directly communicate with the target network device; A fourth unit, configured to construct a device communication relationship graph based on the adjacent device identification information, calculate the anomaly level of the link performance index data, map the anomaly level to the device communication relationship graph to obtain a link anomaly probability distribution graph, perform correlation analysis on the fault feature vector and the link anomaly probability distribution graph, determine the influence scope of the faulty device through recursive calculation, and output the faulty root cause device identification information and the fault occurrence probability; A fifth unit, configured to generate and execute a fault handling instruction according to the faulty root cause device identification information and the fault occurrence probability, record the fault recovery process data, and update the baseline data.
[0012] In a third aspect of the embodiments of the present invention, there is provided an electronic device, including: A processor and a memory for storing instructions executable by the processor, wherein the processor is configured to call the instructions stored in the memory to execute the foregoing method.
[0013] In a fourth aspect of the embodiments of the present invention, there is provided a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the foregoing method is implemented.
[0014] In the present invention, through collecting multi-source network device data and combining the fusion analysis of the fault feature matrices of the hardware layer and the software layer, accurate identification and location of network device faults are realized, the accuracy and efficiency of fault detection are effectively improved, the influence scope of faults is determined through recursive calculation, the fault root cause can be effectively traced, the complexity of network maintenance is reduced, the fault diagnosis time is reduced, the fault recovery process is recorded and the baseline data is dynamically updated, the stability and reliability of the network system are improved, and an intelligent solution is provided for network operation and maintenance. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 It is a schematic flow chart of a multi-source network device fault mode identification and location method according to an embodiment of the present invention; Figure 2 It is a schematic diagram of segmented sliding analysis of link delay; Figure 3 It is a schematic diagram of a node influence propagation chain. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0016] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the drawings in the embodiments of the present invention. Apparently, the described embodiments are only some of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0017] The technical solution of the present invention will be described in detail below with specific embodiments. These several specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.
[0018] Figure 1 It is a schematic flowchart of the method for identifying and locating multi-source network device failure modes in an embodiment of the present invention. As Figure 1 shown, the method includes: Collect the central processor utilization rate, memory utilization rate, network interface traffic data, system log data, and network protocol data of the target network device; Calculate the deviation values of the central processor utilization rate, memory utilization rate, and network interface traffic data from the preset baseline respectively to generate a hardware layer failure feature matrix. Perform semantic segmentation on the system log data, extract log event sequences and process call chain information, and combine the protocol interaction records of the network protocol data to generate a software layer failure feature matrix. Perform fusion calculation on the hardware layer failure feature matrix and the software layer failure feature matrix to obtain the failure feature vector of the target network device; Collect the adjacent device identification information and link performance index data that communicate directly with the target network device; Construct a device communication relationship graph based on the adjacent device identification information, calculate the anomaly level of the link performance index data, map the anomaly level to the device communication relationship graph to obtain a link anomaly probability distribution graph, perform correlation analysis on the failure feature vector and the link anomaly probability distribution graph, and determine the influence range of the faulty device through recursive calculation, and output the faulty root device identification information and the probability of failure occurrence; Generate a fault handling instruction according to the faulty root device identification information and the probability of failure occurrence and execute it, record the fault recovery process data and update the baseline data.
[0019] In an alternative embodiment, Calculating the deviation values of the central processor utilization rate, memory utilization rate, and network interface traffic data from the preset baseline respectively to generate a hardware layer failure feature matrix, performing semantic segmentation on the system log data, extracting log event sequences and process call chain information, and combining the protocol interaction records of the network protocol data to generate a software layer failure feature matrix includes: Perform time series segmentation on the central processor utilization rate, memory utilization rate, and network interface traffic data of the target network device, calculate the baseline values of the central processor utilization rate, memory utilization rate, and network interface traffic data using exponentially weighted moving average, and calculate the standardized deviation values through Z-score standardization; The standardized deviation value is weighted with the time decay factor to obtain the cumulative anomaly score. A three-dimensional feature vector is constructed based on the cumulative anomaly score, and a hardware layer fault feature matrix is constructed by sliding a time window. Hierarchical clustering is used to classify the system log data into log templates. Based on the dynamic programming algorithm, the system log data is semantically segmented and the event sequence entropy value is calculated. The process identifier, process creation time, process end time, and parent-child process identifier are extracted from the system log data to construct the process call chain information. The time-delay weighted value of the critical path in the process call chain information is calculated. The protocol interaction records of the network protocol data are analyzed, and the protocol interaction anomaly score is calculated by combining the violation weight and the violation frequency. The event sequence entropy value, the time-delay weighted value, and the protocol interaction anomaly score are aligned according to the time window and then spliced to construct the software layer feature vector. The software layer feature vector is cumulated in time series to obtain the software layer fault feature matrix.
[0020] For the construction of the hardware layer fault feature matrix, the time series of the central processor usage rate, memory usage rate, and network interface traffic data of the target network device are segmented. Taking a certain network routing device as an example, the performance data of the device every 5 minutes within 24 consecutive hours are collected, including the CPU usage rate, memory usage rate, and interface traffic data. For the CPU usage rate, the exponential weighted moving average algorithm is used to calculate the baseline value. A smoothing factor of 0.2 is selected to calculate the historical data of the most recent 12 hours. For example, if the current CPU usage rate is 85% and the calculated baseline value is 45%, through the Z-score standardization method, considering the standard deviation of the historical data is 10%, the standardized deviation value is (85% - 45%) / 10% = 4. Similarly, the standardized deviation values of the memory usage rate and network traffic data are calculated. For example, the standardized deviation value of the memory usage rate is 2.5, and the standardized deviation value of the network traffic is 3.2.
[0021] The standardized deviation value is weighted with the time decay factor to obtain the cumulative anomaly score. The time decay factor is set to 0.9. For the current moment t and the previous moment t - 1, if the cumulative anomaly score at the moment t - 1 is 3.8, then the cumulative anomaly score at the moment t is 0.9×3.8 + 4 = 7.42. The cumulative anomaly scores of the CPU usage rate, memory usage rate, and network traffic data are calculated respectively to construct a three-dimensional feature vector (7.42, 5.35, 6.88). The sliding window size is selected to be 30 minutes, and it slides 5 minutes each time. Then each sliding window contains the three-dimensional feature vectors of 6 time points, forming a 6×3 hardware layer fault feature matrix.
[0022] For the construction of the software layer fault feature matrix, hierarchical clustering is used to classify log templates for system log data. Taking the system logs of a certain network switch as an example, a data set containing 10,000 log records is collected. By calculating the edit distance of log texts and setting the threshold to 0.7, log entries with a similarity greater than 0.7 are grouped into the same class, and finally 120 log template classes are formed.
[0023] Semantic segmentation is performed on the system log data based on the dynamic programming algorithm. The log sequences are grouped according to the timestamps, and the time interval threshold is set to 5 seconds. When the time interval between adjacent logs exceeds the threshold, a new segment is created. By counting the frequencies of each log template in each segment, the entropy value of the event sequence is calculated. For example, if a segment contains 10 logs belonging to 3 different templates with frequencies of 50%, 30%, and 20% respectively, the entropy value of this segment is -(0.5×log20.5 + 0.3×log20.3 + 0.2×log20.2) = 1.49.
[0024] Extract process information from the system logs, including process identifiers (such as PID = 1234), process creation times (such as 2023 - 06 - 01 08:00:15), process end times (such as 2023 - 06 - 01 08:10:22), and parent - child process identifiers (such as PPID = 1000). Construct process call chain information, draw a process call graph, and calculate the weighted delay value of the critical path. For example, the longest path from the root process to the leaf process contains 5 process nodes, and the execution times of each node are 10 seconds, 15 seconds, 5 seconds, 20 seconds, and 8 seconds respectively. Setting the weight coefficients as (0.1, 0.2, 0.3, 0.3, 0.1), the weighted delay value of this path is 10×0.1 + 15×0.2 + 5×0.3 + 20×0.3 + 8×0.1 = 11.8.
[0025] Analyze the interaction records of network protocol data to identify protocol violations. For example, during the TCP connection establishment process, the violation weights for the normal SYN, SYN - ACK, ACK handshake sequences are set to 0, while the violation weight for the SYN - SYN or SYN - RST sequences is 0.7. Count the frequencies of specific violation sequences within a 30 - minute window. If SYN - RST appears 12 times, the protocol interaction anomaly score for this window is 0.7×12 = 8.4.
[0026] Align the event sequence entropy value, time delay weighted value, and protocol interaction anomaly score according to the time window. For example, the three metrics for a 30-minute window are 1.49, 11.8, and 8.4 respectively. Concatenate and construct the software layer feature vector (1.49, 11.8, 8.4). Perform time series accumulation on the software layer feature vector, set the accumulation coefficient to 0.8. If the software layer feature vector at the previous moment is (1.2, 10.5, 7.8), then the accumulation result at the current moment is (0.8×1.2 + 1.49, 0.8×10.5 + 11.8, 0.8×7.8 + 8.4) = (2.45, 20.2, 14.64). Perform cumulative calculation on the software layer feature vectors of multiple consecutive time windows to form the software layer fault feature matrix.
[0027] In this embodiment, by performing time series segmentation and baseline calculation on the hardware metrics, and constructing the cumulative anomaly score in combination with the time decay factor, it can effectively capture the performance anomaly fluctuations at the hardware level, avoid the limitations of the traditional fixed threshold method, and achieve semantic segmentation through dynamic programming, effectively solving the problems of diverse log formats and complex semantics. The introduction of the event sequence entropy value can quantify the uncertainty of system behavior, providing an important basis for anomaly detection. By constructing the process call chain and extracting the key path time delay, and combining with the protocol interaction anomaly analysis, a multi-dimensional characterization of software layer anomalies is achieved.
[0028] In an alternative embodiment, Perform fusion calculation on the hardware layer fault feature matrix and the software layer fault feature matrix to obtain the fault feature vector of the target network device, including: Determine the features in the hardware layer fault feature matrix that are greater than the preset hardware bottleneck threshold as the hardware bottleneck features, calculate the influence weight of the hardware bottleneck features on the features in the software layer fault feature matrix, and calculate the change rate of adjacent moments of the features in the software layer fault feature matrix based on the influence weight to obtain the hardware bottleneck influence degree; Determine the features in the software layer fault feature matrix that deviate from the feature mean by more than the preset anomaly determination coefficient times the feature standard deviation as the software anomaly features, calculate the influence coefficient of the software anomaly features on the features in the hardware layer fault feature matrix, and calculate the change rate of the relative maximum value of the features in the hardware layer fault feature matrix at the future moment based on the influence coefficient to obtain the degree of hardware resource exhaustion; Calculate the weighted sum of the hardware bottleneck influence degree and the degree of hardware resource exhaustion to obtain the performance bottleneck propagation intensity. Accumulate and calculate the performance bottleneck propagation intensity within the preset time window to obtain the node importance. Multiply the propagation intensities between each node on the propagation path to obtain the propagation path importance. Calculate the weighted sum of the node importance and the propagation path importance to obtain the feature dynamic importance; The dynamic importance of features is normalized exponentially to obtain the propagation perception weight. Based on the product of the propagation perception weight and the corresponding features in the hardware layer fault feature matrix and the software layer fault feature matrix, the fused features are obtained. The fused features are input into a preset activation function to obtain the fault feature vector of the target network device.
[0029] Judge the hardware bottleneck features in the hardware layer fault feature matrix. For each eigenvalue in the hardware layer fault feature matrix, compare it with a preset hardware bottleneck threshold. For example, if the CPU usage rate feature is 95% and the preset hardware bottleneck threshold is 90%, then this CPU usage rate feature is determined as a hardware bottleneck feature. If the memory usage rate is 88%, which is lower than the preset hardware bottleneck threshold of 90%, then this memory usage rate feature is not determined as a hardware bottleneck feature.
[0030] For the determined hardware bottleneck features, calculate the influence weights on the features in the software layer fault feature matrix. Suppose the hardware bottleneck feature is the CPU usage rate of 95%. The influence weight on the service response time feature in the corresponding software layer can be obtained by analyzing historical data. For example, the influence weight is 0.7. Based on this influence weight, calculate the change rate of adjacent moments of the features in the software layer fault feature matrix. Exemplarily, if the service response time is 200 ms at the current moment and 150 ms at the previous moment, then the change rate of adjacent moments is (200 - 150) / 150 = 33.3%. Multiply the change rate by the influence weight to obtain the hardware bottleneck influence degree of 0.7×33.3% = 23.31%.
[0031] Identify software anomaly features from the software layer fault feature matrix, and calculate the mean and standard deviation of each software layer feature. For example, the mean of the service response time is 100 ms, the standard deviation is 20 ms, the preset anomaly determination coefficient is 2.5, and the current service response time is 200 ms, deviating from the mean by 100 ms, exceeding 2.5×20 ms = 50 ms, so it is determined as a software anomaly feature.
[0032] For the determined software anomaly features, calculate the influence coefficients on the features in the hardware layer fault feature matrix. Suppose the influence coefficient of the service response time anomaly on the CPU usage rate is 0.6. Based on this influence coefficient, calculate the change rate of the relative maximum value of the features in the hardware layer fault feature matrix at the future moment. For example, the current CPU usage rate is 95%, the historical maximum value is 98%, and it is predicted that the CPU usage rate will increase to 97% at the future moment. Then the change rate of the relative maximum value at the future moment is (97% - 95%) / (98% - 95%) = 66.7%. Multiply the change rate by the influence coefficient to obtain the degree of hardware resource exhaustion of 0.6×66.7% = 40.02%.
[0033] The weighted sum of the degree of impact of hardware bottlenecks and the degree of exhaustion of hardware resources is calculated to obtain the performance bottleneck propagation intensity. Assuming that the weight of the degree of impact of hardware bottlenecks is 0.4 and the weight of the degree of exhaustion of hardware resources is 0.6, the performance bottleneck propagation intensity is 0.4×23.31% + 0.6×40.02% = 33.25%.
[0034] Within a preset time window (e.g., the past 1 hour), the performance bottleneck propagation intensity is cumulatively calculated to obtain the node importance. Assuming that the average value of the performance bottleneck propagation intensity at 10 sampling points within 1 hour is 30%, the node importance is 30%.
[0035] When calculating the importance of the propagation path, the product of the propagation intensities between nodes on the propagation path is considered. For example, in the propagation path from the CPU to the memory and then to the service response time, the propagation intensity from the CPU to the memory is 25% and the propagation intensity from the memory to the service response time is 35%, then the importance of this propagation path is 25%×35% = 8.75%.
[0036] The weighted sum of the node importance and the propagation path importance is calculated to obtain the feature dynamic importance. Assuming that the weight of the node importance is 0.7 and the weight of the propagation path importance is 0.3, the feature dynamic importance is 0.7×30% + 0.3×8.75% = 23.63%.
[0037] The feature dynamic importance is exponentiated and normalized to obtain the propagation perception weight. For example, if the dynamic importances of three features are 23.63%, 15.42%, and 32.18% respectively, after exponentiated normalization, the propagation perception weights are 0.35, 0.25, and 0.4 respectively.
[0038] Based on the product of the propagation perception weight and the corresponding features in the hardware layer fault feature matrix and the software layer fault feature matrix, the fused features are obtained. For example, for the CPU usage rate feature, if its original value is 95% and the propagation perception weight is 0.35, the fused feature value is 95%×0.35 = 33.25%.
[0039] All the fused features are input into a preset activation function, such as the Sigmoid function, to obtain the fault feature vector of the target network device. For example, if the fused feature value is 33.25%, after being processed by the Sigmoid function, it may obtain 0.67, representing the value of this dimension in the fault feature vector.
[0040] In this embodiment, an impact analysis mechanism of hardware bottleneck features on software layer features is established. By calculating the change rate at adjacent times, the impact degree of hardware resource limitation on software performance is accurately quantified. By introducing the concept of performance bottleneck propagation intensity, the node importance and propagation path importance are analyzed within a preset time window, achieving a comprehensive characterization of the fault propagation characteristics. A propagation-aware feature fusion strategy is adopted to adaptively fuse the features of the hardware layer and the software layer; In the prior art, fault feature extraction usually adopts a single-dimensional threshold judgment method, which analyzes only from the hardware or software level alone and cannot effectively identify complex faults caused by hardware-software interaction. When performing feature fusion, simple weighted superposition is often used, without considering the propagation characteristics of performance bottlenecks in the system, resulting in inaccurate feature expression and affecting the effect of fault location; The feature importance evaluation method considering the propagation effect in this embodiment makes up for the deficiency of only focusing on static features in the prior art. Through the design of propagation-aware weights, the fused features can better reflect the fault propagation law in the system, improving the expression ability and discrimination of the features, analyzing the consumption trend of software abnormal features on hardware resources, realizing the predictive evaluation of the risk of hardware resource exhaustion, breaking through the limitation of traditional one-way analysis, providing more reliable feature support for fault root cause location, and at the same time improving the timeliness of fault warning and enhancing the accuracy of fault location.
[0041] In an alternative embodiment, Collecting the adjacent device identification information and link performance index data directly communicating with the target network device includes: Obtain the adjacent device list of the target network device, and extract the corresponding adjacent device identification information of each adjacent device in the adjacent device list; Establish a communication connection with each adjacent device in the adjacent device list, and collect the link performance index data between the target network device and each adjacent device; Associate the adjacent device identification information with the link performance index data to construct the link communication feature matrix of the target network device.
[0042] Establish a communication connection with the target network device through the management interface. The management interface can be implemented based on network management protocols such as SSH, SNMP, or NETCONF. Send a query command to the target network device to obtain the adjacent device list. For example, for a network device supporting the SNMP protocol, query the LLDP-MIB or CDP-MIB by sending a GetRequest message to obtain the adjacent device information; for a device supporting NETCONF, send a get-config request containing an xpath filter to obtain the adjacent device information.
[0043] After the target network device receives the query request, it returns a response message containing the information of adjacent devices. Parse the list of adjacent devices from the response message. The list of adjacent devices contains all network devices directly physically connected to the target device. Process the obtained list of adjacent devices and extract the identification information of each adjacent device. The identification information of adjacent devices includes, but is not limited to: device name, device type, MAC address, management IP address, model, manufacturer, operating system version, interface identifier, etc. For example, the identification information of adjacent devices obtained from device A may include the following data of device B: the device name is "CoreSwitch01", the device type is "switch", the MAC address is "00:1A:2B:3C:4D:5E", the management IP address is "192.168.1.10", the device model is "MS-7500", the manufacturer is "Network Equipment Manufacturer", the operating system version is "NOS12.2", and the interface identifier connected to device A is "GigabitEthernet1 / 0 / 24". The system stores the extracted identification information of adjacent devices in a temporary data structure, such as a hash table or an array, for subsequent processing.
[0044] After extracting the identification information of adjacent devices, establish communication connections with each device in the list of adjacent devices. According to the extracted management IP addresses, sequentially attempt to establish management connections with each adjacent device using protocols such as SSH, SNMP, or NETCONF, and verify the connection credentials, such as username, password, or key, to ensure that a management session can be successfully established. For adjacent devices with connection failures, record the error information and continue to process the next device.
[0045] After establishing communication connections with adjacent devices, collect the link performance metric data between the target network device and each adjacent device. The link performance metric data includes, but is not limited to, link utilization rate, link bandwidth, packet transmission delay, packet loss rate, packet error rate, packet retransmission rate, traffic throughput, number of interface status changes. Calculate the performance metrics by periodically querying the interface statistic data of the target device and adjacent devices. For example, collect the interface counter data every 30 seconds, calculate the difference between the two collections, and then calculate the performance metric values in combination with the time interval. Specifically, query the ifInOctets and ifOutOctets counters through SNMP to calculate the link utilization rate; query the ifInErrors and ifOutErrors counters to calculate the packet error rate, and obtain the packet transmission delay data through active measurement or passive listening.
[0046] Taking the link utilization calculation as an example, the system collects the ifInOctets and ifOutOctets counter values of the target device interface at time points T1 and T2 (with an interval of 30 seconds). Assume that at time T1, ifInOctets is 1,000,000 bytes and ifOutOctets is 2,000,000 bytes; at time T2, ifInOctets is 1,500,000 bytes and ifOutOctets is 2,800,000 bytes. The interface bandwidth is 1 Gbps (i.e., 125,000,000 bytes / second). The inbound link utilization is calculated as (1,500,000 - 1,000,000) / (125,000,000 × 30) × 100% = 0.4%, and the outbound link utilization is calculated as (2,800,000 - 2,000,000) / (125,000,000 * 30) × 100% = 0.64%. All the collected link performance metric data is stored in a time series database for subsequent analysis and correlation processing.
[0047] After obtaining the adjacent device identification information and link performance metric data, the two are correlated to construct a link communication feature matrix for the target network device. During the correlation process, according to the physical connection relationship between the target device and the adjacent devices, the link performance metric data is mapped to the corresponding adjacent devices. Create a two-dimensional data structure where the rows represent the various adjacent devices connected to the target device, and the columns represent different link performance metrics and device identification information. For example, for the connections of the target device with three adjacent devices (Device B, Device C, and Device D), the link communication feature matrix includes device identification information (such as device name, type, IP address, etc.) and link performance metrics (such as link utilization, latency, error rate, etc.). The specific data may be as follows: the link utilization of Device B is 0.4%, the latency is 2.5 milliseconds, and the error rate is 0.001%; the link utilization of Device C is 1.2%, the latency is 1.8 milliseconds, and the error rate is 0%; the link utilization of Device D is 0.8%, the latency is 3.0 milliseconds, and the error rate is 0.002%.
[0048] In this embodiment, by obtaining the list of adjacent devices of the target network device, a complete device communication topology view is established, breaking through the limitations of traditional single-link analysis, providing a global perspective for subsequent link performance analysis. Based on the data collection method of actual communication connections, compared with traditional passive monitoring methods, it can more accurately reflect the real-time performance status of the link, improving the timeliness and reliability of the data. By performing correlation analysis on the adjacent device identification information and link performance metric data, a link communication feature matrix of the target network device is constructed, enabling link performance analysis to retain both topological structure information and performance status information, providing more complete feature support for subsequent fault location.
[0049] In an alternative implementation, Construct a device communication relationship graph based on the adjacent device identification information, calculate the anomaly level of the link performance index data, map the anomaly level to the device communication relationship graph to obtain a link anomaly probability distribution graph, perform correlation analysis on the fault feature vector and the link anomaly probability distribution graph, and determine the influence scope of the faulty device through recursive calculation, and output the faulty root device identification information and the fault occurrence probability, including: Construct a device communication relationship graph based on the adjacent device identification information, calculate the link weights of different communication types between devices, perform segmented sliding analysis on the link performance index data, calculate the product of the mutation amplitude and the duration of the performance index, and use the deviation degree of the product from the historical baseline corresponding to the current device as the link anomaly level, calculate the link anomaly probability based on the link anomaly level, and map the link anomaly probability to the corresponding link in the device communication relationship graph to obtain a link anomaly probability distribution graph; Extract the frequency domain features and time domain features of the fault feature vector through multi-scale decomposition, combine the extracted features with the link anomaly probability with weights to obtain the node local correlation degree, calculate the node importance based on the shortest path between nodes and the link weights, and obtain the global correlation degree through non-linear mapping of the product of the node local correlation degree and the node importance, and calculate the product of the cumulative attenuation effect of the link anomaly probability in the time series dimension and the node local correlation degree to obtain the propagation intensity matrix; Construct a node influence propagation chain according to the propagation intensity matrix, calculate the node state transition probability through hierarchical recursion until convergence to obtain the final node state, and perform probability fusion on the final node state and the global correlation degree to obtain the faulty root device identification information and the fault occurrence probability.
[0050] Collect the identification information of all devices in the network and their interconnection relationships. For example, in a network containing 100 devices, device A is connected to devices B, C, and D, and device B is connected to devices E and F. Represent the connection relationship as a graph structure, where nodes represent devices and edges represent communication links between devices. For each link, assign different weights according to the communication type, such as TCP connection weight is 0.8, UDP connection weight is 0.6, and ICMP connection weight is 0.4; Perform segmented sliding analysis on the link performance metric data, with a window size of 30 seconds and a slide of 10 seconds each time. Within each window, the system calculates the changes in performance metrics such as latency, packet loss rate, and throughput. When a metric mutation is detected, calculate the product of the mutation amplitude and the duration. For example, if the link latency between device A and device B suddenly increases from 5 milliseconds to 50 milliseconds and lasts for 20 seconds, the mutation amplitude is 9 times, and the product is 9×20 = 180. Compare this product with the historical baseline value of the link, which is calculated as the average of data in the same time period over the past 30 days. If the historical baseline value is 30, the current deviation degree is 180 / 30 = 6, and 6 is used as the anomaly level of this link.
[0051] Calculate the link anomaly probability based on the link anomaly level, using normalization to map the anomaly level to a probability value between 0 and 1. Exemplarily, if the highest anomaly level in the network is 10 and the lowest is 0, the link anomaly probability for an anomaly level of 6 is 6 / 10 = 0.6. Map the anomaly probabilities of all links onto the device communication relationship graph to form a link anomaly probability distribution diagram, where the anomaly probability is visually displayed through the shade of color or the thickness of the line.
[0052] Extract the frequency domain features and time domain features of the fault feature vector, and perform multi-scale decomposition on the collected fault logs and alarm information. The time domain features include the fault occurrence time, duration, repetition frequency, etc.; the frequency domain features are obtained by transforming the time series data and reflect the periodic characteristics of the data. For example, for a periodically occurring network latency increase phenomenon, its main frequency is extracted as 10 am every day and lasts for about 30 minutes.
[0053] Combine the extracted features with the link anomaly probability with weights to obtain the node local relevance. Assign weights to each dimension of the feature vector, such as a time matching degree weight of 0.4, a fault type matching degree weight of 0.3, and a performance metric correlation weight of 0.3. Assume that the time matching degree of device A is 0.8, the type matching degree is 0.7, and the performance metric correlation is 0.9, then its local relevance is 0.8×0.4 + 0.7×0.3 + 0.9×0.3 = 0.8.
[0054] Calculate the node importance based on the shortest path between nodes and the link weights, considering the centrality of the nodes. For each node, calculate the shortest path length to all other nodes and consider the link weights on the path. The higher the node importance value, the greater the influence range of the node in the network. For example, after calculation, the node importance of device A is 0.75.
[0055] Map the product of the node local relevance and the node importance through a non-linear mapping to obtain the global relevance, using a logarithmic function for the mapping to ensure the result is within a reasonable range. For example, the global relevance of device A is the value 0.62 after logarithmic function processing.
[0056] When calculating the product of the cumulative decay effect of the link anomaly probability in the time series dimension and the node local correlation degree to obtain the propagation intensity matrix, the decay of the abnormal state over time is considered. Assuming the decay factor is 0.9, the anomaly probability at time t + 1 is 0.9 times that at time t. Multiply the decayed anomaly probability by the node local correlation degree to obtain a matrix representing the propagation intensity of the abnormal state.
[0057] Construct a node influence propagation chain based on the propagation intensity matrix to identify the propagation path of the abnormal state. Starting from each possible fault source node, determine the propagation direction and intensity of the abnormal state according to the propagation intensity to form a tree-like propagation structure.
[0058] Calculate the node state transition probability recursively by layer until convergence. Starting from the leaf nodes, calculate the state of each node layer by layer from bottom to top. In the initial state, the fault probability of all nodes is set to 0. In each iteration, the new state of a node is jointly determined by its own state and the influence transmitted from all adjacent nodes through the propagation chain. When the state change between two consecutive iterations is less than the threshold of 0.01, the calculation is considered to have converged.
[0059] Perform probability fusion on the converged node state and the global correlation degree. Using the Bayesian fusion method, calculate the posterior probability of each node as the root cause of the fault. Output the device with the highest fault probability as the root cause device of the fault, and at the same time provide the identification information and the probability of the fault occurrence of this device. For example, the system may output "Device A (ID: 10086) is the root cause device of the fault, and the probability of the fault occurrence is 0.85".
[0060] In this embodiment, by constructing a device communication relationship graph and calculating link weights of different types, an accurate expression of the network topology structure is achieved. Combining the mutation characteristics of the performance indicators with the deviation degree from the historical baseline forms a quantitative evaluation mechanism for the link anomaly level, making the identification of the link abnormal state more accurate and objective. The node importance calculated based on the shortest path between nodes and the link weights reflects the status and influence of the nodes in the entire network. The multi-scale decomposition method is used to extract the frequency-domain and time-domain features of the fault feature vector respectively, fully capturing the manifestations of the fault phenomenon in different dimensions. By calculating the time series cumulative decay effect of the link anomaly probability and constructing a propagation intensity matrix in combination with the node local correlation degree, the propagation law of the fault in the network is accurately described.
[0061] Figure 2Schematic diagram of segmented sliding analysis of link delay, which shows the results of segmented sliding analysis of the link delay between devices A and B. In the figure, the blue curve represents the changing trend of link delay over time, and the red dashed line represents the historical baseline value (delay level under normal conditions). By adopting an analysis method with a 30 - second window and a 10 - second slide each time, the system successfully captured a significant delay mutation starting at approximately 120 seconds.
[0062] In the mutation interval (marked in the light red area), the link delay increased sharply from the normal approximately 5 milliseconds to nearly 50 milliseconds, with a mutation amplitude reaching 9 times, and lasted for about 20 seconds. The ratio of the product of the mutation amplitude and the duration (180) to the historical baseline value (30) was calculated to obtain a link anomaly level of 6. This anomaly level can be converted into a link anomaly probability after normalization for subsequent root - cause failure analysis. This method of analyzing performance metrics based on segmented sliding windows can effectively identify mutation phenomena in network performance and provide a key basis for fault location.
[0063] In an alternative implementation Construct a node influence propagation chain according to the propagation intensity matrix. By recursively calculating the node state transition probability layer - by - layer until convergence to obtain the final node state, and probabilistically fusing the final node state with the global relevance, the fault - root device identification information and the fault occurrence probability are obtained, including: Construct a node influence propagation chain according to the propagation intensity matrix. The node influence propagation chain includes a node set, an edge set, and a propagation intensity weight matrix. Sample the historical state sequence of each node within a preset time series window to obtain a node state sequence. Calculate the temporal dependence relationship between node state sequences based on mutual information measure to obtain a temporal dependence matrix. Conduct a Granger causality test on the node state sequence to obtain the direct causality strength. Obtain all paths between nodes and calculate the cumulative product of the direct causality strengths on each path to obtain the indirect causality strength; Calculate the weighted sum of the abnormal state duration and the attenuation coefficient for the node state sequence to obtain the state - duration feature. Calculate the average value of the state differences at adjacent times to obtain the change - trend feature. Calculate the periodicity feature based on the autocorrelation coefficient of the state sequence. Combine the state - duration feature, the change - trend feature, and the periodicity feature to construct a temporal feature vector; Construct a state transition probability matrix according to the propagation intensity weight matrix, the indirect causality strength, and the temporal feature vector. Through layer - by - layer recursive calculation, fuse the state transition probability matrix with the temporal dependence matrix to update the node state until convergence to obtain the final node state; Probabilistically fuse the final node state with the global relevance to obtain the fault - root device identification information and the fault occurrence probability.
[0064] Collect the historical state data of each device node in the network to form a set of state sequences. For a certain telecommunications data center environment, the state data of 100 network device nodes are collected. Each node contains 10 state indicators such as CPU usage rate, memory occupancy, and network traffic. The sampling period is 5 minutes, and a total of 1440 data points (covering 5 days) are collected.
[0065] Based on the collected data, construct a node influence propagation chain. Represent all device nodes with the node set V = {v1, v2,..., v100}, and the edge set E represents the connection relationship between nodes. Normalize the original state data to the interval [0, 1] through preprocessing, and set the time series window size to 12 (i.e., 1 hour). Extract the state sequences for the data within each time window. For example, the state sequence of node v1 at time t is represented as s1(t) = {s1(t - 11), s1(t - 10),..., s1(t)}.
[0066] Calculate the temporal dependence relationship between nodes, and select state sequences for mutual information metric analysis. For the node pair (vi, vj), statistically calculate the joint probability distribution and marginal probability distribution of the corresponding state sequences of each node under different value combinations, and calculate the mutual information value. For example, the calculated mutual information value of nodes v23 and v45 is 0.73, indicating that there is a strong correlation between their state changes. After completing the mutual information calculation for all node pairs, form a temporal dependence matrix D, where D[i][j] represents the dependence strength from node vi to vj.
[0067] Conduct Granger causality test to determine the direct causal relationship. For each pair of nodes (vi, vj), use the lag value to analyze whether the historical state of vi helps predict the future state of vj. Exemplarily, set the lag order to 3 and the significance level to 0.05 in the experiment. The test results form a direct causal strength matrix C. For example, C
[38]
[42] = 0.81 indicates that node v38 has a strong direct causal influence on node v42.
[0068] Based on the network topology and direct causal strength, calculate the indirect causal strength between nodes. Adopt the depth - first search algorithm to find all paths from node vi to vj. Assume the path p = (vi, vk,..., vm, vj), then the indirect causal strength of path p is calculated as the cumulative product of the direct causal strengths along the path. For example, the indirect causal strength of a certain path p = (v12, v25, v56, v87) from node v12 to v87 is C
[12]
[25] × C
[25]
[56] × C
[56]
[87] = 0.65 × 0.72 × 0.58 = 0.27. Take the maximum indirect causal strength among all paths as the indirect causal strength between nodes and record it in the matrix IC.
[0069] Extract temporal features from the node state sequence. When calculating the state duration feature, set the decay coefficient α = 0.9 and perform a weighted sum of the durations of abnormal states. For example, if a certain abnormal state of node v33 lasts for 4 time units, its state duration feature is calculated as 1 + 0.9 + 0.81 + 0.729 = 3.439. When calculating the change trend feature, take the average of the differences between adjacent moments in the state sequence. For example, the change trend feature of node v50 is 0.23, indicating that the state is on an upward trend on average. The periodic feature is determined by the peak positions of the autocorrelation function at different lag values. For example, the autocorrelation coefficient of node v77 has a peak of 0.85 at lag value 6, indicating that there is a periodicity of about 30 minutes in its state changes. Combine these three types of features to form the temporal feature vector F for each node.
[0070] Construct the state transition probability matrix P, where P[i][j] represents the probability that the state change of node vi causes the state transition of node vj. Combine the propagation intensity weight matrix W (determined based on the network topology), the indirect causal intensity matrix IC, and the temporal feature vector F, and comprehensively calculate to obtain P[i][j]=W[i][j]×IC[i][j]×(F[j][1]+F[j][2]×F[j][3]). For example, P
[15]
[28] =0.6×0.43×(3.2+0.15×0.73)=0.86, indicating that there is an 86% probability that the state change of node v15 causes the state transition of node v28.
[0071] Implement node state update through hierarchical recursive calculation, and set the initial node state S0 according to the abnormal alarm information. In each iteration, the node state is updated by fusing the state transition probability matrix P and the temporal dependence matrix D: S(t + 1)=S(t)×(P⊗D), where ⊗ represents normalization after multiplying the corresponding elements of the matrices. Iteratively calculate until the state change is less than the threshold 0.001 or the maximum number of iterations 50 is reached, and obtain the final node state Sf.
[0072] Perform probability fusion on the final node state Sf and the pre-calculated global correlation matrix G to obtain the probability of each node being the root cause of the failure. The global correlation is calculated based on the network topology and historical failure modes. The fusion formula is R = Sf×G, where R[i] represents the probability that node vi is the root cause of the failure. In a certain experiment, the final state of node v42 is 0.92, the global correlation is 0.85, and the probability of the root cause of the failure after fusion is 0.78, and it is successfully identified as the root cause device of the failure.
[0073] In this embodiment, by sampling and analyzing the historical state of nodes within a preset time sequence window, combining mutual information measurement and Granger causality test, not only the time sequence dependence relationship between nodes is captured, but also the direct and indirect causal influence intensities are quantified, providing a reliable data basis for the identification of fault propagation paths. By introducing an attenuation coefficient to calculate the state persistence feature, both the persistent influence of abnormal states is retained and the over-accumulation of historical data is avoided. The change trend feature obtained by analyzing the difference in states at adjacent moments can capture the development trend of faults in a timely manner. By adopting a hierarchical recursive calculation method, the state transition probability and the time sequence dependence relationship are dynamically fused, realizing the accurate update of node states and ensuring the accuracy of the analysis of the fault propagation process.
[0074] Figure 3 It is a schematic diagram of the node influence propagation chain. The node v42, as the core root node (red), is located at the center and is directly connected to four first-level related nodes (orange: v38, v45, v23, v56). These first-level nodes are further connected to eight second-level related nodes (yellow). The dashed lines in the figure indicate weak connection relationships between nodes. Although these connections do not form the main propagation paths, they may form alternative channels for fault propagation in a complex network environment. Overall, this network diagram reveals the fault propagation structure centered on v42, providing a visual basis for fault root cause location and influence scope assessment, and helping to quickly identify key influencing nodes and formulate precise fault response strategies.
[0075] In an alternative embodiment, Generating and executing a fault handling instruction based on the fault root cause device identification information and the fault occurrence probability, and recording the fault recovery process data and updating the baseline data includes: Based on the fault root cause device identification information and the fault occurrence probability, matching the fault handling strategy corresponding to the fault root cause device identification information based on a preset fault type library, and sorting the fault handling strategy according to the fault occurrence probability to generate a fault handling instruction; Executing the fault handling instruction, collecting the fault recovery process data, and recording the execution effect of the fault handling instruction; Updating the baseline data based on the fault recovery process data, and feeding back the execution effect of the fault handling instruction to the fault type library to complete the dynamic optimization of the fault handling strategy.
[0076] After receiving the device monitoring alarm, obtain the device identification information of the root cause of the fault and the probability of the fault occurrence. For example, in a server room environment, the device identification information can be "Storage Server - S001", the fault type is "Disk read / write error", and the probability of the fault occurrence is 85%. Match this information with a preset fault type library, which contains various device types, their possible fault modes, and corresponding handling strategies. For the fault type of "Storage Server - Disk read / write error", the system may match three handling strategies: disk restart recovery, disk data migration, and disk replacement.
[0077] Rank these handling strategies according to the probability of the fault occurrence. Since the probability of the fault occurrence reaches 85%, it is determined as a high-probability fault. Therefore, a more reliable handling method is preferred. By querying the historical handling records, it is found that in similar high-probability cases, the success rate of disk replacement is 95%, the success rate of disk data migration is 85%, and the success rate of disk restart recovery is only 60%. Therefore, generate a sequence of handling instructions in the order of disk replacement, disk data migration, and disk restart recovery.
[0078] The generated fault handling instructions contain detailed operation steps. For example, the disk replacement instruction includes specific operation steps such as stopping the relevant services, confirming the current data backup status, pulling out the faulty disk, inserting a new disk, starting the data recovery program, verifying the data integrity, and restarting the relevant services. Send these instructions to the on-site technicians through the operation and maintenance terminal, or directly trigger the execution in an automated environment.
[0079] During the execution of the fault handling instructions, continuously collect key data points to form data on the fault recovery process, record the changes in performance metrics before and after disk replacement, including the read / write speed increasing from the original 50MB / s to 120MB / s, the error rate decreasing from 10 times per hour to 0 times, and the service response time decreasing from 200ms to 80ms, etc. Record the time node data during the handling process, such as the instruction issuance time is 10:15:30, the start execution time is 10:16:05, the completion execution time is 10:25:40, and the service recovery normal time is 10:26:15, etc.
[0080] Record the execution effect of the fault handling instructions, including information such as the execution result (success / failure), the degree of service recovery (complete recovery / partial recovery), and whether secondary faults occur. After disk replacement, the system records the execution result as "success", the degree of service recovery as "complete recovery", no secondary faults occur, and the overall score is 95 points (out of 100).
[0081] Update the baseline data based on the collected data during the fault recovery process. The baseline data includes a set of reference metrics for the normal operation of the device. For example, the normal read and write speed of the storage server should be no less than 100 MB / s, the error rate should be no higher than 1 time per hour, and the service response time should be no higher than 100 ms, etc. During the update process, comprehensively consider the historical baseline data and the actual performance data after this fault recovery, and obtain the new baseline value through the method of weighted average. For example, the updated baseline of the read and write speed is adjusted from the original 90 MB / s to 95 MB / s, reflecting the improvement of the overall system performance after replacing the new disk.
[0082] Feed back the execution effect of the fault handling instruction to the fault type library to complete the dynamic optimization of the fault handling strategy. Since the execution effect of the disk replacement strategy is good (score 95 points), increase the weight of this strategy for this type of fault from the original 0.75 to 0.80, and the weights of other strategies will be adjusted accordingly to ensure that the sum of all strategy weights is 1.
[0083] Update the correlation data between fault characteristics and fault types. For example, by analyzing the characteristic data of this fault (such as precursor signals such as abnormal temperature and power consumption fluctuations before disk read and write errors), it is found that when the temperature exceeds 65 °C and lasts for more than 30 minutes, the occurrence probability of disk read and write errors will increase by 20%.
[0084] In this embodiment, by introducing the fault occurrence probability as the basis for priority sorting, it is ensured that the execution order of the processing strategy matches the degree of fault risk, improving the pertinence and efficiency of fault handling. By recording the execution effect of each fault handling instruction, a complete fault handling process tracking mechanism is established, making the fault handling process more controllable and visible. By feeding back the execution effect of the fault handling instruction to the fault type library, the adaptive optimization of the fault handling strategy is realized, enabling the fault handling solution to continuously improve and evolve.
[0085] In the second aspect of the embodiment of the present invention, a multi-source network device fault mode recognition and location system is provided, including: A first unit for collecting the usage rate of the central processing unit, memory usage rate, network interface traffic data, system log data, and network protocol data of the target network device; A second unit for respectively calculating the deviation values of the usage rate of the central processing unit, memory usage rate, and network interface traffic data from the preset baseline, generating a hardware layer fault feature matrix, performing semantic segmentation on the system log data, extracting log event sequences and process call chain information, combining with the protocol interaction records of the network protocol data, generating a software layer fault feature matrix, and performing fusion calculation on the hardware layer fault feature matrix and the software layer fault feature matrix to obtain the fault feature vector of the target network device; A third unit, configured to collect adjacent device identification information directly communicating with a target network device and link performance metric data; A fourth unit, configured to construct a device communication relationship graph based on the adjacent device identification information, calculate an anomaly level of the link performance metric data, map the anomaly level to the device communication relationship graph to obtain a link anomaly probability distribution graph, perform correlation analysis on the fault feature vector and the link anomaly probability distribution graph, determine a fault device influence range through recursive calculation, and output fault root cause device identification information and a fault occurrence probability; A fifth unit, configured to generate and execute a fault handling instruction according to the fault root cause device identification information and the fault occurrence probability, record fault recovery process data, and update baseline data.
[0086] In a third aspect of embodiments of the present invention, there is provided an electronic device, including: A processor and a memory for storing processor-executable instructions, wherein the processor is configured to call the instructions stored in the memory to execute the method described above.
[0087] In a fourth aspect of embodiments of the present invention, there is provided a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method described above is implemented.
[0088] The present invention may be a method, an apparatus, a system, and / or a computer program product. The computer program product may include a computer-readable storage medium, on which computer-readable program instructions for executing various aspects of the present invention are loaded.
[0089] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for multi-source network device fault mode recognition and location, characterized in that Including: Collecting the central processing unit (CPU) utilization rate, memory utilization rate, network interface traffic data, system log data, and network protocol data of the target network device; Calculating the deviation values of the CPU utilization rate, memory utilization rate, and network interface traffic data from the preset baseline respectively to generate a hardware layer fault feature matrix, performing semantic segmentation on the system log data, extracting log event sequences and process call chain information, combining with the protocol interaction records of the network protocol data to generate a software layer fault feature matrix, and performing fusion calculation on the hardware layer fault feature matrix and the software layer fault feature matrix to obtain the fault feature vector of the target network device; Collecting the adjacent device identification information and link performance index data that communicate directly with the target network device; Constructing a device communication relationship graph based on the adjacent device identification information, calculating the anomaly level of the link performance index data, mapping the anomaly level to the device communication relationship graph to obtain a link anomaly probability distribution graph, performing correlation analysis on the fault feature vector and the link anomaly probability distribution graph, and determining the influence range of the faulty device through recursive calculation, and outputting the faulty root device identification information and the fault occurrence probability; Generating a fault handling instruction according to the faulty root device identification information and the fault occurrence probability and executing it, recording the fault recovery process data and updating the baseline data.
2. The method according to claim 1, wherein Calculating the deviation values of the CPU utilization rate, memory utilization rate, and network interface traffic data from the preset baseline respectively to generate a hardware layer fault feature matrix, performing semantic segmentation on the system log data, extracting log event sequences and process call chain information, combining with the protocol interaction records of the network protocol data to generate a software layer fault feature matrix includes: Performing time series segmentation on the CPU utilization rate, memory utilization rate, and network interface traffic data of the target network device, calculating the baseline values of the CPU utilization rate, memory utilization rate, and network interface traffic data using exponential weighted moving average, and calculating the standardized deviation values through Z-score standardization; Performing weighted calculation on the standardized deviation values and the time decay factor to obtain the cumulative anomaly score, constructing a three-dimensional feature vector based on the cumulative anomaly score, and constructing a hardware layer fault feature matrix through a sliding time window; Classifying the system log data using hierarchical clustering for log template classification, performing semantic segmentation on the system log data based on the dynamic programming algorithm and calculating the event sequence entropy value, extracting process identifiers, process creation times, process end times, and parent-child process identifiers from the system log data to construct process call chain information, calculating the delay weighted value of the critical path in the process call chain information, analyzing the protocol interaction records of the network protocol data, and calculating the protocol interaction anomaly score by combining the violation weight and the violation frequency; Aligning and splicing the event sequence entropy value, delay weighted value, and protocol interaction anomaly score according to the time window to construct a software layer feature vector, and performing time series accumulation on the software layer feature vector to obtain a software layer fault feature matrix.
3. The method according to claim 1, wherein Performing fusion calculation on the hardware layer fault feature matrix and the software layer fault feature matrix to obtain the fault feature vector of the target network device includes: Determine the features in the hardware layer fault feature matrix that are greater than the preset hardware bottleneck threshold as hardware bottleneck features, calculate the influence weights of the hardware bottleneck features on the features in the software layer fault feature matrix, calculate the change rate of the features in the software layer fault feature matrix at adjacent times based on the influence weights, and obtain the hardware bottleneck influence degree; Determine the features in the software layer fault feature matrix that deviate from the feature mean by more than the preset anomaly determination coefficient and the feature standard deviation as software anomaly features, calculate the influence coefficients of the software anomaly features on the features in the hardware layer fault feature matrix, calculate the change rate of the relative maximum value of the features in the hardware layer fault feature matrix at future times based on the influence coefficients, and obtain the degree of hardware resource exhaustion; Calculate the weighted sum of the hardware bottleneck influence degree and the hardware resource exhaustion degree to obtain the performance bottleneck propagation intensity, cumulatively calculate the performance bottleneck propagation intensity within a preset time window to obtain the node importance, multiply the propagation intensities between each pair of nodes on the propagation path to obtain the propagation path importance, and calculate the weighted sum of the node importance and the propagation path importance to obtain the feature dynamic importance; Obtain the propagation perception weight by exponential normalization of the feature dynamic importance, obtain the fusion feature based on the product of the propagation perception weight and the corresponding features in the hardware layer fault feature matrix and the software layer fault feature matrix, and input the fusion feature into a preset activation function to obtain the fault feature vector of the target network device.
4. The method according to claim 1, wherein Collect the adjacent device identification information and link performance index data that communicate directly with the target network device, including: Obtain the list of adjacent devices of the target network device, and extract the corresponding adjacent device identification information of each adjacent device in the list of adjacent devices; Establish communication connections with each adjacent device in the list of adjacent devices, and collect the link performance index data between the target network device and each adjacent device; Associate the adjacent device identification information with the link performance index data to construct the link communication feature matrix of the target network device.
5. The method according to claim 1, wherein Construct a device communication relationship graph based on the adjacent device identification information, calculate the anomaly level of the link performance index data, map the anomaly level to the device communication relationship graph to obtain the link anomaly probability distribution graph, perform correlation analysis on the fault feature vector and the link anomaly probability distribution graph, and determine the influence range of the faulty device through recursive calculation, and output the faulty root device identification information and the fault occurrence probability, including: Construct a device communication relationship graph based on the adjacent device identification information, calculate the link weights of different communication types between devices, perform segmented sliding analysis on the link performance index data, calculate the product of the mutation amplitude and the duration of the performance index, and use the deviation degree of the product from the historical baseline corresponding to the current device as the link anomaly level, calculate the link anomaly probability based on the link anomaly level, and map the link anomaly probability to the corresponding links in the device communication relationship graph to obtain the link anomaly probability distribution graph; Extract the frequency-domain features and time-domain features of the fault feature vector through multi-scale decomposition, and combine the extracted features with the link anomaly probability with weights to obtain the local relevance of the node. Calculate the node importance based on the shortest path between nodes and the link weights. Map the product of the local relevance of the node and the node importance non-linearly to obtain the global relevance. Calculate the product of the cumulative decay effect of the link anomaly probability in the time series dimension and the local relevance of the node to obtain the propagation intensity matrix; Construct a node influence propagation chain according to the propagation intensity matrix, calculate the node state transition probability through hierarchical recursion until convergence to obtain the final node state, and perform probability fusion on the final node state and the global relevance to obtain the fault root cause device identification information and the fault occurrence probability.
6. The method according to claim 1, wherein Construct a node influence propagation chain according to the propagation intensity matrix, calculate the node state transition probability through hierarchical recursion until convergence to obtain the final node state, and perform probability fusion on the final node state and the global relevance to obtain the fault root cause device identification information and the fault occurrence probability, including: Construct a node influence propagation chain according to the propagation intensity matrix. The node influence propagation chain includes a node set, an edge set, and a propagation intensity weight matrix. Sample the historical state sequence of each node within a preset time series window to obtain a node state sequence. Calculate the time series dependence relationship between node state sequences based on mutual information measurement to obtain a time series dependence matrix. Conduct a Granger causality test on the node state sequence to obtain the direct causality intensity. Obtain all paths between nodes and calculate the cumulative product of the direct causality intensity on each path to obtain the indirect causality intensity; Calculate the weighted sum of the abnormal state duration and the decay coefficient for the node state sequence to obtain the state duration feature. Calculate the average value of the state difference between adjacent moments to obtain the change trend feature. Calculate the periodic feature based on the autocorrelation coefficient of the state sequence. Combine the state duration feature, the change trend feature, and the periodic feature to construct a time series feature vector; Construct a state transition probability matrix according to the propagation intensity weight matrix, the indirect causality intensity, and the time series feature vector. Update the node state by fusing the state transition probability matrix and the time series dependence matrix through hierarchical recursion until convergence to obtain the final node state; Perform probability fusion on the final node state and the global relevance to obtain the fault root cause device identification information and the fault occurrence probability.
7. The method according to claim 1, wherein Generate and execute a fault handling instruction according to the fault root cause device identification information and the fault occurrence probability, record the fault recovery process data, and update the baseline data, including: Based on the fault root cause device identification information and the fault occurrence probability, match the fault handling strategy corresponding to the fault root cause device identification information based on a preset fault type library, and sort the fault handling strategies according to the fault occurrence probability to generate a fault handling instruction; Execute the fault handling instruction, collect the fault recovery process data, and record the execution effect of the fault handling instruction; Update the baseline data based on the fault recovery process data, and feedback the execution effect of the fault handling instruction to the fault type library to complete the dynamic optimization of the fault handling strategy.
8. A multi-source network device fault mode recognition and location system for implementing the method described in any one of the foregoing claims 1-7, characterized in that, Including: The first unit is used to collect the central processor utilization rate, memory utilization rate, network interface traffic data, system log data, and network protocol data of the target network device; The second unit is used to calculate the deviation values of the central processor utilization rate, memory utilization rate, and network interface traffic data from the preset baseline respectively, generate a hardware layer fault feature matrix, perform semantic segmentation on the system log data, extract log event sequences and process call chain information, combine the protocol interaction records of the network protocol data, generate a software layer fault feature matrix, and perform fusion calculation on the hardware layer fault feature matrix and the software layer fault feature matrix to obtain the fault feature vector of the target network device; The third unit is used to collect the adjacent device identification information and link performance index data that communicate directly with the target network device; The fourth unit is used to construct a device communication relationship graph based on the adjacent device identification information, calculate the abnormal level of the link performance index data, map the abnormal level to the device communication relationship graph to obtain a link abnormal probability distribution graph, perform correlation analysis on the fault feature vector and the link abnormal probability distribution graph, determine the influence range of the faulty device through recursive calculation, and output the identification information of the faulty root cause device and the fault occurrence probability; The fifth unit is used to generate and execute a fault handling instruction according to the identification information of the faulty root cause device and the fault occurrence probability, record the fault recovery process data, and update the baseline data.
9. An electronic device, characterized in that, Comprising: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to call the instructions stored in the memory to execute the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Network fault diagnosis method and system based on log data
CN118138442A
Safety early warning method, device, equipment, medium and program product
CN119341888A
Communication network operation and maintenance fault positioning and tracking method and system
CN119420639A
Virtualized network fault rapid positioning method fusing multi-source information
CN119520229A
Adaptive operation and maintenance root cause positioning method and system based on deep learning
CN119691576A
Cited By
Integrated fault monitoring method and platform for power supply circuit
CN120559397A
Failure prediction method and device of PCIE link, equipment and storage medium
CN121217547A
Intelligent fault diagnosis method and system based on multi-source data
CN121940274A
A fault intelligent diagnosis method and system based on multi-source data
CN121940274B