A TianQing system fault management method based on deep learning

Through deep learning methods built by graph convolutional networks and recurrent neural networks, multi-node linkage faults in Tianqing system are identified and isolated, which solves the problem of insufficient accuracy and timeliness of fault diagnosis in traditional methods, and achieves efficient fault management and rapid repair.

CN120301785BActive Publication Date: 2025-08-08FUJIAN METEOROLOGICAL DATA CENTER (FUJIAN METEOROLOGICAL OBSERVATION CENTER FUJIAN METEOROLOGICAL ARCHIVES)
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510771802.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-11
Publication Date
2025-08-08
Estimated Expiration
2045-06-11

AI Technical Summary

Technical Problem

When facing the complex failure mode of Tianqing system, traditional rules-based fault diagnosis methods cannot accurately identify multi-node linkage fault modes, resulting in low accuracy and timeliness of fault diagnosis.

Method used

Deep learning methods based on graph convolutional networks and recurrent neural networks are adopted to analyze the dynamic interaction and fault propagation paths between nodes, build a multi-node fault management model, and identify and isolate fault nodes and influencing nodes by training and updating model parameters to achieve fault repair.

Benefits of technology

It improves the accuracy and efficiency of fault diagnosis, shortens fault detection and recovery time, enhances the system's adaptability and reliability, reduces the risks brought by human errors, and ensures that the system maintains efficient operation in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120301785B_ABST
    Figure CN120301785B_ABST
Patent Text Reader

Abstract

The present invention provides a TianQing system fault management method based on deep learning, which relates to the field of data processing technology. The method includes: analyzing the dynamic interactions and fault propagation paths between different nodes according to a graph convolutional network and a recurrent neural network, identifying potential multi-node linkage fault modes, and constructing an initial multi-node fault management model; training the initial multi-node fault management model according to historical operation data, dynamically adjusting the parameters of the initial multi-node fault management model, identifying the linkage effects generated when each node fails, and obtaining a multi-node fault management model; obtaining a real-time operation data set of the TianQing system, identifying the current faulty node and the influencing node, determining the current multi-node linkage fault mode, isolating the faulty node and the influencing node, eliminating the influence of the faulty node on the influencing node, and obtaining a fault repair result. The present invention manages faults in the TianQing system through deep learning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing technology, and in particular to a TianQing system fault management method based on deep learning. Background Art

[0002] In existing fault management methods, rule-based monitoring systems are usually used to monitor the system status in real time through predefined rules. However, traditional rule-based fault diagnosis methods may have certain limitations when facing complex fault modes in the system. When a node fails, the failure of the node may cause abnormalities in multiple related nodes. The traditional rule system may not be able to accurately identify and locate the source of the fault, resulting in low fault diagnosis accuracy.

[0003] For example, in the computing nodes of the TianQing system, a computing node experienced a service interruption due to a hardware failure, resulting in a performance bottleneck in the storage node that the node depended on. Traditional rule-based systems determine failures by monitoring the status of individual nodes, but may not be able to promptly detect the expansion effect of the failure and its impact on the entire system, and cannot effectively identify complex failure modes with chain reactions, thus affecting the timeliness and accuracy of failure recovery. Summary of the Invention

[0004] The purpose of the present invention is to provide a TianQing system fault management method based on deep learning, aiming to solve the problems mentioned in the background technology.

[0005] In order to solve the above technical problems, the technical solutions of the present invention are as follows:

[0006] A TianQing system fault management method based on deep learning, the method comprising:

[0007] Obtain node status data, performance monitoring data, and fault log data for each node in the TianQing system to obtain a historical operation data set;

[0008] Using graph convolutional networks and recurrent neural networks, we analyze the dynamic interactions and fault propagation paths between different nodes, identify potential multi-node linkage failure modes, and build an initial multi-node fault management model.

[0009] The initial multi-node fault management model is trained based on historical operation data, and the parameters of the initial multi-node fault management model are dynamically adjusted to identify the linkage effects caused by failures of each node, thereby obtaining a multi-node fault management model.

[0010] Obtain the real-time operating data set of the TianQing system and input it into the multi-node fault management model to identify the current faulty node and affected nodes, determine the current multi-node linkage fault mode, and obtain the fault status data set;

[0011] Based on the fault status data set, the faulty node and the affected nodes are isolated, the impact of the faulty node on the affected nodes is eliminated, and the fault repair result is obtained;

[0012] The multi-node fault management model is updated according to the fault repair results to improve the accuracy of the multi-node fault management model in identifying multi-node linkage failure modes.

[0013] Furthermore, based on graph convolutional networks and recurrent neural networks, we analyze the dynamic interactions and fault propagation paths between different nodes, identify potential multi-node linkage failure modes, and build an initial multi-node fault management model, including:

[0014] Construct an input layer to receive historical operation data sets, extract node status data and performance monitoring data of each node at different time steps, and perform feature extraction on the node status data and performance monitoring data of each node at different time steps to obtain node feature sequences;

[0015] A graph convolutional network layer is constructed to receive node feature sequences and generate a static adjacency matrix based on the static structural relationship of data transmission between nodes in the TianQing system. The current feature vector of each node is then aggregated with the feature vectors of its adjacent nodes to obtain a set of structural association sequences.

[0016] Furthermore, based on graph convolutional networks and recurrent neural networks, we analyze the dynamic interactions and fault propagation paths between different nodes, identify potential multi-node linkage failure modes, and build an initial multi-node fault management model, including:

[0017] Construct a recurrent neural network layer to calculate the hidden state of each node in chronological order based on the set of structural association sequences, identify the state change relationship of each node between different time steps, and obtain the node state time series dataset;

[0018] Construct a linkage fault identification layer to calculate the state correlation of different nodes based on the node state time series data set. When the state correlation of multiple nodes exceeds the preset correlation threshold, a fault propagation path is established between the nodes, the set of nodes with linkage influence is identified, and a multi-node linkage fault mode data set is generated;

[0019] Construct an output layer to receive and output the multi-node linkage failure mode dataset;

[0020] The initial multi-node fault management model is constructed by connecting the input layer, graph convolutional network layer, recurrent neural network layer, linkage fault identification layer and output layer in sequence.

[0021] Furthermore, a graph convolutional network layer is constructed to receive node feature sequences and generate a static adjacency matrix based on the static structural relationship of data transmission between nodes in the TianQing system. The current feature vector of each node is then aggregated with the feature vectors of its adjacent nodes to obtain a set of structural association sequences, including:

[0022] Based on the static connection relationship between nodes in the TianQing system during data transmission, a static adjacency matrix is constructed. The static adjacency matrix is a symmetric matrix, in which each non-zero element indicates a direct interaction relationship between two nodes.

[0023] According to the static adjacency matrix, the feature vectors of the current node and each adjacent node are transformed by matrix multiplication to obtain the structure-aware feature vector of the current node;

[0024] By performing nonlinear activation on the structure-aware feature vector, the interaction structure characteristics of each node in the current time period are obtained;

[0025] The interaction structural features of all nodes at different time steps are arranged and merged according to the node dimension and time dimension to form a structural association sequence dataset.

[0026] Furthermore, a recurrent neural network layer is constructed to calculate the hidden state of each node in chronological order based on the set of structural association sequences, identify the state change relationship of each node between different time steps, and obtain a node state time series dataset, including:

[0027] Input the structure association sequence set into the recurrent neural network layer in chronological order, and initialize the initial hidden states of all nodes to the default initial states when the TianQing system just starts running;

[0028] At each time step, the node structure perception feature vector of the current time period and the hidden state of the previous time step are used as input to calculate the hidden state of the current time step;

[0029] By calculating the hidden state of all time steps, the hidden state of each node at each time step is determined to obtain the hidden state data set;

[0030] Based on the hidden state data set, the hidden state difference of each node between adjacent time steps is calculated, and the state changes of each node between different time steps are analyzed based on the hidden state difference. When the hidden state difference between adjacent time steps exceeds the preset difference threshold, the node state is identified as a significant change and marked as an abnormal state fluctuation.

[0031] According to the abnormal state fluctuations, the hidden state changes of each node between all time steps are extracted to obtain the node state time series dataset.

[0032] Furthermore, a linkage fault identification layer is constructed to calculate the state correlation of different nodes based on the node state time series data set. When the state correlation of multiple nodes exceeds the preset correlation threshold, a fault propagation path is established between the nodes, and the set of nodes with linkage influence is identified to generate a multi-node linkage fault mode data set, including:

[0033] Based on the node state time series dataset, the hidden state vectors of the two nodes at each time step are extracted respectively. The first similarity value is calculated by performing dot product and normalization on the two vectors. The change amplitude of the hidden state vectors of the two nodes at different time steps relative to the previous time step is calculated, and the difference between the change amplitudes of the two nodes is calculated to obtain the second similarity value.

[0034] Determine whether two nodes experience abnormal state fluctuations at the same time in the current time step. When both nodes are in abnormal state, the abnormal response coupling term of this time step takes a fixed value. When there is no abnormality at any node, the term is zero, and the third similarity value is obtained.

[0035] Calculate the state correlation between the state change trends of different nodes in the entire time series according to the first similarity value, the first similarity value, and the first similarity value to obtain a state correlation data set;

[0036] When the state correlation between two nodes exceeds the preset correlation threshold, it is determined that there is a fault coupling relationship between the two nodes, and a fault propagation path is established between the two nodes to obtain the linkage-affected node set;

[0037] According to the linkage impact node set and its propagation path, a multi-node linkage fault mode dataset is generated.

[0038] Furthermore, the initial multi-node fault management model is trained based on historical operation data, and the parameters of the initial multi-node fault management model are dynamically adjusted to identify the linkage effects caused by failures of each node. The multi-node fault management model is obtained, including:

[0039] The training set and validation set are obtained by randomly selecting node status data and performance monitoring data from the historical operation data set;

[0040] According to the fault log data, the fault propagation path between each node is extracted to obtain the actual fault mode;

[0041] Inputting the training set into the initial multi-node fault management model, obtaining the multi-node linkage failure mode predicted by the initial multi-node fault management model, and obtaining the training failure mode;

[0042] By comparing the differences between the training failure mode and the real failure mode, the training deviation path is obtained;

[0043] According to the training deviation path, the parameters in the initial multi-node fault management model are updated and adjusted to gradually reduce the training deviation path and obtain the training deviation path change rate;

[0044] Through multiple rounds of training iterations, when the training deviation path change rate is less than the preset change rate threshold, a multi-node fault management model is obtained.

[0045] Furthermore, through multiple rounds of training iterations, when the training deviation path change rate is less than a preset change rate threshold, a multi-node fault management model is obtained, including:

[0046] After each round of training is completed, the validation set is input into the updated multi-node fault management model of the current round to obtain the verified fault mode;

[0047] Compare the verification failure mode with the actual failure mode to obtain the verification deviation path;

[0048] According to the verification deviation path, determine whether the verification deviation path of the current round of training is smaller than that of the previous round of training. If the verification deviation path of the preset number of consecutive rounds is not lower than the preset verification deviation threshold, the training process is terminated, otherwise the training is continued;

[0049] When the training deviation path change rate is less than the preset change rate threshold and the verification deviation path for a preset number of consecutive rounds is not less than the preset verification deviation threshold, the training process is terminated and a multi-node fault management model is obtained.

[0050] Furthermore, based on the fault status data set, the faulty node and the affected nodes are isolated to eliminate the impact of the faulty node on the affected nodes, and the fault repair results are obtained, including:

[0051] Based on the fault status data set, determine whether the current operating status of the fault node meets the fault isolation conditions. When the abnormal state duration of the fault node exceeds the preset abnormal time threshold, an isolation signal is triggered;

[0052] According to the isolation signal, the data transmission between the faulty node and the affected node is cut off, and the data transmission path is reallocated to the affected node to obtain the reconstructed transmission state;

[0053] According to the reconstructed transmission status, confirm whether the affected node has resumed normal operation after isolating the faulty node. When the status of the affected node is normal, mark the isolation process as successful and generate a fault repair result.

[0054] Furthermore, the multi-node fault management model is updated based on the fault repair results to improve the accuracy of the multi-node fault management model in identifying multi-node linkage fault modes, including:

[0055] Based on the fault repair results, the fault propagation path in the historical operation trajectory is marked with the fault node as the center to obtain a supplementary training set;

[0056] An incremental data set is obtained based on the supplementary training set and the training set, and is input into the multi-node fault management model for parameter fine-tuning to obtain an updated multi-node fault management model;

[0057] Perform a validation test on the updated multi-node fault management model. If its recognition accuracy on the validation set improves or does not decrease, retain the updated multi-node fault management model as the current valid model configuration.

[0058] When the recognition accuracy of the updated multi-node fault management model on the validation set tends to be stable and does not improve any further, the updating process is stopped to avoid overfitting.

[0059] The above solution of the present invention includes at least the following beneficial effects:

[0060] The present invention combines graph convolutional networks and recurrent neural networks to provide a deep learning model that can simultaneously consider the static structural relationship and temporal dynamic changes between nodes. The graph convolutional network can make full use of the static structural information between nodes in the TianQing system, accurately model the interaction characteristics between nodes, and thus identify the mutual influence between nodes. By constructing a static adjacency matrix and aggregating the feature information of adjacent nodes, the graph convolutional network can effectively capture the relationship between nodes and optimize the representation of the state characteristics of the nodes. The recurrent neural network is responsible for processing time series data and identifying the state changes of nodes between different time steps. This combination enables the system to accurately capture the laws of node state changes in the time dimension, thereby improving the ability to predict the occurrence of faults. This combined method has a strong fault pattern recognition capability, can promptly detect multi-node linkage faults, and improve the accuracy of fault diagnosis.

[0061] The present invention analyzes the dynamic interactions and fault propagation paths between nodes through graph convolutional networks and recurrent neural networks, and constructs an initial multi-node fault management model. Through in-depth analysis of node status data, potential multi-node linkage failure modes can be identified. Compared with traditional rule-based monitoring methods, this deep learning method has higher adaptability and accuracy. Since multiple nodes in the system are often in a state of interdependence, the failure of a single node may trigger a chain reaction, affecting the normal operation of other nodes. By using graph convolutional networks to capture the structural relationship between nodes and combining the learning of time series data by recurrent neural networks, the initial model can quickly identify these linkage failure modes, so that fault diagnosis is no longer limited to detecting anomalies of a single node, but can identify and predict the occurrence and propagation paths of multi-node faults from a global perspective, thereby improving the accuracy and efficiency of fault diagnosis.

[0062] By combining graph convolutional networks and recurrent neural networks, the present invention can accurately identify the current fault node and its affected nodes in a real-time environment. The system can immediately identify the impact range and respond at the early stage of the fault, rather than just waiting for the fault to fully manifest. This real-time performance is impossible for traditional rule systems to achieve. It can greatly shorten the time for fault detection and recovery, and improve the system's operation and maintenance efficiency. Through the input of real-time data, the system can flexibly respond to complex and dynamically changing fault modes and identify those unforeseen linkage faults, so that the system can still maintain a high level of fault detection and repair capabilities when facing complex network environments and large-scale distributed systems. Real-time data processing not only improves the system's response speed, but also enhances the accuracy of fault recovery and reduces the losses that may be caused by delayed responses.

[0063] The present invention prevents the spread of faults by isolating faulty nodes and affected nodes. Traditional fault handling often relies on manual intervention or rule setting, but in complex distributed systems, the limitations of manual intervention and rules often cannot effectively deal with sudden and complex faults. Through automated isolation operations, not only can the faulty nodes be quickly isolated, but also the data transmission path can be reconstructed to ensure that other nodes can resume normal operation. This operation not only reduces the impact of the fault on the entire system, but also improves the system's recovery speed, reduces downtime, and improves the reliability of the overall system. By isolating the faulty nodes from the affected nodes, the system can ensure that the scope of the fault impact is minimized, avoids serious damage to the entire system caused by the propagation of the fault, improves processing efficiency, and reduces the risks caused by human errors, so that the system has stronger fault tolerance and higher availability.

[0064] The present invention ensures that the system maintains high fault identification accuracy and real-time performance during long-term operation by continuously updating and iteratively optimizing the multi-node fault management model. As the system operating environment continues to change and new fault modes emerge, traditional fault management methods are often unable to adapt to changes in a timely manner. By continuously feeding back new fault repair results into the model for optimization, the fault management system has strong adaptability. Through continuous training and verification, the system can continuously adjust the model parameters and gradually improve the accuracy of fault identification. This self-update and iterative optimization mechanism ensures that the system can cope with increasingly complex fault modes, thereby improving reliability in long-term operation and maintenance, while avoiding the risk of overfitting, enabling the model to maintain efficient generalization capabilities, ensuring that it can exhibit high fault diagnosis capabilities under different operating conditions, and enabling the system to always be in the best operating state, effectively improving the overall performance and long-term stability of the fault management system. BRIEF DESCRIPTION OF THE DRAWINGS

[0065] Figure 1 This is a flowchart of a TianQing system fault management method based on deep learning provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0066] Exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. Rather, these embodiments are provided to enable a more thorough understanding of the present disclosure and to fully convey the scope of the present disclosure to those skilled in the art.

[0067] like Figure 1 As shown, an embodiment of the present invention proposes a TianQing system fault management method based on deep learning, the method comprising:

[0068] Obtain node status data, performance monitoring data, and fault log data for each node in the TianQing system to obtain a historical operation data set;

[0069] Using graph convolutional networks and recurrent neural networks, we analyze the dynamic interactions and fault propagation paths between different nodes, identify potential multi-node linkage failure modes, and build an initial multi-node fault management model.

[0070] The initial multi-node fault management model is trained based on historical operation data, and the parameters of the initial multi-node fault management model are dynamically adjusted to identify the linkage effects caused by failures of each node, thereby obtaining a multi-node fault management model.

[0071] Obtain the real-time operating data set of the TianQing system and input it into the multi-node fault management model to identify the current faulty node and affected nodes, determine the current multi-node linkage fault mode, and obtain the fault status data set;

[0072] Based on the fault status data set, the faulty node and the affected nodes are isolated, the impact of the faulty node on the affected nodes is eliminated, and the fault repair result is obtained;

[0073] The multi-node fault management model is updated according to the fault repair results to improve the accuracy of the multi-node fault management model in identifying multi-node linkage failure modes.

[0074] In an embodiment of the present invention, node status data, performance monitoring data and fault log data of each node in the TianQing system are obtained to obtain a historical operation data set, and comprehensive and accurate historical data can be obtained, which provides solid data support for subsequent fault mode analysis and prediction; based on the graph convolutional network and the recurrent neural network, the dynamic interactions and fault propagation paths between different nodes are analyzed, potential multi-node linkage fault modes are identified, and an initial multi-node fault management model is constructed, so that the system can accurately analyze the dynamic interaction relationship between nodes and identify multi-node linkage fault modes; the initial multi-node fault management model is trained according to the historical operation data, the parameters of the initial multi-node fault management model are dynamically adjusted, the linkage impact generated when each node fails is identified, and the multi-node fault management model is obtained. By training the initial multi-node fault management model, the system can optimize the parameters so that the model can be more adaptable and accurate when facing new fault modes.

[0075] The real-time operation data set of the TianQing system is obtained and input into the multi-node fault management model to identify the current fault node and the influencing node, determine the current multi-node linkage fault mode, and obtain the fault status data set, thereby ensuring the timeliness of fault detection and response, and enabling the system to respond in the early stages of the fault; based on the fault status data set, the fault node and the influencing node are isolated and processed, the impact of the fault node on the influencing node is eliminated, and the fault repair result is obtained. Through fault isolation processing, the system can eliminate the impact of the fault node on other nodes in the shortest time and prevent the fault from spreading to more nodes; based on the fault repair result, the multi-node fault management model is updated to improve the accuracy of the multi-node fault management model in identifying the multi-node linkage fault mode. By continuously updating the multi-node fault management model based on the latest fault repair data, the system can maintain a high level of fault identification accuracy and can gradually adapt to changes in the system environment.

[0076] The node status data, performance monitoring data, and fault log data of each node in the TianQing system are obtained to obtain the historical operation data set, which specifically includes:

[0077] The system deploys monitoring tools and sensors to collect various key data from each node of the TianQing system in real time. Node status data usually includes basic operating information of each node, such as CPU utilization, memory usage, network bandwidth, storage usage, etc., which can reflect the overall health of the node; performance monitoring data provides performance data about the node under different loads, such as the number of tasks carried by each node during operation, response time, etc. These data help the system evaluate the real-time performance of the node and identify whether there are bottlenecks or abnormal performance; fault log data contains information such as the type, time and repair process of faults that have occurred in the past. This information is crucial for analyzing historical failure modes, identifying potential root causes of failures and repair strategies.

[0078] The collected data will be transmitted to a central storage or cloud platform for centralized processing. To ensure the accuracy and availability of the data, the system first pre-processes the data, which includes removing redundant data, filling in missing data, standardizing data formats, and other operations. Through these data cleaning steps, the system can ensure the quality of the historical operation data set. The cleaned data will be integrated into a historical data set, covering the operating status, performance monitoring data, and fault records of all nodes in different time periods. This data set provides the necessary foundation for the subsequent training of deep learning models, helping the model identify the operating rules of different nodes and analyze the interdependence between nodes and potential failure modes. Through this process, the system can accumulate a large amount of real historical data. This data can not only help train the fault management model, but also provide the system with an accurate reference basis to identify and predict possible future failure modes. The historical operation data set has become the cornerstone of fault diagnosis, prediction, and repair decisions.

[0079] The real-time operation data set of the TianQing system is obtained and input into the multi-node fault management model to identify the current faulty node and the affected node, determine the current multi-node linkage fault mode, and obtain the fault status data set, which specifically includes:

[0080] The system needs to obtain the real-time operation data of the TianQing system so that it can respond in time when a fault occurs. The real-time operation data set is composed of real-time monitoring data of each node in the TianQing system, including the node's current status information, performance data and instant feedback on the fault status. This data includes the node's real-time resource consumption such as CPU, memory, storage, network status, service response time, data traffic, etc. Through real-time monitoring tools, the system can continuously collect this data and transmit it to the central system for centralized processing.

[0081] After obtaining real-time data, the system will input this data into a trained multi-node fault management model for analysis. Based on the input data, the multi-node fault management model will identify the current fault node and the affected nodes caused by the fault node, while considering the spatial relationship and temporal changes between nodes. Through the graph convolutional network, the model can effectively capture the static structural relationship between nodes and analyze the interaction of nodes during failures. At the same time, the recurrent neural network can process the time series data of node status, identify the dynamic changes of node status, and thus capture the temporal characteristics of node fault propagation.

[0082] After identifying the faulty node, the system will further determine the fault propagation path and analyze the status of the affected nodes. The affected nodes are those nodes that may experience performance degradation or service interruption due to the abnormality of the faulty node. Through the operation of the multi-node fault management model, the system can accurately identify the current multi-node linkage failure mode, describe the faulty node and its impact range. This process not only helps to detect the faulty node in a timely manner, but also can accurately determine which nodes will be affected by the abnormality of the faulty node.

[0083] Ultimately, the system will generate a fault status dataset containing detailed information about the current fault node and affected nodes, as well as the linkage relationship and fault propagation path between them. The generation of the fault status dataset provides a detailed basis for subsequent fault isolation and repair processing. Through this real-time dataset, the system can quickly respond to current faults, identify potential system risks, and take appropriate measures to isolate and repair faults in a timely manner, thereby providing effective fault management support.

[0084] In a preferred embodiment of the present invention, based on graph convolutional networks and recurrent neural networks, the dynamic interactions and fault propagation paths between different nodes are analyzed, potential multi-node linkage failure modes are identified, and an initial multi-node fault management model is constructed, including:

[0085] Construct an input layer to receive historical operation data sets, extract node status data and performance monitoring data of each node at different time steps, and perform feature extraction on the node status data and performance monitoring data of each node at different time steps to obtain node feature sequences;

[0086] A graph convolutional network layer is constructed to receive node feature sequences and generate a static adjacency matrix based on the static structural relationship of data transmission between nodes in the TianQing system. The current feature vector of each node is then aggregated with the feature vectors of its adjacent nodes to obtain a set of structural association sequences.

[0087] In an embodiment of the present invention, an input layer is constructed to receive historical operation data sets, extract node status data and performance monitoring data of each node at different time steps, and perform feature extraction on the node status data and performance monitoring data of each node at different time steps to obtain a node feature sequence. By constructing an input layer and performing feature extraction on the data, the system can efficiently process a large amount of complex historical data and extract key feature information of each node;

[0088] A graph convolutional network layer is constructed to receive node feature sequences and generate a static adjacency matrix based on the static structural relationship of data transmission between nodes in the TianQing system. The current feature vector of each node is then aggregated with the feature vectors of its adjacent nodes to obtain a set of structural association sequences. By aggregating the features of adjacent nodes, the structural perception features of each node can be constructed, thereby capturing the complex relationships between nodes in the network.

[0089] Among them, the input layer is constructed to receive the historical operation data set, extract the node status data and performance monitoring data of each node at different time steps, and perform feature extraction on the node status data and performance monitoring data of each node at different time steps to obtain the node feature sequence, which specifically includes:

[0090] In the process of building the input layer, we first need to collect the historical operation data of each node from the TianQing system. The historical operation data of each node usually includes the node's status data and performance monitoring data. These data are recorded and stored in the form of time series. The status and performance data of each node at different time points have corresponding timestamps, forming a series of time step data. The system will organize these data into a multidimensional data matrix according to the time series. Each row represents a record of a time step, and each column represents the status data and performance data of a node. Since the status data and performance monitoring data of each node are different, the data of each time step will provide the system with key information about the node status changes.

[0091] Next, the system will extract features from the status data and performance monitoring data of each node at different time steps, and convert the original status data and performance data into more concise feature vectors that can reflect the node's operating characteristics. The system will extract the key features of each node from the original data, such as the average value, variance, maximum value, minimum value, and change trend. These statistical features can help the model better understand the node's operating status over a period of time. For example, the average value of the node's CPU usage can reflect the node's load level, while the variance can reflect the volatility of the load. In addition, the node's response time change trend can also provide the system with important information on whether the node's performance is stable.

[0092] By extracting features from the status data and performance monitoring data of each node at different time steps, the system will eventually obtain a feature sequence for each node. These feature sequences not only contain the status information of the node itself, but also cover aspects such as the node's performance fluctuations and operational stability, and can provide more efficient input data for subsequent deep learning models. By extracting these key features, the system effectively represents the status of each node in the time dimension, providing rich input data for the subsequent processing of graph convolutional networks and recurrent neural network layers.

[0093] In a preferred embodiment of the present invention, based on graph convolutional networks and recurrent neural networks, the dynamic interactions and fault propagation paths between different nodes are analyzed, potential multi-node linkage failure modes are identified, and an initial multi-node fault management model is constructed, including:

[0094] Construct a recurrent neural network layer to calculate the hidden state of each node in chronological order based on the set of structural association sequences, identify the state change relationship of each node between different time steps, and obtain the node state time series dataset;

[0095] Construct a linkage fault identification layer to calculate the state correlation of different nodes based on the node state time series data set. When the state correlation of multiple nodes exceeds the preset correlation threshold, a fault propagation path is established between the nodes, the set of nodes with linkage influence is identified, and a multi-node linkage fault mode data set is generated;

[0096] Construct an output layer to receive and output the multi-node linkage failure mode dataset;

[0097] The initial multi-node fault management model is constructed by connecting the input layer, graph convolutional network layer, recurrent neural network layer, linkage fault identification layer and output layer in sequence.

[0098] In an embodiment of the present invention, a recurrent neural network layer is constructed to calculate the hidden state of each node in chronological order based on a set of structural association sequences, identify the state change relationship of each node between different time steps, and obtain a node state time series data set. The system can capture the dynamic characteristics of node states changing over time. A linkage fault identification layer is constructed to calculate the state correlation of different nodes based on the node state time series data set. When the state correlation of multiple nodes exceeds a preset correlation threshold, a fault propagation path is established between the nodes, a set of nodes with linkage influence is identified, and a multi-node linkage fault mode data set is generated. By calculating the state correlation between the nodes, the linkage relationship between the multiple nodes is effectively identified, revealing potential multi-node linkage fault modes. An output layer is constructed to receive and output the multi-node linkage fault mode data set, intuitively presenting the location of the fault, the propagation path, and the affected node information. An initial multi-node fault management model is constructed by sequentially connecting the input layer, graph convolutional network layer, recurrent neural network layer, linkage fault identification layer, and output layer. Through this connection, the system can monitor and diagnose faults in the TianQing system from all aspects, not only accurately identifying single node faults, but also revealing the linkage effects between nodes, thereby effectively improving the accuracy and response speed of fault diagnosis.

[0099] The initial multi-node fault management model is constructed by sequentially connecting the input layer, graph convolutional network layer, recurrent neural network layer, linkage fault identification layer, and output layer. Specifically, it includes:

[0100] First, the system collects node status data and performance monitoring data from the historical operation data of the TianQing system, and organizes it into a feature set containing a time step sequence. The data of each node at each time step is processed into a node feature vector. These feature vectors will serve as the output data of the input layer, providing a basis for subsequent deep learning processing.

[0101] Next, these output data enter the graph convolutional network layer. Through the graph convolution operation, the feature vector of each node is weighted averaged with the feature vector of its adjacent nodes to obtain a new feature vector. This process can be regarded as a process of information propagation. At each node, the feature vector is affected not only by its own state, but also by the state of its adjacent nodes. This aggregation process is the core of the graph convolutional network and can help the system understand the structural relationship and interaction pattern between nodes.

[0102] Then, the node features will be passed to the recurrent neural network layer, and the hidden state of the current time step will be calculated based on the feature vector of the current node and the hidden state of the previous time step. Through such a recursive process, the changing trend and time correlation of the node state can be identified. The node state of each time step is recorded as a hidden state. Ultimately, these hidden states constitute the dynamic behavior of the node in the entire time series.

[0103] Subsequently, the node status data will enter the linkage fault identification layer. By calculating the similarity between the state vectors of different nodes and analyzing their state change trends, the node status data in the time series will be used to mine the potential fault propagation chain, ensuring that the system can promptly identify the fault propagation path and predict which nodes may be affected.

[0104] Finally, the data processed by the linkage fault identification layer enters the output layer. Based on the status data of the current node and the linkage fault mode, a detailed fault status data set is generated, indicating the faulty node and the affected nodes. Through this output result, the system can intuitively provide detailed information about the fault to the operation and maintenance personnel, including the fault source, propagation path and the set of affected nodes.

[0105] In a preferred embodiment of the present invention, a graph convolutional network layer is constructed to receive node feature sequences and generate a static adjacency matrix based on the static structural relationship of data transmission between nodes in the TianQing system. The current feature vector of each node is then aggregated with the feature vectors of its adjacent nodes to obtain a set of structural association sequences, including:

[0106] Based on the static connection relationship between nodes in the TianQing system during data transmission, a static adjacency matrix is constructed. The static adjacency matrix is a symmetric matrix, in which each non-zero element indicates a direct interaction relationship between two nodes.

[0107] According to the static adjacency matrix, the feature vectors of the current node and each adjacent node are transformed by matrix multiplication to obtain the structure-aware feature vector of the current node;

[0108] By performing nonlinear activation on the structure-aware feature vector, the interaction structure characteristics of each node in the current time period are obtained;

[0109] The interaction structural features of all nodes at different time steps are arranged and merged according to the node dimension and time dimension to form a structural association sequence dataset.

[0110] In an embodiment of the present invention, a static adjacency matrix is constructed based on the static connection relationship between nodes in the TianQing system during the data transmission process. The static adjacency matrix is a symmetric matrix in which each non-zero element indicates a direct interaction relationship between two nodes. By constructing the static adjacency matrix, the system can fully capture the structural relationship between nodes, accurately represent the dependency between nodes, and provide a reliable foundation for subsequent graph convolution calculations. According to the static adjacency matrix, the feature vectors of the current node and each adjacent node are transformed by matrix multiplication to obtain the structure-aware feature vector of the current node, which can integrate the interaction information between nodes, so that the feature vector of each node can reflect the relative importance and role of the node in the entire network. By performing nonlinear activation on the structure-aware feature vector, the interaction structure characteristics of each node in the current time period are obtained. The system can avoid the limitations of linear models in dealing with complex fault modes and improve the fitting ability of the model. The interaction structure characteristics of all nodes at different time steps are arranged and merged according to the node dimension and time dimension to form a structure-related sequence data set, which can reflect the dynamic changes of the node itself and capture the linkage changes between nodes.

[0111] Among them, based on the static connection relationship between each node in the TianQing system during the data transmission process, a static adjacency matrix is constructed. The static adjacency matrix is a symmetric matrix, in which each non-zero element indicates a direct interaction relationship between two nodes, specifically including:

[0112] First, the system needs to construct a static adjacency matrix based on the data transmission relationship between each node in the TianQing system. This adjacency matrix is used to represent the static connection relationship between each node. In the TianQing system, there may be different data transmission paths between each node, and the connectivity between nodes can be static, that is, it does not change over time. The system constructs a symmetric adjacency matrix based on the existing node connection information, where each non-zero element of the matrix indicates a direct interaction relationship between two nodes. The adjacency matrix construction process actually represents each node and its connection relationship in the system through a graph structure, providing a basis for the subsequent feature aggregation of the graph convolutional network.

[0113] Among them, according to the static adjacency matrix, the feature vectors of the current node and each adjacent node are transformed by matrix multiplication to obtain the structure-aware feature vector of the current node, which specifically includes:

[0114] After obtaining the adjacency matrix, the system weights and aggregates the eigenvectors of the adjacent nodes into the eigenvector of the current node. This process is completed through matrix multiplication. The eigenvector of each node is multiplied by the corresponding row of the adjacency matrix to obtain the structural perception eigenvector of the current node. In this way, the characteristics of the node not only reflect its own state, but also include the interaction and connection with other nodes.

[0115] Among them, by performing nonlinear activation on the structure-aware feature vector, the interactive structural features of each node in the current time period are obtained, including:

[0116] By applying a linear rectification function, the system can perform nonlinear transformations on the structural perception features of each node, enabling the feature vector to better adapt to complex failure modes and data changes. The linear rectification function performs threshold processing on each input feature vector, suppressing negative values and retaining positive values, thereby accelerating training and reducing the risk of overfitting.

[0117] The linear rectification function expression is: ,in, Is the input value, usually each element in the structure-aware feature vector of the node. The role of the linear rectification function is to map all negative values to zero, and the positive values remain unchanged. That is, for each input feature, if its value is greater than zero, the original value is retained; if its value is less than or equal to zero, it is mapped to zero.

[0118] After activation function processing, the structural perception feature vector of each node will become a new nonlinear feature representation, reflecting the interaction structure characteristics of the node in the current time period. These new features can not only more accurately represent the state changes of the node, but also enhance the network model's understanding of the nonlinear relationship between nodes, thereby improving the ability to identify and predict fault modes. Through this nonlinear processing, the system can better capture the complex interaction patterns between nodes, especially when facing multi-node linkage failures, and can identify the potential impacts and associations between nodes in advance.

[0119] Among them, the interactive structural features of all nodes at different time steps are arranged and merged according to the node dimension and time dimension to form a structural association sequence dataset, which specifically includes:

[0120] The system arranges the interaction structural features of all nodes at different time steps according to the time and node dimensions to form a structural association sequence dataset. This dataset contains the state information of each node at different time steps and encodes the interaction influence between nodes based on the adjacency relationship. The structural association sequence dataset actually combines the time series data with the structural dependency between nodes to form a time series dataset that can be processed by subsequent deep learning models.

[0121] In a preferred embodiment of the present invention, a recurrent neural network layer is constructed to calculate the hidden state of each node in chronological order based on the set of structural association sequences, identify the state change relationship of each node between different time steps, and obtain a node state time series data set, including:

[0122] Input the structure association sequence set into the recurrent neural network layer in chronological order, and initialize the initial hidden states of all nodes to the default initial states when the TianQing system just starts running;

[0123] At each time step, the node structure perception feature vector of the current time period and the hidden state of the previous time step are used as input to calculate the hidden state of the current time step;

[0124] By calculating the hidden state of all time steps, the hidden state of each node at each time step is determined to obtain the hidden state data set;

[0125] Based on the hidden state data set, the hidden state difference of each node between adjacent time steps is calculated, and the state changes of each node between different time steps are analyzed based on the hidden state difference. When the hidden state difference between adjacent time steps exceeds the preset difference threshold, the node state is identified as a significant change and marked as an abnormal state fluctuation.

[0126] According to the abnormal state fluctuations, the hidden state changes of each node between all time steps are extracted to obtain the node state time series dataset.

[0127] In an embodiment of the present invention, the structure association sequence set is input into the recurrent neural network layer in chronological order, and the initial hidden states of all nodes are initialized to the default initial states when the TianQing system starts running, which provides a unified starting point for the system and ensures that each node is consistent when training starts; at each time step, the node structure perception feature vector of the current time period and the hidden state of the previous time step are used as input to calculate the hidden state of the current time step, which can capture the state changes of the node between different time steps; by calculating the hidden states of all time steps, the hidden state of each node at each time step is determined, and a hidden state data set is obtained. By updating the hidden state, the system can gradually enhance the fault detection capability and improve the fault prediction accuracy; based on the hidden state data set, the hidden state difference of each node between adjacent time steps is calculated, and based on it, the state changes of each node between different time steps are analyzed. When the hidden state difference between adjacent time steps exceeds the preset difference threshold, the node state is identified to have changed significantly and marked as abnormal state fluctuation. By calculating the node state change relationship and identifying significant changes, the system can accurately capture the abnormal fluctuation of the node state and then discover potential faults or other abnormal states; based on the abnormal state fluctuation, the hidden state changes of each node between all time steps are extracted to obtain the node state time series data set. By collecting node state information of multiple time steps, the system can fully understand the performance of the node in the entire time process.

[0128] The calculation formula of the hidden state is:

[0129] ,

[0130] in, For nodes In the The hidden state of time steps, is the index of the node, is the index of the time step, For nodes In the The structure-aware feature vector of time steps, For nodes In the The hidden state vector of time steps, is the hyperbolic tangent function, For nodes In the The modulus of the structure-aware feature vector at each time step, For nodes In the The modulus of the hidden state vector at time steps, is the cosine function, is the coefficient.

[0131] In a preferred embodiment of the present invention, a linkage fault identification layer is constructed to calculate the state correlation of different nodes based on the node state time series data set. When the state correlation of multiple nodes exceeds a preset correlation threshold, a fault propagation path is established between the nodes, a set of nodes with linkage influence is identified, and a multi-node linkage fault mode data set is generated, including:

[0132] Based on the node state time series dataset, the hidden state vectors of the two nodes at each time step are extracted respectively. The first similarity value is calculated by performing dot product and normalization on the two vectors. The change amplitude of the hidden state vectors of the two nodes at different time steps relative to the previous time step is calculated, and the difference between the change amplitudes of the two nodes is calculated to obtain the second similarity value.

[0133] Determine whether two nodes experience abnormal state fluctuations at the same time in the current time step. When both nodes are in abnormal state, the abnormal response coupling term of this time step takes a fixed value. When there is no abnormality at any node, the term is zero, and the third similarity value is obtained.

[0134] Calculate the state correlation between the state change trends of different nodes in the entire time series according to the first similarity value, the first similarity value, and the first similarity value to obtain a state correlation data set;

[0135] When the state correlation between two nodes exceeds the preset correlation threshold, it is determined that there is a fault coupling relationship between the two nodes, and a fault propagation path is established between the two nodes to obtain the linkage-affected node set;

[0136] According to the linkage impact node set and its propagation path, a multi-node linkage fault mode dataset is generated.

[0137] In an embodiment of the present invention, according to the node state time series data set, the hidden state vectors of the two nodes at each time step are extracted respectively, and the first similarity value is calculated by performing dot product and normalization processing on the two vectors; the change amplitude of the hidden state vectors of the two nodes at different time steps relative to the previous time step is calculated, and then the difference in the change amplitude of the two nodes is calculated to obtain the second similarity value; it is determined whether the two nodes have abnormal state fluctuations at the same time in the current time step. When both nodes are in abnormal state, the abnormal response coupling item of the time step takes a fixed value. When there is no abnormality in any node, the item is zero, and a third similarity value is obtained; according to the first similarity value, the first similarity value and the first similarity value, the second similarity value is obtained. The state correlation between the state change trends of different nodes in the entire time series is calculated to obtain a state correlation dataset, which can quantify and capture the temporal state change laws between different nodes. When the state correlation between two nodes exceeds the preset correlation threshold, it is determined that there is a fault coupling relationship between the two nodes, and a fault propagation path is established between the two nodes to obtain a set of linkage-affected nodes. A dynamic graph structure containing fault propagation paths can be constructed. Based on the set of linkage-affected nodes and their propagation paths, a multi-node linkage fault mode dataset is generated, which can effectively describe the fault propagation process and provide the system with more detailed fault mode information.

[0138] The calculation formula of the state correlation is:

[0139] ,

[0140] in, For nodes and nodes The state correlation between and is the index of the node, is the length of the entire time series, is the index of the time step, For nodes In the The hidden state vector of time steps, For nodes In the The hidden state vector of time steps, For nodes In the time step and the The magnitude of the change in the hidden state vector between time steps, , For the node time step and the The magnitude of the change in the hidden state vector between time steps, , For nodes In the The judgment value of whether abnormal state fluctuation occurs in a time step, when the node In the When abnormal state fluctuation occurs in a time step, the judgment value is 1, otherwise the judgment value is 0, For nodes In the The judgment value of whether abnormal state fluctuation occurs in the time step, the judgment value The value and judgment value of Similarly, is the coefficient.

[0141] in, is the weight coefficient, They correspond to the state similarity term, the state change trend term and the abnormal response coupling term respectively. The three together determine the comprehensive correlation measurement between nodes in the multi-time step state sequence. To ensure the effectiveness of each factor under different types of fault modes, the three weight coefficients must satisfy the constraint relationship , its specific value is dynamically set according to the system operating environment, node coupling characteristics and target detection fault type.

[0142] During the deployment of the TianQing system, if the main manifestation of the failure is that multiple nodes are abnormal at the same time, for example, a central node failure causes multiple downstream computing nodes to interrupt tasks, it is recommended to set is the dominant weight, The value is 0.2, The value is 0.2, A value of 0.2 is used to improve the ability to respond to sudden linkage anomalies. If the abnormalities of nodes in the system are often manifested as status indicators that continue to degrade over a long period of time with similar fluctuation trends, for example, multiple storage nodes gradually enter performance bottlenecks after the I / O load increases, it is appropriate to increase the proportion of trend consistency items. The value is 0.2, The value is 0.7, The value 0.1 is used to enhance the coupled modeling of state evolution trajectories. If the operation goal is to identify the synchronization of the operation status of the entire system, such as determining whether a computing node deviates from the global scheduling rhythm of the cluster, it is recommended to increase the proportion of the state similarity item. Take the value 0.6, Take the value 0.3, Set the value to 0.1 to strengthen the measurement effect of state consistency between nodes.

[0143] Among them, when the state correlation between two nodes exceeds the preset correlation threshold, it is determined that the two nodes have a fault coupling relationship, and a fault propagation path is established between the two nodes to obtain a set of linkage-affected nodes. Specifically, it includes:

[0144] When the state correlation between nodes exceeds the preset correlation threshold, the system will establish a fault propagation path between the two nodes. The establishment of the fault propagation path is based on the idea of graph theory. Each node is regarded as a vertex in the graph, and the fault coupling relationship between nodes is expressed as edges in the graph. These edges represent the fault propagation path, that is, the failure of one node may affect other related nodes. By establishing this propagation path based on the mutual influence between nodes, the system can fully understand how the fault propagates between nodes.

[0145] The establishment of a fault propagation path is not only to record the fault coupling relationship between two nodes, but also to help the system predict and identify potential fault expansion. When the system detects the existence of a fault propagation path between multiple nodes, it can issue a timely warning to alert the system administrator to the possible spread of the fault. By establishing this path, the system can accurately capture the linkage effect between each node, so as to take fault isolation or repair measures in advance and reduce the impact of the fault on the entire system.

[0146] In a preferred embodiment of the present invention, an initial multi-node fault management model is trained based on historical operation data, parameters of the initial multi-node fault management model are dynamically adjusted, and the linkage effects generated when each node fails are identified to obtain a multi-node fault management model, including:

[0147] The training set and validation set are obtained by randomly selecting node status data and performance monitoring data from the historical operation data set;

[0148] According to the fault log data, the fault propagation path between each node is extracted to obtain the actual fault mode;

[0149] Inputting the training set into the initial multi-node fault management model, obtaining the multi-node linkage failure mode predicted by the initial multi-node fault management model, and obtaining the training failure mode;

[0150] By comparing the differences between the training failure mode and the real failure mode, the training deviation path is obtained;

[0151] According to the training deviation path, the parameters in the initial multi-node fault management model are updated and adjusted to gradually reduce the training deviation path and obtain the training deviation path change rate;

[0152] Through multiple rounds of training iterations, when the training deviation path change rate is less than the preset change rate threshold, a multi-node fault management model is obtained.

[0153] In the embodiment of the present invention, the node status data and performance monitoring data in the historical operation data set are randomly selected to obtain the training set and the verification set. By randomly selecting the data set for training and verification, the overfitting phenomenon is effectively prevented, making the training process more representative; according to the fault log data, the fault propagation path between each node is extracted to obtain the real fault mode. By analyzing the fault log, the system can identify the dependency relationship between multiple nodes and the fault propagation path; the training set is input into the initial multi-node fault management model to obtain the multi-node linkage fault mode predicted by the initial multi-node fault management model, and the training fault mode is obtained. Predict multi-node linkage failure modes that may occur in the future; obtain the training deviation path by comparing the differences between the training failure mode and the actual failure mode, and by evaluating the training deviation path, provide a specific basis for subsequent model optimization; according to the training deviation path, update and adjust the parameters in the initial multi-node fault management model to gradually reduce the training deviation path, and obtain the training deviation path change rate, which can minimize the prediction error by continuously adjusting the model parameters during the training process; through multiple rounds of training iterations, when the training deviation path change rate is less than the preset change rate threshold, a multi-node fault management model is obtained, ensuring that the fault management model achieves the best performance.

[0154] According to the training deviation path, the parameters in the initial multi-node fault management model are updated and adjusted to gradually reduce the training deviation path, and the training deviation path change rate is obtained, which specifically includes:

[0155] First, after the model completes a prediction, it is necessary to compare the model's prediction results, namely the training failure mode, with the actual failure mode to quantify the difference between the two. The training failure mode refers to the multi-node linkage failure path output by the model after inputting the training set data, including the predicted fault node set and its mutual propagation relationship, while the real failure mode is the actual node failure relationship obtained by analyzing historical fault logs. In order to quantify the difference between the two, it is necessary to introduce a deviation measurement mechanism. A common practice is to use a graph structure similarity measurement method, such as cosine similarity based on the adjacency matrix, and use the intersection-union ratio at the node level to measure the matching degree of the fault node set, and comprehensively consider the accuracy of the node connection.

[0156] After the deviation calculation is completed, the system will design a loss function based on the deviation. The essence of the loss function is to convert the deviation into an optimizable objective function to reflect the distance between the current model prediction and the target. In this system, the loss function can be designed as a weighted combination of multiple losses, such as , represents the loss between the predicted set of fault nodes and the true set, which is obtained using the mean square error, and The structural similarity of the fault propagation path is measured, which is usually obtained by calculating the graph structure distance. The parameter is a hyperparameter used to balance the weights of the two types of losses.

[0157] In the actual training process, the initial setting is 0.7, is 0.3, which means that it focuses more on the accurate identification of node fault status while maintaining basic sensitivity to the propagation path structure. This setting is suitable for the early stage of the system, and its main task is to quickly lock the key fault node. In the experimental verification, the sample data set of the TianQing system, including about 100,000 node status samples and 1,200 groups of multi-node linkage failure samples, is used for training. When only node loss is used for supervision, that is, is 1.0, When it is 0, the fault node identification accuracy can reach 92.3%, but the path identification accuracy is only 68.1%; when the reverse deployment is 0.3, When is 0.7, although the path structure recovery accuracy can reach 81.6%, the node recognition rate drops to 85.2%. is 0.7, When the value is 0.3, the two accuracies are stable at 90.1% and 76.4% respectively, achieving the best comprehensive recognition performance. As the training rounds progress, if the system goal turns to improving the ability to understand the linkage propagation mechanism, it can be gradually improved. The proportion of is 0.5, When the value is 0.5, the model focuses on fault path relationship modeling, achieving more precise propagation path prediction and restoration, which is especially suitable for dealing with system scenarios with long fault impact chains and complex structural coupling. The value setting of should be dynamically adjusted in combination with the training objectives of the current system stage and the actual scenario requirements, and can be automatically optimized with the help of validation set accuracy feedback to achieve a balance between node-level and structure-level fault identification, thereby obtaining a multi-node fault management model with higher robustness and generalization capabilities.

[0158] After the loss function is calculated, the gradient of the parameters of each layer in the model with respect to the total loss can be calculated through the backpropagation algorithm. During this process, the system will propagate the error information from the output layer to the input layer layer by layer, and derive the weight matrix and bias terms in the model based on the chain rule. For example, in the graph convolutional network layer, the gradient will be propagated to the feature extraction layer of each node according to the static adjacency matrix, and further affect the weighted aggregation process of the features; in the recurrent neural network layer, the gradient will expand along the time dimension and propagate to the hidden state of each time step, thereby adjusting the memory and forgetting ability of the parameters in the time series modeling process.

[0159] Next, based on the calculated gradient information, the system uses an optimization algorithm to update the model parameters. Common optimization algorithms include stochastic gradient descent and Adam. The Adam optimizer is particularly suitable for training neural networks with time series and structured data. The basic formula for parameter update is: , represents the parameters to be optimized, is the time index, is the learning rate, It is the gradient of the parameter with respect to the loss function. Each parameter update is based on the gradient obtained by backpropagation of the current input features and labels of the training set, gradually approaching the optimal parameter configuration.

[0160] After completing a parameter update, the system will recalculate the training deviation path and compare it with the deviation of the previous round to calculate the training deviation path change rate. The training deviation path change rate is defined as the relative change between the deviations of the previous and next two rounds, and the formula is , For the Training deviation at the moment, For the -1 moment of training deviation, If the value is less than the preset threshold, the model is considered to have basically converged and the training can be ended; otherwise, the next round of training will be continued.

[0161] Through this process, the model can continuously optimize parameters in a data-driven manner, reduce training bias, and gradually improve prediction accuracy and generalization capabilities. The resulting multi-node fault management model not only has good fault node identification capabilities, but can also accurately restore the fault propagation path. It can be effectively applied to the actual operation of the TianQing system to achieve efficient management of complex linked faults.

[0162] In a preferred embodiment of the present invention, through multiple rounds of training iterations, when the training deviation path change rate is less than a preset change rate threshold, a multi-node fault management model is obtained, including:

[0163] After each round of training is completed, the validation set is input into the updated multi-node fault management model of the current round to obtain the verified fault mode;

[0164] Compare the verification failure mode with the actual failure mode to obtain the verification deviation path;

[0165] According to the verification deviation path, determine whether the verification deviation path of the current round of training is smaller than that of the previous round of training. If the verification deviation path of the preset number of consecutive rounds is not lower than the preset verification deviation threshold, the training process is terminated, otherwise the training is continued;

[0166] When the training deviation path change rate is less than the preset change rate threshold and the verification deviation path for a preset number of consecutive rounds is not less than the preset verification deviation threshold, the training process is terminated and a multi-node fault management model is obtained.

[0167] In an embodiment of the present invention, after each round of training is completed, the verification set is input into the multi-node fault management model updated in the current round to obtain a verification failure mode. By comparing and analyzing the verification failure mode, the deficiencies of the model in certain specific situations can be discovered, thereby providing a basis for subsequent optimization. The verification failure mode is compared with the actual failure mode to obtain a verification deviation path. By calculating the verification deviation path, it is possible to clearly know in which aspects the current model has deviations, the size of the deviation and its changing trend. According to the verification deviation path, it is determined whether the verification deviation path of the current round of training is reduced compared with the previous round of training. When the verification deviation path for a consecutive preset number of rounds is not lower than the preset verification deviation threshold, the training process is terminated, otherwise the training is continued, which can effectively avoid the occurrence of model overfitting during the training process. When the training deviation path change rate is less than the preset change rate threshold and the verification deviation path for a consecutive preset number of rounds is not lower than the preset verification deviation threshold, the training process is terminated to obtain a multi-node fault management model, ensuring the efficiency of the training process and avoiding excessive invalid training rounds.

[0168] In a preferred embodiment of the present invention, based on the fault status data set, the fault node and the affected nodes are isolated, the impact of the fault node on the affected nodes is eliminated, and the fault repair result is obtained, including:

[0169] Based on the fault status data set, determine whether the current operating status of the fault node meets the fault isolation conditions. When the abnormal state duration of the fault node exceeds the preset abnormal time threshold, an isolation signal is triggered;

[0170] According to the isolation signal, the data transmission between the faulty node and the affected node is cut off, and the data transmission path is reallocated to the affected node to obtain the reconstructed transmission state;

[0171] According to the reconstructed transmission status, confirm whether the affected node has resumed normal operation after isolating the faulty node. When the status of the affected node is normal, mark the isolation process as successful and generate a fault repair result.

[0172] In an embodiment of the present invention, based on the fault status data set, it is determined whether the current operating status of the fault node meets the fault isolation conditions. When the abnormal status duration of the fault node exceeds the preset abnormal time threshold, the isolation signal is triggered, which can automatically identify and isolate the fault node, avoiding the uncertainty and delay of manual operation; according to the isolation signal, the data transmission between the fault node and the affected node is cut off, and the data transmission path of the affected node is reallocated to obtain a reconstructed transmission state, which can quickly and effectively reduce the interference of the fault node on other system components, avoid data loss and the expansion of performance bottlenecks; according to the reconstructed transmission state, it is confirmed whether the affected node has resumed normal operation after isolating the fault node. When the status of the affected node is normal, the isolation processing is marked as successful, and a fault repair result is generated to ensure that the system can verify whether normal operation has been restored after fault isolation, and evaluate the repair effect through continuous monitoring.

[0173] According to the isolation signal, the data transmission between the faulty node and the affected node is cut off, and the data transmission path is reallocated to the affected node to obtain the reconstructed transmission state, which specifically includes:

[0174] The system interrupts the connection of the faulty node through the network switch or router, updates the network routing table, and ensures that the data flow is no longer transmitted through the path of the faulty node. The router will select a new data transmission path based on the newly configured network topology to prevent the abnormal status of the faulty node from continuing to affect other normally operating nodes.

[0175] After cutting off the data transmission between the faulty node and the affected nodes, the system needs to reallocate the transmission paths of the affected nodes. To ensure the normal operation of these affected nodes, the system will find the best available path for each affected node based on the current network topology. By using the network scheduling algorithm, it analyzes the connection status and load conditions between other nodes in the system in real time, and dynamically selects the best data transmission path that does not pass through the faulty node. This path reconstruction needs to consider multiple factors, such as path bandwidth, latency, load balancing, etc., to ensure that data can be transmitted to the target node in the shortest time and at the lowest cost.

[0176] After reconfiguring the transmission path, the system will update the routing tables of all relevant nodes according to the new network topology to ensure that the affected nodes can receive data through the new path. At the same time, the system will continue to monitor the status changes of these affected nodes to ensure that they can resume normal operation after isolating the faulty nodes. If any node is detected to have transmission problems on the new path, the system will adjust the path again to further optimize the data flow.

[0177] In a preferred embodiment of the present invention, the multi-node fault management model is updated according to the fault repair result to improve the accuracy of the multi-node fault management model in identifying the multi-node linkage fault mode, including:

[0178] Based on the fault repair results, the fault propagation path in the historical operation trajectory is marked with the fault node as the center to obtain a supplementary training set;

[0179] An incremental data set is obtained based on the supplementary training set and the training set, and is input into the multi-node fault management model for parameter fine-tuning to obtain an updated multi-node fault management model;

[0180] Perform a validation test on the updated multi-node fault management model. If its recognition accuracy on the validation set improves or does not decrease, retain the updated multi-node fault management model as the current valid model configuration.

[0181] When the recognition accuracy of the updated multi-node fault management model on the validation set tends to be stable and does not improve any further, the updating process is stopped to avoid overfitting.

[0182] In an embodiment of the present invention, based on the fault repair results, the fault propagation path of the fault node in the historical operation trajectory is marked with the fault node as the center to obtain a supplementary training set. By marking the propagation path of each fault node, the system can accurately reflect the expansion range of the fault and the characteristics of the affected nodes, providing an accurate data basis for subsequent model updates and fine-tuning. Based on the supplementary training set and the training set, an incremental data set is obtained and input into the multi-node fault management model for parameter fine-tuning to obtain an updated multi-node fault management model. By introducing new fault propagation path data, complex fault modes can be better identified during the training process. The updated multi-node fault management model is verified and tested. When its recognition accuracy on the verification set improves or does not decrease, the updated multi-node fault management model is retained as the current valid model configuration. Through the verification process, the system can effectively avoid the risk of model overfitting or performance degradation, ensuring that the final model has high practical application value. When the recognition accuracy of the updated multi-node fault management model on the verification set tends to be stable and does not increase further, the update process is stopped to avoid overfitting, ensure that the training process does not overfit, and ensure that the final model has high generalization ability.

[0183] The above is a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as within the scope of protection of the present invention.

Claims

1. A TianQing system fault management method based on deep learning, characterized in that: The method comprises: Obtain node status data, performance monitoring data, and fault log data for each node in the TianQing system to obtain a historical operation data set; Based on graph convolutional networks and recurrent neural networks, we analyze the dynamic interactions and fault propagation paths between different nodes, identify potential multi-node linkage failure modes, and build an initial multi-node fault management model, including: Construct a recurrent neural network layer to calculate the hidden state of each node in chronological order based on the set of structural association sequences, identify the state change relationship of each node between different time steps, and obtain the node state time series dataset; A linkage fault identification layer is constructed to calculate the state correlation of different nodes based on the node state time series data set. When the state correlation of multiple nodes exceeds the preset correlation threshold, a fault propagation path is established between the nodes, and the set of nodes with linkage influence is identified to generate a multi-node linkage fault mode data set, including: Based on the node state time series dataset, the hidden state vectors of the two nodes at each time step are extracted respectively. The first similarity value is calculated by performing dot product and normalization on the two vectors. The change amplitude of the hidden state vectors of the two nodes at different time steps relative to the previous time step is calculated, and the difference between the change amplitudes of the two nodes is calculated to obtain the second similarity value. Determine whether two nodes experience abnormal state fluctuations at the same time in the current time step. When both nodes are in abnormal state, the abnormal response coupling term of this time step takes a fixed value. When there is no abnormality at any node, the term is zero, and the third similarity value is obtained. According to the first similarity value, the second similarity value, and the third similarity value, the state correlation between the state change trends of different nodes in the entire time series is calculated to obtain a state correlation data set; When the state correlation between two nodes exceeds the preset correlation threshold, it is determined that there is a fault coupling relationship between the two nodes, and a fault propagation path is established between the two nodes to obtain the linkage-affected node set; Generate a multi-node linkage fault mode dataset based on the linkage impact node set and its propagation path; The initial multi-node fault management model is trained based on historical operation data, and the parameters of the initial multi-node fault management model are dynamically adjusted to identify the linkage effects caused by failures of each node, thereby obtaining a multi-node fault management model. Obtain the real-time operating data set of the TianQing system and input it into the multi-node fault management model to identify the current faulty node and affected nodes, determine the current multi-node linkage fault mode, and obtain the fault status data set; Based on the fault status data set, the faulty node and the affected nodes are isolated, the impact of the faulty node on the affected nodes is eliminated, and the fault repair result is obtained; The multi-node fault management model is updated according to the fault repair results to improve the accuracy of the multi-node fault management model in identifying multi-node linkage failure modes.

2. The TianQing system fault management method based on deep learning according to claim 1, characterized in that: Based on graph convolutional networks and recurrent neural networks, we analyze the dynamic interactions and fault propagation paths between different nodes, identify potential multi-node linkage failure modes, and build an initial multi-node fault management model, including: Construct an input layer to receive historical operation data sets, extract node status data and performance monitoring data of each node at different time steps, and perform feature extraction on the node status data and performance monitoring data of each node at different time steps to obtain node feature sequences; A graph convolutional network layer is constructed to receive node feature sequences and generate a static adjacency matrix based on the static structural relationship of data transmission between nodes in the TianQing system. The current feature vector of each node is then aggregated with the feature vectors of its adjacent nodes to obtain a set of structural association sequences.

3. The TianQing system fault management method based on deep learning according to claim 2, characterized in that: Based on graph convolutional networks and recurrent neural networks, we analyze the dynamic interactions and fault propagation paths between different nodes, identify potential multi-node linkage failure modes, and build an initial multi-node fault management model, including: Construct an output layer to receive and output the multi-node linkage failure mode dataset; The initial multi-node fault management model is constructed by connecting the input layer, graph convolutional network layer, recurrent neural network layer, linkage fault identification layer and output layer in sequence.

4. The TianQing system fault management method based on deep learning according to claim 3, characterized in that: Construct a graph convolutional network layer to receive node feature sequences and generate a static adjacency matrix based on the static structural relationship of data transmission between nodes in the TianQing system. Then, aggregate the current feature vector of each node with the feature vectors of its adjacent nodes to obtain a set of structural association sequences, including: Based on the static connection relationship between nodes in the TianQing system during data transmission, a static adjacency matrix is constructed. The static adjacency matrix is a symmetric matrix, in which each non-zero element indicates a direct interaction relationship between two nodes. According to the static adjacency matrix, the feature vectors of the current node and each adjacent node are transformed by matrix multiplication to obtain the structure-aware feature vector of the current node; By performing nonlinear activation on the structure-aware feature vector, the interaction structure characteristics of each node in the current time period are obtained; The interaction structural features of all nodes at different time steps are arranged and merged according to the node dimension and time dimension to form a structural association sequence dataset.

5. The TianQing system fault management method based on deep learning according to claim 4, characterized in that: Construct a recurrent neural network layer to calculate the hidden state of each node in chronological order based on the set of structural association sequences, identify the state change relationship of each node between different time steps, and obtain a node state time series dataset, including: Input the structure association sequence set into the recurrent neural network layer in chronological order, and initialize the initial hidden states of all nodes to the default initial states when the TianQing system just starts running; At each time step, the node structure perception feature vector of the current time period and the hidden state of the previous time step are used as input to calculate the hidden state of the current time step; By calculating the hidden state of all time steps, the hidden state of each node at each time step is determined to obtain the hidden state data set; Based on the hidden state data set, the hidden state difference of each node between adjacent time steps is calculated, and the state changes of each node between different time steps are analyzed based on the hidden state difference. When the hidden state difference between adjacent time steps exceeds the preset difference threshold, the node state is identified as a significant change and marked as an abnormal state fluctuation. According to the abnormal state fluctuations, the hidden state changes of each node between all time steps are extracted to obtain the node state time series dataset.

6. The TianQing system fault management method based on deep learning according to claim 5, characterized in that: The initial multi-node fault management model is trained based on historical operation data, and its parameters are dynamically adjusted to identify the linkage effects caused by failures at each node. The resulting multi-node fault management model includes: The training set and validation set are obtained by randomly selecting node status data and performance monitoring data from the historical operation data set; According to the fault log data, the fault propagation path between each node is extracted to obtain the actual fault mode; Inputting the training set into the initial multi-node fault management model, obtaining the multi-node linkage failure mode predicted by the initial multi-node fault management model, and obtaining the training failure mode; By comparing the differences between the training failure mode and the real failure mode, the training deviation path is obtained; According to the training deviation path, the parameters in the initial multi-node fault management model are updated and adjusted to gradually reduce the training deviation path and obtain the training deviation path change rate; Through multiple rounds of training iterations, when the training deviation path change rate is less than the preset change rate threshold, a multi-node fault management model is obtained.

7. The TianQing system fault management method based on deep learning according to claim 6, characterized in that: Through multiple rounds of training iterations, when the training deviation path change rate is less than the preset change rate threshold, a multi-node fault management model is obtained, including: After each round of training is completed, the validation set is input into the updated multi-node fault management model of the current round to obtain the verified fault mode; Compare the verification failure mode with the actual failure mode to obtain the verification deviation path; According to the verification deviation path, determine whether the verification deviation path of the current round of training is smaller than that of the previous round of training. If the verification deviation path of the preset number of consecutive rounds is not lower than the preset verification deviation threshold, the training process is terminated, otherwise the training is continued; When the training deviation path change rate is less than the preset change rate threshold and the verification deviation path for a preset number of consecutive rounds is not less than the preset verification deviation threshold, the training process is terminated and a multi-node fault management model is obtained.

8. The TianQing system fault management method based on deep learning according to claim 7, characterized in that: Based on the fault status data set, the faulty node and the affected nodes are isolated, the impact of the faulty node on the affected nodes is eliminated, and the fault repair results are obtained, including: Based on the fault status data set, determine whether the current operating status of the fault node meets the fault isolation conditions. When the abnormal state duration of the fault node exceeds the preset abnormal time threshold, an isolation signal is triggered; According to the isolation signal, the data transmission between the faulty node and the affected node is cut off, and the data transmission path is reallocated to the affected node to obtain the reconstructed transmission state; According to the reconstructed transmission status, confirm whether the affected node resumes normal operation after isolating the faulty node. If the status of the affected node is normal, mark the isolation process as successful and generate the fault repair result.

9. The TianQing system fault management method based on deep learning according to claim 8, characterized in that: Update the multi-node fault management model based on the fault repair results to improve the accuracy of the multi-node fault management model in identifying multi-node linkage fault modes, including: Based on the fault repair results, the fault propagation path in the historical operation trajectory is marked with the fault node as the center to obtain a supplementary training set; An incremental data set is obtained based on the supplementary training set and the training set, and is input into the multi-node fault management model for parameter fine-tuning to obtain an updated multi-node fault management model; Perform a validation test on the updated multi-node fault management model. If its recognition accuracy on the validation set improves or does not decrease, retain the updated multi-node fault management model as the current valid model configuration. When the recognition accuracy of the updated multi-node fault management model on the validation set tends to be stable and does not improve any further, the updating process is stopped to avoid overfitting.

Citation Information

Patent Citations

  • Autonomous real-time fault isolation method based on event log

    CN117640350A

  • Cloud computing node fault prediction method based on improved graph neural network

    CN118764395A