Fault diagnosis method for server
By building a server virtual model and using a two-way long and short-term memory network and dynamic fault knowledge graph, the problem of comprehensive analysis of multi-source data in server fault diagnosis is solved, efficient fault location and processing is achieved, and the accuracy and real-time diagnosis is improved.
Patent Information
- Application Number
- CN202510915613.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-03
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2045-07-03
AI Technical Summary
Existing server fault diagnosis technology lacks the ability to analyze multi-source data in a comprehensive way, and it is difficult to accurately locate the root causes of complex faults. It is inefficient and prone to misjudgment and misjudgment. It is unable to effectively capture the early features and propagation rules of the fault, resulting in lag in response.
Build a virtual server model, collect hardware underlying, operating system kernel and network interaction traffic data, extract time- and space-related fault characteristics through a two-way long and short-term memory network, combine dynamic fault knowledge graphs and graph neural networks for real-time analysis, and generate processing strategies.
Real-time processing of multi-source data is realized, the accuracy and real-timeness of fault diagnosis are improved, the error judgment rate is reduced, and predictive maintenance and intelligent decision-making of complex server systems are supported.
Smart Images

Figure CN120407269A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of server fault diagnosis, and particularly relates to a fault diagnosis method for servers. Background Art
[0002] At present, with the rapid development of information technology, servers, as the core carriers for data storage and processing, their stability and reliability are crucial for enterprise operations and network services. However, the server operating environment is complex, with frequent failures and diverse causes, and traditional fault diagnosis technologies have many limitations. Moreover, with the wide application of cloud computing and big data technologies, the scale of server clusters is continuously expanding and the architecture is becoming increasingly complex, and problems such as hardware failures, software vulnerabilities, and network attacks are intertwined.
[0003] Traditional server fault diagnosis is based on threshold judgment. Due to the lack of comprehensive analysis ability for multi-source data, it is difficult to accurately locate the root causes of complex faults. The diagnosis mode relying on manual experience is not only inefficient but also prone to missed judgments and misjudgments. In addition, the dynamic evolution characteristics of server faults require that the diagnosis technology has real-time and forward-looking capabilities, but existing technologies often cannot effectively capture the early characteristics and propagation laws of faults, resulting in a lag in fault response and causing serious economic losses.
[0004] Therefore, it is urgent to improve the existing server fault diagnosis methods to solve the technical problems of lacking comprehensive analysis ability for multi-source data, being difficult to accurately locate the root causes of complex faults, being inefficient, and being prone to missed judgments and misjudgments. Summary of the Invention
[0005] The purpose of the present invention is to provide a fault diagnosis method for servers to solve the technical problems of lacking comprehensive analysis ability for multi-source data, being difficult to accurately locate the root causes of complex faults, being inefficient, and being prone to missed judgments and misjudgments.
[0006] To solve the above technical problems, the technical solution adopted by the present invention is as follows: A fault diagnosis method for servers includes the following steps: S1: Build server virtual models at each node respectively, and establish the connection relationships between the virtual models based on the real-time connection relationships of each node; S2: Collect the hardware underlying data, operating system kernel data, application program running data, and network interaction traffic data of the server, and map each item of data to the corresponding virtual model in real time after preprocessing; S3: Obtain the historical fault diagnosis data and the correlation degrees between the preprocessed data items, and perform weighted fusion on each item of data based on the correlation degrees to obtain fusion data; S4: Construct a bidirectional long short-term memory network model and input the fused data based on the time series. Extract the fault features including time dependence and spatial correlation from the fused data through the bidirectional long short-term memory network, and convert them into fault feature vectors. S5: Construct a dynamic fault knowledge graph, and perform real-time dynamic updates based on real-time fault data and empirical rules. Input the fault feature vectors into the dynamic fault knowledge graph, and obtain the fault probability distribution through inference calculation by the graph neural network. S6: Evaluate the fault type, fault level, fault location and probability according to the fault probability distribution through a preset fault assessment model and visualize them in the corresponding virtual model, and evaluate the fault triggering results based on the connection relationship. S7: Analyze the multi-dimensional behavior deviations of the virtual model in real time through the neural network, generate a processing strategy according to the behavior deviations and the fault assessment results, and execute it.
[0007] Preferably, the specific process of constructing the server virtual model at each node in step S1 is as follows: S11: Obtain the point cloud data of the server's shell and internal specified components, model the standardized components through the parametric model library, and construct the three-dimensional model of the server based on the point cloud data and the standardized component modeling. S12: Perform equivalent circuit modeling in the three-dimensional model. The equivalent circuit modeling includes constructing a power factor correction model, a DC-DC conversion model, and a filter circuit module in the three-dimensional model. S13: Build a three-dimensional thermal resistance network based on the heat conduction equation to establish a CPU / GPU thermal model in the three-dimensional model. The nodes of the three-dimensional thermal resistance network correspond to chip cores, packages, and radiator components. S14: Establish a virtual mechanical hard disk to simulate the vibration characteristics and power consumption curves of the spindle motor, the head arm, and the disk platter, and establish a temperature-performance-life association model for the flash memory chip and the controller, so as to realize the establishment of a storage system model in the three-dimensional model.
[0008] Preferably, the specific process of real-time mapping the preprocessed data to the corresponding virtual model in step S2 is as follows: S21: Unify the data into a specified format, establish a time series data set for the data in ascending order of time, identify the abnormal data in the time series data set, and delete or replace the abnormal data. S22: Set a sliding window of a specified size, calculate the mean and standard deviation of the specified data within the sliding window within a specified time, perform Fourier transform on the specified data, and extract the data values within a specified interval. S23: Establish a hierarchical transmission strategy, and split the time series data set into different data subsets, including a real-time control data set, a log data set, and a file-type data set. The real-time control data set is transmitted through MQTT, the log data set is transmitted asynchronously through Kafka, and the file-type data set is encrypted and transmitted through FTP / SFTP; S24: Establish a hierarchical mapping model and set the structural mapping relationship between the physical data and the virtual model. Direct mapping is performed on the real-time control data set and the file-type data set, and filtered mapping is performed on the log data set.
[0009] Preferably, the specific process of filtered mapping of the log data set in step S24 is as follows: S241: Filter out invalid logs through regular expressions, predict normal log patterns through a preset machine learning model, filter out deviation values, and define log semantic rules based on domain knowledge to remove redundant information; S242: Extract key events, scope of influence, and timestamps from the filtered logs to generate an event chain; S243: Map the event chain to the log node of the virtual model, store it in time series, and associate it with the state change of the virtual model; S244: Associate the log events with the specified data of the virtual model to form a complete fault traceability chain.
[0010] Preferably, the specific process of obtaining the correlation degree between the preprocessed data items in step S3 and performing weighted fusion on the data items based on the correlation degree to obtain the fusion data is as follows: S31: Split the preprocessed data into different data blocks, screen the specified features in each data block, and construct a feature set for each data block; S32: Construct feature combinations based on the feature set, allocate the feature combinations to different computing nodes, each node processes several feature pairs, and calculate the correlation degree of all feature combinations in parallel to generate a structured correlation matrix; S33: Incrementally update the correlation matrix using a sliding window and perform normalization processing on the correlation matrix.
[0011] Preferably, the specific process of calculating the correlation degree of all feature combinations in parallel in step S32 is as follows: S321: Let the feature combination be ( X i , X j ), calculate X i and X j 's marginal frequency and joint frequency. The formula for calculating the marginal frequency is as follows: ; Among them, count (·) represents the number satisfying the condition in the dataset, represents the feature X i when taking the value x i k appears, i, j, k, l are different serial numbers, N is the global sample number; X i take x i k and X j take x i l The joint probability calculation formula is as follows: ; Among them, X j = x i l represents the feature X j when taking the value x i l appears, N is the global sample number; S322: Calculate the local mutual information and the global mutual information: The local mutual information calculation formula is as follows: I local ( X i ; Y j ) = ∑ x,y ( count ( x , y ) / N block ) log count ( x , y ) · N block / count ( x ) · count ( y )]; x and y are different feature sets, X i is the feature set xThe features in N block is the number of block samples; The formula for global mutual information is as follows: I ( X i ; Y j ) = ∑ x,y ( count ( x , y ) / N ) log count ( x , y ) · N / count ( x ) · count ( y )]; N is the number of global samples.
[0012] Preferably, the bidirectional long short-term memory network is provided with dual input channels, one input channel is connected to the forward LSTM layer, and the other input channel is connected to the reverse LSTM layer. The output ends of the forward LSTM layer and the reverse LSTM layer are both connected to the merging layer, the merging layer is connected to the fully connected layer, and the fully connected layer is connected to the output layer.
[0013] Preferably, the process of extracting the fault features including time dependence and spatial correlation from the fusion data by the bidirectional long short-term memory network and converting them into fault feature vectors is as follows: S41: Input the fusion data into the forward LSTM layer and the reverse LSTM layer respectively through the dual input channels. The forward LSTM layer processes the fusion data from t = 0 to t = T to capture historical features; the reverse LSTM layer processes the fusion data from t = T to t = 0 to capture future features; S42: Learn the spatial relationship of the feature parts through the weight matrix between the forward LSTM layer and the reverse LSTM layer; S43: Extract the hidden state at the last time step: , h t is the merged hidden state vector, which is used to extract the global features of the time series data. and are the hidden states of the forward LSTM layer and the reverse LSTM layer at the time step t respectively; S44: Map to a fault feature vector of a fixed dimension through a fully connected layer: v = W·h t +b , where W is the weight matrix, b is the bias term.
[0014] Preferably, the specific process of step S5 is as follows: S51: Define entities including equipment components, fault types, fault symptoms, repair measures, and association rules, as well as their relationships and key attributes; S52: Use real-time fault data to train an incremental learning model, predict new fault association relationships, and add them to the knowledge graph after filtering through a threshold; S53: Perform dimensionality reduction and normalization processing on the fault feature vector, and calculate the cosine similarity between the fault feature vector and the embedding vector of the fault type node in the knowledge graph e i : ; S54: Construct an attention model f ( h t , e i ), output the weights of the features for each entity, and generate a soft mapping probability: ; S55: Use the fault node mapped from the fault feature vector as a seed node through the input layer of the graph neural network, initialize the fault probability of the seed node p i = 1, and for other nodes p j = 0. The node passes the fault influence message to its neighbors, and the weight is determined by the confidence of the edge. After probability aggregation, the probability distribution is output by the output layer.
[0015] Preferably, the specific process of step S7 is as follows: S71: Obtain multi-source data in real time from the virtual model, including physical dimensions: sensor data such as temperature, pressure, vibration frequency, current, and voltage; logical dimensions: model state transition probability, algorithm execution efficiency, parameter convergence, etc.; spatial dimensions: topological structure associated with the equipment location, signal transmission delay, spatial coupling effect, etc.; S72: Establish a normal behavior baseline: Define the normal range through unsupervised learning, and calculate deviation indicators: absolute deviation: |current value - baseline value|; dynamic deviation rate: (current value - baseline value) / baseline value × 100%; spatio-temporal correlation deviation: Combine the equipment location and the signal propagation path to calculate the deviation propagation coefficient in the spatial dimension; S73: Map the behavior deviation data into the initial amplitude, encode the real - valued features as probability amplitudes through embedding, capture the correlations of multi - dimensional deviations, simulate the non - linear coupling effect of complex faults, and set up a parameterized circuit. Optimize the circuit parameters through variational algorithms to adapt to the dynamic changes of real - time data; S74: Map the correlation between the probability amplitude and the multi - dimensional deviation into the priority score of the processing strategy, and combine the fault evaluation results to screen the optimal strategy.
[0016] The beneficial effects of the present invention include: The fault diagnosis method for servers provided by the present invention realizes efficient fault location and processing by constructing a multi - dimensional intelligent diagnosis system. Build server virtual models at each node and establish connection relationships, collect various data and map them to the virtual models; perform weighted fusion of data through correlation analysis, and use bidirectional long - short - term memory networks to extract fault features containing temporal and spatial correlations; infer the fault probability distribution with the help of a dynamic fault knowledge graph and a graph neural network, and combine the virtual model to visualize the fault evaluation results; finally, analyze the behavior deviation through a neural network and generate a processing strategy. It realizes the real - time processing of multi - source data, the intelligent extraction of fault features and the dynamic deduction of fault propagation, effectively improves the accuracy, real - time performance and automation level of server fault diagnosis, and realizes the predictive maintenance and intelligent decision - making of complex server systems.
[0017] First, by collecting data at the hardware bottom layer, system kernel, applications and network traffic, construct a monitoring system covering the entire server stack, avoid the diagnostic blind spots caused by single - dimensional data, perform weighted fusion of multi - source data based on correlation analysis, highlight key features, suppress noise interference, and thus effectively improve the accuracy of fault recognition. Second, simultaneously process the time - series data through forward and backward LSTM layers, synchronously capture the historical trends before the fault and the future patterns after the fault, and realize the deduction of the fault probability distribution through a dynamic knowledge graph combined with a graph neural network, effectively reducing the prediction error of fault propagation.
[0018] Finally, integrate the fault type, level and propagation probability into the virtual model. The operation and maintenance personnel can quickly locate the key fault nodes through the interactive interface, thereby effectively improving the decision - making efficiency, and continuously enhance the new hardware fault diagnosis ability based on real - time fault data and experience rules to dynamically update the knowledge graph. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 It is a schematic flow chart of the fault diagnosis method for servers of the present invention.
[0020] Figure 2 It is a schematic architecture diagram of the bidirectional long - short - term memory network of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0021] The following will further elaborate on the present invention in conjunction with the appended Figures 1 - 2 drawings: Embodiment 1 Refer to the appended Figure 1 drawings. A fault diagnosis method for a server includes the following steps: S1: Build server virtual models, three-dimensional models, equivalent circuit models, thermal models, and virtual mechanical hard disks respectively at each node, and establish the connection relationships between the virtual models based on the real-time connection relationships of each node.
[0022] S2: Collect the hardware low-level data, operating system kernel data, application program running data, and network interaction traffic data of the server, and map the preprocessed data to the corresponding virtual models in real time. The hardware low-level data includes the basic information of hardware components, hardware operation status data, and performance monitoring data. The operating system kernel data includes process and thread data, memory management data, storage and file system data, and system scheduling and security data. The application program running data includes program basic running data, business logic data, performance and exception data. The network interaction traffic data includes network interface and protocol data, traffic performance and security data, and connection and session data.
[0023] S3: Obtain the historical fault diagnosis data and the correlation degrees between the preprocessed data items, and perform weighted fusion on the data items based on the correlation degrees to obtain the fusion data.
[0024] S4: Build a bidirectional long short-term memory network model and input the fusion data based on the time series. Extract the fault features including time dependence and spatial correlation of the fusion data through the bidirectional long short-term memory network, and convert them into fault feature vectors. Refer to Figure 2 , the bidirectional long short-term memory network is provided with two input channels. One input channel is connected to the forward LSTM layer, and the other input channel is connected to the reverse LSTM layer. The output ends of the forward LSTM layer and the reverse LSTM layer are both connected to the merging layer, the merging layer is connected to the fully connected layer, and the fully connected layer is connected to the output layer.
[0025] S5: Build a dynamic fault knowledge graph, and perform real-time dynamic update based on real-time fault data and experience rules. Input the fault feature vectors into the dynamic fault knowledge graph, and obtain the fault probability distribution through graph neural network inference calculation; S6: Evaluate the fault type, fault level, fault location, and probability according to the fault probability distribution through a preset fault evaluation model and visualize them in the corresponding virtual models, and evaluate the fault triggering results based on the connection relationships; S7: Analyze the multi-dimensional behavior deviation of the virtual model in real time through a neural network, generate a processing strategy according to the behavior deviation and the fault evaluation result, and execute it.
[0026] In this embodiment, the specific process of constructing the server virtual model at each node in step S1 is as follows: S11: Obtain the point cloud data of the shell and internal specified components of the server, model the standardized components through a parametric model library, and construct a three-dimensional model of the server based on the point cloud data and the standardized component modeling; S12: Perform equivalent circuit modeling in the three-dimensional model. The equivalent circuit modeling includes constructing a power factor correction model, a DC-DC conversion model, and a filter circuit module in the three-dimensional model; S13: Based on the heat conduction equation, construct a three-dimensional thermal resistance network to establish a CPU / GPU thermal model in the three-dimensional model. The nodes of the three-dimensional thermal resistance network correspond to chip cores, packages, and radiator components; S14: Establish a virtual mechanical hard disk to simulate the vibration characteristics and power consumption curves of the spindle motor, head arm, and disk, and establish a temperature-performance-life association model for flash memory chips and controllers, so as to realize the establishment of a storage system model in the three-dimensional model.
[0027] In the present invention, thousands of servers are deployed in the cloud computing data center of an Internet company to form a distributed storage and computing cluster. A scanner is used to perform layer-by-layer scanning on the server shell and internal components to obtain point cloud data. Shell scanning: Scan the surface of the chassis at a specified point spacing length, and focus on capturing the geometric features of ventilation holes and heat dissipation grilles; after disassembling the server, scan non-standard components such as the CPU radiator and hard disk bracket, with a scanning angle covering 360° and a point cloud density of 10 points / mm. Use specified software to remove noise, and fit basic geometries such as planes and cylinders through the RANSAC algorithm. Based on the measured parameters of the server power supply, construct a totem pole PFC circuit including a main inductor, equivalent series resistance, switching tube, on-resistance, and control IC. Package the PFC circuit model into a recognizable VTK format and locate it to the three-dimensional space coordinates of the power module. Define thermal resistance nodes and their corresponding components and thermal resistance parameters, set the heat conduction equation and solve it.
[0028] Through the combination of point cloud scanning and parametric modeling, a three-dimensional geometric model of the server is constructed, and equivalent circuit, thermal resistance network, and storage multi-physical field models are embedded. The model accuracy meets the requirements of fault diagnosis, providing a quantitative basis for heat dissipation optimization, hardware selection, and fault prediction in the cloud computing data center. Subsequently, real-time sensor data can be further integrated to realize the dynamic update of the server virtual model, improving the real-time performance and accuracy of fault diagnosis.
[0029] Embodiment 2 Based on Embodiment 1, the specific process of step S2 for preprocessing each item of data and then real-time mapping it to the corresponding virtual model is as follows: S21: Unify each item of data into a specified format, perform timestamp calibration, synchronize all data to UTC through NTP, establish a time series data set for each item of data in ascending order of time, identify abnormal data in the time series data set, and delete or replace the abnormal data among them. For continuous data, use the IQR method to calculate the first quartile Q1 and the third quartile Q3, and the abnormal threshold: lower limit = Q1 - 1.5×IQR, upper limit = Q3 + 1.5×IQR. For discrete data, statistically analyze the frequency distribution, and determine the low-frequency values lower than the 3σ principle as abnormal. For single-point anomalies among them, use the weighted average of the previous and next 5 valid points for replacement, and mark continuous anomalies as invalid data and trigger an alarm.
[0030] S22: Set a sliding window of a specified size, calculate the mean and standard deviation of the specified data within the sliding window within a specified time, perform Fourier transform on the specified data, transform it to the frequency domain, and extract the data values within a specified interval.
[0031] S23: Establish a hierarchical transmission strategy, divide the time series data set into different data subsets, including real-time control data set, log data set, and file type data set. The real-time control data set is transmitted through MQTT, the log data set is transmitted asynchronously through Kafka, and the file type data set is encrypted and transmitted through FTP / SFTP; S24: Establish a hierarchical mapping model and set the structure mapping relationship between physical data and the virtual model, directly map the real-time control data set and the file type data set, and perform filtering and then mapping on the log data set.
[0032] The specific process of filtering and then mapping the log data set in step S24 is as follows: S241: Filter out invalid logs through regular expressions, predict normal log patterns through a preset machine learning model, filter and count deviation values, and define log semantic rules based on domain knowledge to remove redundant information; S242: Extract key events, scope of influence, and timestamps from the filtered logs to generate an event chain, parse timestamps, event types, and scope of influence from the logs, sort events by timestamps, and associate causal relationships.
[0033] S243: Map the event chain to the log node of the virtual model, store it in time series, and associate the state changes of the virtual model. Map each event in the event chain to the corresponding component of the virtual model, and the time series relationship needs to be retained during storage, and the state of the virtual model is dynamically updated.
[0034] S244: Associate the log events with the specified data of the virtual model to form a complete fault traceability chain. Associating the log events with the monitoring metrics and configuration change records of the virtual model, the change records in the configuration management database can be associated through the device restart events in the logs. Kibana or Grafana can be used to display the event chain, supporting timeline backtracking and root cause localization.
[0035] Embodiment 3 Based on Embodiment 1 or Embodiment 2, the specific process of obtaining the correlation degrees between the preprocessed data items in step S3 and performing weighted fusion on the data items based on the correlation degrees to obtain the fusion data is as follows: S31: Split the preprocessed data into different data blocks, screen the specified features in each data block, and construct the feature set of each data block. Here, it is evenly split based on the number of records to ensure that each data block contains the same number of samples, realizing the segmented processing of time series data. In a distributed system, the splitting is automatically performed by the scheduler, and the split size needs to balance the load, avoiding overly large blocks causing memory overflow and overly small blocks increasing communication overhead. Usually, the block size is set to 64MB - 128MB.
[0036] Apply feature screening to each data block independently, remove redundant or irrelevant features, and focus on key variables. The screening method combines domain knowledge and statistical metrics, and quickly evaluates the importance of features based on statistical metrics. Calculate the correlation between the features and the target variable, and retain the TOP-K features. If the feature distributions between data blocks are significantly different, the statistics of each block need to be calculated independently to adapt to local characteristics, and a threshold is set to filter out noise features; when dealing with categorical features, they need to be encoded and then screened. Recombine the screened features in the original order, remove the discarded columns, and generate a refined data set. For example, screen out the specified number of key features from the original specified number of features to construct a new data block. Convert to a unified format to ensure consistent feature dimensions, add identifiers to the feature sets of each data block, and store them in a columnar format to optimize query efficiency.
[0037] S32: Construct feature combinations based on the feature sets, allocate the feature combinations to different computing nodes, and each node processes several feature pairs to calculate the correlation degrees of all feature combinations in parallel, generating a structured correlation matrix.
[0038] The specific process of calculating the correlation degrees of all feature combinations in parallel in step S32 is as follows: S321: Let the feature combination be ( X i , X j ), calculate X i and X j 's marginal frequencies and joint frequencies. The formula for calculating the marginal frequencies is as follows: ; wherein, count (·) represents the number satisfying the condition in the dataset, represents the feature X i when taking the value of x i k appears, i, j, k, l is a different serial number, N is the global sample number; X i takes x i k and X j takes x i l The joint probability calculation formula is as follows: ; wherein, X j = x i l represents the feature X j when taking the value of x i l appears, N is the global sample number; S322: Calculate the local mutual information and the global mutual information: The local mutual information calculation formula is as follows: I local ( X i ; Y j ) = ∑ x,y ( count ( x , y ) / N block ) log count ( x , y ) · N block / count ( x ) · count ( y )]; x [[ID=IOO]]and y are different feature sets, Xi is the feature set x in the features N block is the number of block samples; The global mutual information calculation formula is as follows: I ( X i ; Y j ) = ∑ x,y ( count ( x , y ) / N ) log count ( x , y ) · N / count ( x ) · count ( y )]; N is the global number of samples.
[0039] S33: The association matrix is incrementally updated using a sliding window, and the association matrix is normalized.
[0040] In this embodiment, the specific process of extracting the fault features including time dependence and spatial association from the fusion data through the bidirectional long short-term memory network in step S4 and converting them into fault feature vectors is as follows: S41: The fusion data is respectively input into the forward LSTM layer and the reverse LSTM layer through dual input channels. The forward LSTM layer processes the fusion data from t = 0 to t = T to capture historical features; the reverse LSTM layer processes the fusion data from t = T to t = 0 to capture future features.
[0041] The fusion data is time series data with a shape of (T, D), where T is the number of time steps and D is the feature dimension. The forward LSTM layer processes in order from t = 0 to t = T, and calculates the hidden state h → t , depending on the previous moment state h → t -1 and the current input. The processing order of the reverse LSTM layer: from t = T to t = 0, calculate the hidden state h ← t , depending on the next moment state h ← t + 1 and the current input.
[0042] S42: Learn the spatial relationship of the feature parts through the weight matrix between the forward LSTM layer and the backward LSTM layer. The hidden states of the forward LSTM layer and the backward LSTM layer interact through a shared or independent weight matrix to learn the spatial dependence in the time series data, and jointly optimize the weight matrix of the bidirectional LSTM through backpropagation to minimize the loss function.
[0043] S43: Extract the hidden state at the last time step: , h t is the merged hidden state vector, which is used to extract the global features of the time series data. and are the hidden states of the forward LSTM layer and the backward LSTM layer at time step t respectively; S44: Map it to a fault feature vector with a fixed dimension through a fully connected layer: , where W is the weight matrix, b is the bias term.
[0044] Embodiment 4 Based on Embodiment 1 or Embodiment 2 or Embodiment 3, the specific process of step S5 is as follows: S51: Define entities including equipment components, fault types, fault symptoms, maintenance measures, and association rules, as well as their relationships and key attributes. Equipment components include component ID, name, model, and installation location, which are used to identify the physical composition of the equipment. Fault types include fault codes, severity levels, and impact scopes.
[0045] S52: Use real-time fault data to train an incremental learning model to predict new fault association relationships, and add them to the knowledge graph after filtering through a threshold. The incremental learning model uses a random forest model, which inputs the fault feature vector and outputs the predicted fault association probability. Only the association relationships with probabilities higher than the threshold are retained to avoid noise interference. Add new nodes and edges through a graph database.
[0046] S53: Perform dimensionality reduction and normalization processing on the fault feature vector, and calculate the cosine similarity between the fault feature vector h t and the embedding vector e i of the fault type node in the knowledge graph: ; S54: Construct an attention model f ( h t , e i) Output the weights of the feature pairs for each entity and generate the soft mapping probability: ; S55: Use the fault nodes obtained by mapping the fault feature vectors through the input layer of the graph neural network as seed nodes, and initialize the fault probabilities of the seed nodes p i = 1, for other nodes p j = 0. The nodes pass the fault influence messages to their neighbors. The weights are determined by the confidence of the edges. After probability aggregation, the output layer outputs the probability distribution.
[0047] The specific process of step S7 is as follows: S71: Obtain multi-source data from the virtual model in real time, including physical dimensions: sensor data such as temperature, pressure, vibration frequency, current and voltage; logical dimensions: model state transition probability, algorithm execution efficiency, parameter convergence, etc.; spatial dimensions: topological structure associated with the device location, signal transmission delay, spatial coupling effect, etc. S72: Establish a normal behavior baseline: Define the normal range through unsupervised learning and calculate the deviation indicators: absolute deviation: |current value - baseline value|; dynamic deviation rate: (current value - baseline value) / baseline value × 100%; spatio-temporal correlation deviation: Combine the device location and the signal propagation path to calculate the deviation propagation coefficient in the spatial dimension. S73: Map the behavior deviation data to the initial amplitude, encode the real-valued features as probability amplitudes through embedding, capture the correlations of multi-dimensional deviations, simulate the non-linear coupling effect of complex faults, and set up a parameterized circuit to optimize the circuit parameters through variational algorithms to adapt to the dynamic changes of real-time data. S74: Map the correlations between the probability amplitudes and the multi-dimensional deviations to the priority scores of the processing strategies, and screen the optimal strategies in combination with the fault evaluation results.
[0048] In summary, the fault diagnosis method for servers provided by the present invention realizes efficient fault location and processing by constructing a multi-dimensional intelligent diagnosis system. Build a virtual model of the server and establish connection relationships at each node, collect various data and map them to the virtual model; perform weighted fusion of the data through correlation analysis, and use a bidirectional long short-term memory network to extract fault features including time and space correlations; infer the fault probability distribution with the help of a dynamic fault knowledge graph and a graph neural network, and visualize the fault evaluation results in combination with the virtual model; finally, analyze the behavior deviations through a neural network and generate processing strategies. It realizes the real-time processing of multi-source data, the intelligent extraction of fault features and the dynamic deduction of fault propagation, effectively improves the accuracy, real-time performance and automation level of server fault diagnosis, and realizes the predictive maintenance and intelligent decision-making of complex server systems.
Claims
1. A fault diagnosis method for a server, characterized in that, It includes the following steps: S1: Build server virtual models at each node respectively, and establish the connection relationships between the virtual models based on the real-time connection relationships of each node; S2: Collect the hardware low-level data, operating system kernel data, application program running data and network interaction traffic data of the server, and map each item of data to the corresponding virtual model in real time after preprocessing; S3: Obtain the historical fault diagnosis data and the correlation degrees between the preprocessed data items, and perform weighted fusion on each item of data based on the correlation degrees to obtain the fusion data; S4: Build a bidirectional long short-term memory network model and input the fusion data based on the time series. Extract the fault features including time dependence and spatial correlation of the fusion data through the bidirectional long short-term memory network, and convert them into fault feature vectors; S5: Build a dynamic fault knowledge graph, and perform real-time dynamic update based on real-time fault data and experience rules. Input the fault feature vector into the dynamic fault knowledge graph, and obtain the fault probability distribution through graph neural network inference calculation; S6: Evaluate the fault type, fault level, fault location and probability according to the fault probability distribution through a preset fault evaluation model and visualize them in the corresponding virtual model, and evaluate the fault triggering results based on the connection relationship; S7: Analyze the multi-dimensional behavior deviations of the virtual model in real time through a neural network, and generate and execute a processing strategy according to the behavior deviations and the fault evaluation results.
2. The fault diagnosis method for a server according to claim 1, characterized in that, The specific process of building server virtual models at each node in step S1 is as follows: S11: Obtain the point cloud data of the shell and internal specified components of the server, model the standardized components through a parametric model library, and build a three-dimensional model of the server based on the point cloud data and the standardized component modeling; S12: Perform equivalent circuit modeling in the three-dimensional model. The equivalent circuit modeling includes building a power factor correction model, a DC-DC conversion model, and a filter circuit module in the three-dimensional model; S13: Build a three-dimensional thermal resistance network based on the heat conduction equation, and establish a CPU / GPU thermal model in the three-dimensional model. The nodes of the three-dimensional thermal resistance network correspond to chip cores, packages, and radiator components; S14: Establish a virtual mechanical hard disk to simulate the vibration characteristics and power consumption curves of the spindle motor, head arm, and disk, and establish a temperature-performance-life association model for flash memory chips and controllers, so as to realize the establishment of a storage system model in the three-dimensional model.
3. A fault diagnosis method for a server according to claim 1, characterized in that The specific process of mapping each item of data to the corresponding virtual model in real time after preprocessing in step S2 is as follows: S21: Unify each item of data into a specified format, establish a time series data set for each item of data in ascending order of time, identify the abnormal data in the time series data set, and delete or replace the abnormal data; S22: Set a sliding window of a specified size, calculate the mean and standard deviation of the specified data within the sliding window within a specified time, perform Fourier transform on the specified data, and extract the data values within a specified interval; S23: Establish a hierarchical transmission strategy, split the time series data set into different data subsets, including a real-time control data set, a log data set, and a file-type data set. The real-time control data set is transmitted via MQTT, the log data set is transmitted asynchronously via Kafka, and the file-type data set is encrypted and transmitted via FTP / SFTP; S24: Establish a hierarchical mapping model and set the structural mapping relationship between the physical data and the virtual model. Perform direct mapping on the real-time control data set and the file-type data set, and perform filtered mapping on the log data set.
4. A fault diagnosis method for a server according to claim 3, characterized in that, The specific process of performing filtered mapping on the log data set in step S24 is as follows: S241: Filter out invalid logs through regular expressions, predict normal log patterns through a preset machine learning model, filter out deviation values, and define log semantic rules based on domain knowledge to remove redundant information; S242: Extract key events, scope of influence, and timestamps from the filtered logs to generate an event chain; S243: Map the event chain to the log node of the virtual model, store it in time series, and associate it with the state change of the virtual model; S244: Associate the log events with the specified data of the virtual model to form a complete fault traceability chain.
5. A fault diagnosis method for a server according to claim 1, characterized in that, The specific process of obtaining the correlation degree between the preprocessed data items in step S3 and performing weighted fusion on the data items based on the correlation degree to obtain the fusion data is as follows: S31: Split the preprocessed data into different data blocks, screen the specified features in each data block, and construct the feature set of each data block; S32: Construct feature combinations based on the feature set, allocate the feature combinations to different computing nodes, each node processes several feature pairs, and calculate the correlation degree of all feature combinations in parallel to generate a structured correlation matrix; S33: Incrementally update the correlation matrix using a sliding window and perform normalization processing on the correlation matrix.
6. The fault diagnosis method for a server according to claim 5, wherein, The specific process of calculating the correlation degree of all feature combinations in parallel in step S32 is as follows: S321: Let the feature combination be ( X i , X j ), calculate the X i and X j edge frequencies and joint frequencies. The formula for calculating the edge frequencies is as follows: ; Among them, count (·) represents the number that meets the conditions in the dataset, represents the feature X i when the value is x i k appears, i, j, k, l are different serial numbers, N is the global number of samples; X i Take x i k And X j Take x i l The joint probability calculation formula of ; Among them, X j = x i l represents a feature X j when the value is x i l appears on, N which is the global number of samples; S322: Calculate the local mutual information and the global mutual information: The formula for calculating the local mutual information is as follows: I local ( X i ; Y j ) = ∑ x,y ( count ( x , y ) / N block ) log count ( x , y ) · N block / count ( x ) · count ( y )]; x and y are different feature sets, X i is a feature in the feature set x ; N block is the number of block samples; The formula for calculating the global mutual information is as follows: I ( X i ; Y j ) = ∑ x,y ( count ( x , y ) / N ) log count ( x , y ) · N / count ( x ) · count ( y )]; N is the global number of samples.
7. A fault diagnosis method for a server according to claim 1, characterized in that The bidirectional long short-term memory network is provided with two input channels. One input channel is connected to the forward LSTM layer, and the other input channel is connected to the reverse LSTM layer. The output ends of the forward LSTM layer and the reverse LSTM layer are both connected to the merging layer, and the merging layer is connected to the fully connected layer, and the fully connected layer is connected to the output layer.
8. A fault diagnosis method for a server according to claim 7, characterized in that, The specific process of extracting the fault features including time dependence and spatial correlation from the fusion data through the bidirectional long short-term memory network in step S4 and converting them into fault feature vectors is as follows: S41: Input the fusion data into the forward LSTM layer and the backward LSTM layer respectively through the dual-input channels. The forward LSTM layer processes the fusion data from t = 0 to t = T to capture historical features; The reverse LSTM layer processes the fusion data from t =T to t =0 to capture future features; S42: Learn the spatial relationship of the feature events through the weight matrix between the forward LSTM layer and the reverse LSTM layer; S43: Extract the hidden state of the last time step; S44: Map it to a fault feature vector with a fixed dimension through the fully connected layer.
9. The fault diagnosis method for a server according to claim 8, wherein The specific process of step S5 is as follows: S51: Define entities including device components, fault types, fault symptoms, repair measures, as well as their relationships and key attributes, and the association rules; S52: Use real-time fault data to train an incremental learning model, predict new fault association relationships, and add them to the knowledge graph after filtering by a threshold; S53: Perform dimensionality reduction and normalization on the fault feature vectors, and calculate the cosine similarity between the fault feature vectors and the embedding vectors of the fault type nodes in the knowledge graph: S54: Construct an attention model, output the weights of the features for each entity, and generate soft mapping probabilities; S55: Use the fault nodes mapped from the fault feature vectors through the input layer of the graph neural network as seed nodes, initialize the fault probabilities of the seed nodes and the fault probabilities of other nodes. The nodes pass fault influence messages to their neighbors, and the weights are determined by the confidence of the edges. After probability aggregation, the output layer outputs the probability distribution.
10. A fault diagnosis method for a server according to claim 8, characterized in that The specific process of step S7 is as follows: S71: Obtain multi-source data in real time from the virtual model, including physical dimensions: sensor data such as temperature, pressure, vibration frequency, current, and voltage; logical dimensions: model state transition probabilities, algorithm execution efficiency, parameter convergence, etc.; spatial dimensions: topological structures associated with device locations, signal transmission delays, spatial coupling effects, etc.; S72: Establish a normal behavior baseline: Define the normal range through unsupervised learning, and calculate deviation indicators such as absolute deviation, dynamic deviation rate, and spatio-temporal association deviation; S73: Map the behavior deviation data to initial amplitudes, encode real-valued features as probability amplitudes through embedding, capture the associations of multi-dimensional deviations, simulate the non-linear coupling effects of complex faults, and set up a parameterized circuit to optimize the circuit parameters through variational algorithms to adapt to the dynamic changes of real-time data; S74: Map the associations between probability amplitudes and multi-dimensional deviations to priority scores of processing strategies, and select the optimal strategy in combination with the fault assessment results.
Citation Information
Patent Citations
Vehicle fault diagnosis and maintenance evaluation method based on artificial intelligence
CN119087962A
Intelligent auxiliary diagnosis and maintenance method and system based on multi-path recall
CN119357787A
Complex production line reliability analysis method based on combination of fault operation and maintenance knowledge graph and fault tree
CN119903430A
Fault diagnosis method, electronic device, and storage medium
WO2025081844A1
Cited By
Resource cross-layer collaborative management method and substrate manager
CN120849140A
Resource cross-layer cooperative management method and baseboard manager
CN120849140B
Fault diagnosis agent data preprocessing method and system based on multi-modal alignment
CN120892239A
Fault diagnosis agent data preprocessing method and system based on multi-modal alignment
CN120892239B
Industrial equipment fault intelligent diagnosis system based on deep learning algorithm
CN120995335A