Transfer learning-based cross-server fault prediction system and method

Through a cross-server fault prediction method based on transfer learning, the data of multiple servers is used to train the neural network model in deep domains and fine-tune it, which solves the low-precision problem caused by insufficient data by traditional fault prediction methods, and achieves higher fault prediction accuracy and server stability.

WO2025119083A1PCT designated stage expired Publication Date: 2025-06-12CHINA TELECOM CLOUD TECH CO LTD

Patent Information

Application Number
PCT/CN2024/135493
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-05
Filing Date
2024-11-29
Publication Date
2025-06-12

AI Technical Summary

Technical Problem

The traditional fault prediction method cannot obtain an accurate prediction model due to insufficient data from the target server, resulting in low prediction accuracy, affecting the reliability and performance of the server.

Method used

A cross-server failure prediction method based on transfer learning is adopted, and the data of multiple servers is trained to build a neural network model with deep domain confusion, and fine-tune the model to improve the generalization ability and accuracy of the prediction model.

Benefits of technology

It improves the accuracy of fault prediction, can adapt to the characteristics and changes of different servers, helps administrators to detect potential faults in a timely manner, improves server stability and reliability, and reduces downtime and maintenance costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024135493_12062025_PF_FP_ABST
    Figure CN2024135493_12062025_PF_FP_ABST
Patent Text Reader

Abstract

The present application is suitable for the technical field of fault prediction, and in particular relates to a transfer learning-based cross-server fault prediction system and method. The method comprises: carrying out data collection, collecting server data, and preprocessing the server data to obtain preprocessed data; constructing a neural network model, constructing a data set, and, on the basis of the data set, pre-training the neural network model to obtain a pre-trained model; on the basis of a deep domain confusion mechanism, aligning feature distributions of data domains of different servers to obtain a deep domain confusion neural network model; collecting sample data, and tuning the neural network model to obtain a prediction neural network model; and performing data information testing on a server requiring prediction to obtain a test result, and, on the basis of the test result, performing visualization processing. The present application takes preventive measures to improve the stability and reliability of servers, shortening shutdown time and reducing maintenance costs, thus improving the overall performance of systems and user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Cross-server fault prediction system and method based on transfer learning

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority to the Chinese patent application filed with the China Patent Office on December 5, 2023, with application number 202311655249.1 and invention name “Cross-server fault prediction system and method based on transfer learning”, the entire contents of which are incorporated by reference into this application. Technical Field

[0003] The present application belongs to the field of fault prediction technology, and in particular to a cross-server fault prediction system and method based on transfer learning. Background Art

[0004] In modern computer systems, the normal operation of servers is crucial to maintaining system stability and efficiency. Existing technologies generally calculate the failure probability of a server based on its own status data using a neural prediction model. The details are as follows:

[0005] A neural network prediction model is pre-trained and used to predict whether the server will fail at the current moment. If the server is not failing at the current moment, a target moment and at least two future historical moments for prediction are determined. The state values ​​of each associated parameter corresponding to each historical moment and the current state value are then obtained. Finally, the state values ​​of each associated parameter corresponding to the target moment and the neural network prediction model are used to predict whether the server will fail at the target moment.

[0006] Problems with existing technologies: Due to insufficient data on certain target servers, traditional fault prediction methods often cannot obtain accurate prediction models, resulting in low prediction accuracy, which may affect the reliability and performance of the server. Summary of the Invention

[0007] The purpose of this application is to provide a cross-server fault prediction method based on transfer learning, aiming to solve the problem that traditional fault prediction methods often cannot obtain accurate prediction models, resulting in low prediction accuracy, which may affect the reliability and performance of the server.

[0008] This application provides a cross-server fault prediction method based on transfer learning, the method comprising:

[0009] Perform data collection, collect server data, pre-process the server data, and obtain pre-processed data;

[0010] Build a neural network model, build a data set based on the preprocessed data, and pre-train the neural network model based on the data set to obtain a pre-trained model;

[0011] Based on the deep domain obfuscation mechanism, the feature distributions between different server data domains are aligned to obtain a neural network model for deep domain obfuscation.

[0012] Collect sample data, divide it into training data and verification data, adjust the neural network model based on the training data and verification data, and obtain a predictive neural network model;

[0013] Perform data information testing on the server that needs to be predicted, obtain test results, and perform visualization based on the test results.

[0014] Optionally, the step of collecting data, collecting server data, and preprocessing the server data to obtain preprocessed data specifically includes:

[0015] Collect data from multiple servers to obtain server data, wherein the server data includes at least operation data, log files, and fault information;

[0016] Clean the server data, remove invalid data, and fill in missing data;

[0017] Normalize the server data to the same data range and unit;

[0018] Construct a data set and divide the normalized server data into a training set and a test set;

[0019] Generate corresponding tags based on the fault information and the server's operating status.

[0020] Optionally, the steps of constructing a neural network model, constructing a data set based on preprocessed data, and pre-training the neural network model based on the data set to obtain a pre-trained model specifically include: constructing a neural network model, the neural network model is based on LSTM and self-attention mechanism, freezing some parameters in the neural network model or all parameters of the LSTM layer, using the preprocessed data as training data, training the neural network model, updating parameters of the prediction layer in the neural network model, and saving the neural network model after parameter update to obtain a pre-trained model.

[0021] Optionally, the step of aligning the feature distributions between different server data domains based on the deep domain obfuscation mechanism to obtain a neural network model for deep domain obfuscation specifically includes:

[0022] The pre-trained model contains an adaptive layer that extracts feature representations from the data of the source and target servers;

[0023] The adaptive layer calculates adaptive weights based on the distribution of data;

[0024] The feature representation is adjusted based on the calculated weights.

[0025] Optionally, the step of collecting sample data, dividing the sample data into training data and verification data, and adjusting the neural network model based on the training data and verification data to obtain the prediction neural network model specifically includes:

[0026] Collect sample data from the target server and use it as training data for fine-tuning and validation data for fine-tuning;

[0027] Freeze some parameters in the neural network model of deep domain confusion;

[0028] Adjust the parameters of the adaptive layer and the fully connected layer based on the training data for fine-tuning and the validation data for fine-tuning;

[0029] By performing repeated iterations and parameter adjustments, the neural network model of deep domain confusion is adjusted to obtain a prediction neural network model.

[0030] Optionally, the step of performing data information testing on the server that needs to be predicted to obtain test results specifically includes: constructing data input, where the data input comes from preprocessed data; importing the data input into the prediction neural network model, performing forward propagation calculations to obtain prediction results, and determining components with failure risks based on the prediction results.

[0031] Optionally, the step of performing visualization based on the test results specifically includes constructing a visualization interface, generating a chart based on the prediction results, displaying the chart on the visualization interface, and regularly updating the visualization interface.

[0032] Optionally, whether to trigger an alarm is determined based on the prediction result, and when an alarm is triggered, an alarm notification is sent to an administrator.

[0033] Optionally, the following formula is used to calculate the adaptability weight: m t =β1·m t-1 +(1-β1)·g t

[0034] Among them, m t is the momentum term at time step t, β1 is the momentum parameter, g t is the gradient at the current time step.

[0035] Another object of the present application is to provide a cross-server fault prediction system based on transfer learning, the system comprising:

[0036] A data collection module, wherein the data collection module is used to collect data from each server;

[0037] A data preprocessing module, wherein the data preprocessing module is used to preprocess the server data to obtain preprocessed data;

[0038] A neural network model pre-training module is used to construct a neural network model, construct a data set based on the pre-processed data, and pre-train the neural network model based on the data set to obtain a pre-trained model;

[0039] A domain obfuscation module, which is used to align the feature distributions between different server data domains based on a deep domain obfuscation mechanism to obtain a neural network model for deep domain obfuscation;

[0040] A model fine-tuning module is used to collect sample data, divide it into training data and verification data, and adjust the neural network model based on the training data and verification data to obtain a predictive neural network model;

[0041] A fault prediction module is used to perform data information testing on the server that needs to be predicted and obtain test results;

[0042] The result display module is used to perform visualization based on the test results.

[0043] This application provides a cross-server fault prediction method based on transfer learning. By using data from multiple servers for training, the prediction model has good generalization capabilities and can adapt to the characteristics and changes of different servers. At the same time, the strategy of transfer learning is adopted to apply existing knowledge to the target server, avoiding the problems caused by insufficient data on the target server, thereby effectively improving the accuracy of fault prediction. Such a method can help administrators detect potential faults in a timely manner during actual operation and maintenance, take preventive measures, improve the stability and reliability of the server, reduce downtime and maintenance costs, and improve the overall performance of the system and user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] FIG1 is a flow chart of a cross-server fault prediction method based on transfer learning provided in an embodiment of the present application;

[0045] FIG2 is an architecture diagram of a cross-server fault prediction system based on transfer learning provided in an embodiment of the present application;

[0046] FIG3 is an architecture diagram of a neural network model provided in an embodiment of the present application;

[0047] FIG4 is an architecture diagram of a domain obfuscation module provided in an embodiment of the present application;

[0048] FIG5 is a schematic diagram showing the principle of the model fine-tuning module provided in an embodiment of the present application. DETAILED DESCRIPTION

[0049] In order to make the purpose, technical solutions and advantages of this application more clearly understood, the present application is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0050] It is understood that the terms "first," "second," etc., used herein may be used to describe various elements, but unless otherwise specified, these elements are not limited by these terms. These terms are only used to distinguish a first element from another element. For example, a first xx script may be referred to as a second xx script, and similarly, a second xx script may be referred to as a first xx script without departing from the scope of this application.

[0051] As shown in FIG1 , a flow chart of a cross-server fault prediction method based on transfer learning provided in an embodiment of the present application is provided. The method includes:

[0052] Perform data collection, collect server data, pre-process the server data, and obtain pre-processed data;

[0053] Build a neural network model, build a data set based on the preprocessed data, and pre-train the neural network model based on the data set to obtain a pre-trained model;

[0054] Based on the deep domain obfuscation mechanism, the feature distributions between different server data domains are aligned to obtain a neural network model for deep domain obfuscation.

[0055] Collect sample data, divide it into training data and verification data, adjust the neural network model based on the training data and verification data, and obtain a predictive neural network model;

[0056] Perform data information testing on the server that needs to be predicted, obtain test results, and perform visualization based on the test results.

[0057] In this embodiment, the method specifically includes:

[0058] Step 1: Obtain the operating data, log files, fault information and other data of multiple servers;

[0059] Step 2: Analyze the server data and remove information data that is not related to the server failure;

[0060] Step 3: Analyze the data information and divide it into information data of each component of the server;

[0061] Step 4: Perform training based on the data information to obtain a basic training model with initial weight parameters;

[0062] Step 5: Add a deep domain obfuscation mechanism to the basic training model to align the feature distributions between different server data domains and obtain a deep domain obfuscation neural network model;

[0063] Step 6: Use a small amount of data information of the known target server to fine-tune the neural training model to obtain a predictive neural network model;

[0064] Step 7: Based on the prediction neural network model, a test can be performed according to the data information of the target server, and the output is the failure probability of each component of the target server at a certain moment in the future.

[0065] As shown in FIG2 , an embodiment of the present application further provides a cross-server fault prediction system based on transfer learning, the system comprising:

[0066] A data collection module, wherein the data collection module is used to collect data from each server;

[0067] A data preprocessing module, wherein the data preprocessing module is used to preprocess the server data to obtain preprocessed data;

[0068] A neural network model pre-training module is used to construct a neural network model, construct a data set based on the pre-processed data, and pre-train the neural network model based on the data set to obtain a pre-trained model;

[0069] A domain obfuscation module, which is used to align the feature distributions between different server data domains based on a deep domain obfuscation mechanism to obtain a neural network model for deep domain obfuscation;

[0070] A model fine-tuning module is used to collect sample data, divide it into training data and verification data, and adjust the neural network model based on the training data and verification data to obtain a predictive neural network model;

[0071] A fault prediction module is used to perform data information testing on the server that needs to be predicted and obtain test results;

[0072] The result display module is used to perform visualization based on the test results.

[0073] In this embodiment, 1) Data Collection Module: This module is responsible for collecting data such as operation data, log files, and fault information from multiple servers. The data may include server performance indicators, system logs, fault records, etc.

[0074] 2) Data preprocessing module: After data collection, this module preprocesses the collected data, removes information data not related to server failures, and cleans and normalizes the data for subsequent model training.

[0075] 3) Neural Network Model Pretraining Module: After data preprocessing, this module uses data collected from multiple source servers to train the neural network model. This model is a deep learning model based on LSTM and self-attention mechanisms. Through pretraining using transfer learning, a fault prediction model with initial weights is obtained.

[0076] 4) Domain Obfuscation Module: Introducing a deep domain obfuscation mechanism into the neural network model enables the model to learn to remove domain-specific information from features, which helps to reduce the data differences between the source server data domain and the target server data domain.

[0077] 5) Model fine-tuning module: Apply the domain-confused neural network model to the target server and fine-tune the model parameters to make them suitable for the target server to solve the problem of insufficient sample data on the target server.

[0078] 6) Fault Prediction Module: After the model fine-tuning module is complete, the model will be tested on the target server's data using the predictive neural network model. Based on the target server's data, the model will output the probability of failure of each component at a certain point in the future.

[0079] 7) Results Display Module: This module is responsible for displaying the fault prediction results to administrators or maintenance personnel. The results can be displayed through a visual interface to intuitively demonstrate the failure probability of the target server, helping administrators to promptly identify potential faults and take appropriate measures.

[0080] In this embodiment, the entire process can be described as:

[0081] (1) Data collection and preprocessing:

[0082] 1) Data Collection: This collects operational data, log files, and fault information from multiple servers. Operational data may include various server performance metrics, such as CPU usage, memory usage, and network traffic; log files include server system logs and error logs; and fault information records different types of server failures.

[0083] 2) Data cleaning: Clean the collected data to remove invalid or missing data. You may encounter some missing data or outliers, which require appropriate methods to handle, such as using interpolation to fill missing values ​​or correcting outliers based on historical data and rules.

[0084] 3) Data Normalization: Because different servers may have different data ranges and units, data normalization is necessary to ensure better convergence and optimization during model training. Common normalization methods include scaling the data to a range of 0 to 1 or using standardization to achieve a mean of 0 and a variance of 1.

[0085] 4) Dataset Partitioning: Before training, the collected dataset needs to be divided into training and test sets. The training set is used for model training and parameter optimization, while the test set is used to evaluate model performance. Typically, random or chronological partitioning can be used to ensure that the data distribution of the training and test sets is roughly the same.

[0086] 5) Label Generation: Generate labels for data samples based on the fault information and server operating status. Labels are multi-class labels that represent different types of faults at a given moment. Label generation should be tailored to the specific situation and problem definition.

[0087] (2) In the neural network model pre-training module, the pre-training method in transfer learning is used to build a fault prediction model. Transfer learning is a learning method that uses knowledge learned from one task to assist another related task. The specific pre-training process is as follows:

[0088] 1) Pre-trained Model Design: A deep learning model based on LSTM and self-attention mechanisms was constructed. As shown in Figures 1, 2, and 3, the model's input consists of source server status data and log files from different time periods. The embedding layer maps the source server data into other vector representations, converting high-dimensional, sparse feature vectors into low-dimensional, dense feature vectors. The closer the relationship between source server data, the closer the distance between the feature vectors after embedding. In the LSTM layer of the model, the model uses LSTM to analyze server data files to detect possible failure types and timing. LSTM (Long Short-Term Memory) is a variant of the recurrent neural network (RNN) that addresses the problems of vanishing and exploding gradients in traditional RNN models. The core concept of LSTM is the introduction of memory cells and a series of gating mechanisms to effectively capture and transmit long-term dependencies in sequential data, making it effectively applicable to server failure prediction. LSTM consists of a memory cell, an input gate, a forget gate, and an output gate. The output of the LSTM model is the important features of the input data, which are then passed to the self-attention mechanism to further model the relevance and contextual information of the data. The self-attention mechanism focuses on the relationship between each input in the server data. First, the weight matrix W is initialized. Q 、W K and W V , the dimensions of these three matrices are the same. Multiply the feature matrix E by the three weight matrices respectively to obtain the query matrix Q, key matrix K and value matrix V. The formula is as follows: (Q,K,V)=E·(W Q ,W K ,W V );

[0089] The self-attention score is then calculated and adjusted to make the neural network more stable during training. The adjusted self-attention score is then subjected to a softmax function to convert all scores into positive numbers, with the sum of the scores being 1. The next step is to multiply the obtained score by the corresponding V to adjust the weights and reduce the weight of data not related to the server failure. Finally, all V are added together, using the following formula:

[0090] The obtained attention scores are then fed into a fully connected layer, followed by a softmax calculation to determine the type and probability of a server failure occurring at a certain point in the future. This model performs well when sufficient information about source server failures is available. This pre-trained model is typically pre-trained on large-scale datasets and has achieved good performance on related tasks.

[0091] 2) Freeze some parameters: To preserve the learning ability of the base model for the original task, freeze some of the model's parameters, typically the first few layers or all LSTM layers. This ensures that the base model's feature extraction capabilities are preserved, while only the prediction layer is trained.

[0092] 3) Data Preparation: Use information data collected from multiple servers at different times and input it into the pre-trained model as training data. Since some parameters are frozen, only the parameters of the prediction layer will be updated.

[0093] 4) Pre-training: While freezing some parameters, the parameters of the prediction layer are trained using the back-propagation algorithm and optimization methods. The goal of pre-training is to obtain a fault prediction model with initial weights based on the data from the source server.

[0094] 5) Save the pre-trained model: After the pre-training is completed, save the obtained pre-trained model as the initial model for transfer learning.

[0095] Through the above pre-training method, a fault prediction model with initial weights was obtained. This pre-trained model was applied to the target server, and the model was further fine-tuned through transfer learning to solve the problem of low prediction accuracy caused by insufficient data on the target server.

[0096] (3) As shown in Figures 1, 2, 3, and 4, in the domain confusion module, an adaptive layer is added before the fully connected layer. The adaptive layer calculates the distance between the two distributions and adds this distance as a loss function to the model training process. It also adjusts the input features appropriately to reduce the specific differences between the source server data domain and the target server data domain, thereby reducing the generalization ability of the model. The specific process is as follows:

[0097] 1) Feature representation: Extract feature representations from the data of the source and target servers, such as the output of the LSTM layer neural network.

[0098] 2) Adaptive Weight Calculation: The adaptive layer calculates adaptive weights based on the data distribution. These weights are adjusted based on the difference in data distribution between the source and target domains. A common method, Maximum Mean Discrepancy (MMD), is used to calculate these weights. MMD involves mapping the sample features of the two distributions into a high-dimensional space and calculating the distance in that space. The formula for MMD is as follows:

[0099] Where p and q represent the distribution of the two data domains, n and m represent the number of samples between the two distributions, Represents the feature mapping function. The smaller the MMD, the more similar the two distributions are. The purpose of domain confusion is to minimize the MMD between different data domains. Therefore, the total loss function of deep domain confusion is: total =L cls +λL MMD

[0100] Among them, L cls Represents the classification loss; λ is a hyperparameter used to balance the classification loss and the loss of the adaptive layer; L MMD represents the loss of the adaptive layer, calculated using MMD. During the final backpropagation, the entire network is updated using both MMD and loss. Therefore, by optimizing the total loss function, we can align feature distributions across different domains, thereby improving the generalization performance of the neural network model.

[0101] 3) Feature Adjustment: The calculated adaptive weights are applied to the feature representation to adjust the feature expression. This means that some parts of the feature are emphasized while others are suppressed, thereby reducing the differences between different data domains.

[0102] (4) In the model fine-tuning module, the fine-tuning method in transfer learning is used to apply the trained domain confusion training model to the target server to solve the problem of insufficient sample data on the target server. When adopting the fine-tuning strategy, the previous fine-tuning strategy is improved. Only the adaptive layer and fully connected layer of the model are fine-tuned, which can speed up the model training and make the model converge better. The specific fine-tuning process is as follows:

[0103] 1) Data Preparation: First, collect a small amount of sample state data from the target server at different times and a small amount of sample data from the source server at different times to serve as training and validation data for fine-tuning. Since the target server has less sample data, this data may not be sufficient for training the original model, so fine-tuning is required to adapt it to the target server's data.

[0104] 2) Freeze Some Parameters: Similar to pre-training, this method freezes the parameters of all LSTM layers and the attention mechanism layer of the deep domain confusion neural network model, and only fine-tunes the model's adaptive and fully connected layers. This preserves the feature extraction capabilities learned by the domain confusion network model on large-scale data, speeds up model training, and ensures better convergence.

[0105] 3) Data input and fine-tuning: The training data of the target server is input into the domain confusion pre-trained model, and the parameters of the adaptive layer and the fully connected layer are fine-tuned. During the fine-tuning process, only the weight parameters of the adaptive layer and the fully connected layer are updated through the backpropagation algorithm and optimization method, so that the model can adapt to the data distribution of the target server. During the fine-tuning process, a small learning rate is used to avoid losing useful knowledge of the domain confusion model. Among them, the Adam (Adaptive Moment Estimation) optimization algorithm is used to adjust the weight parameters of the model; combining the ideas of momentum and adaptive learning rate, it accelerates the convergence process and avoids falling into local minima. The following are the main features of the Adam algorithm:

[0106] Momentum term: Adam introduces the concept of momentum, which is similar to the traditional momentum optimization algorithm. The momentum term is used to maintain the direction of the previous gradient to reduce the oscillation of gradient descent. The momentum term is calculated as follows: m t =β1·m t-1 +(1-β1)·g t

[0107] Among them, m t is the momentum term at time step t, β1 is the momentum parameter, g t is the gradient at the current time step.

[0108] Adaptive learning rate: Adam also introduces an adaptive learning rate, which is different from the traditional fixed learning rate. The core idea of ​​the adaptive learning rate is to automatically adjust the learning rate based on the gradient of each parameter. The adaptive learning rate is calculated as follows:

[0109] where v t is the moving average of the squared gradient term at time step t, and β2 is the decay rate parameter of the squared gradient.

[0110] Bias correction: Since the estimated values ​​of the momentum term and the adaptive learning rate term are biased towards zero at the beginning, Adam performs bias correction to correct this problem. The corrected momentum and learning rate are calculated as follows:

[0111] Parameter Update: Finally, the parameters are updated using the modified momentum and adaptive learning rate, as well as a small constant ∈ to avoid division by zero errors:

[0112] Among them, θ t is the model parameter at time step t, α is the learning rate, and ε is the smoothing term

[0113] 4) Iterative fine-tuning: Repeat the iterative fine-tuning process multiple times until the model achieves good performance on the target server's data. Since the target server has limited sample data, the number of iterative fine-tuning times should not be excessive, and appropriate adjustments can be made based on the target server's sample data.

[0114] 5) Save the fine-tuned model: After fine-tuning is completed, save the fine-tuned model as the final fault prediction model; obtain a fault prediction model that performs well on the target server, which can be used to accurately predict the failure probability of each component of the target server at a certain time in the future.

[0115] (5) In the fault prediction module, the fine-tuned prediction neural network model is used to test the target server's data information. This module aims to predict the target server's faults using the prediction model and obtain the failure probability of each component. The specific process is as follows:

[0116] 1) Data Input: We extract operational data, log files, and other relevant information from the target server's test group data at specific times as input. This data is prepared during the data preprocessing and model fine-tuning phases, ensuring data availability and accuracy.

[0117] 2) Forward Propagation: The input test data is fed into the fine-tuned predictive neural network model for forward propagation calculations. During the forward propagation process, the model extracts features and calculates the test data to extract information relevant to fault prediction.

[0118] 3) Output: After forward propagation, the model generates predictions. These predictions include the probability of failure of each component at a certain time in the future. Based on the target server's test data and fine-tuned parameters, the model predicts failures for the target server.

[0119] 4) Fault prediction result analysis: After obtaining the prediction results, further analysis and interpretation can be performed. For example, based on the probability of failure, components at risk can be identified and preventive maintenance and optimization can be performed based on the predicted probability of failure.

[0120] 5) Feedback and Optimization: Prediction results can provide important insights for actual server management and maintenance. If the prediction results don't match the actual failure scenario, the model can be optimized and improved through the feedback mechanism to increase prediction accuracy and reliability.

[0121] The fault prediction module can fully utilize the fine-tuned prediction model to predict faults on the test group data of the target server, output valuable prediction results on the probability of server failure in the future, provide support for server management and maintenance, and optimize system stability and performance.

[0122] (6) In the result display module, the prediction results obtained by the fault prediction module will be presented to the administrator or operation and maintenance personnel. The main task of this module is to display information such as the predicted fault probability and the type of fault through a visual interface so that the administrator can intuitively understand the status of the target server, discover potential faults in a timely manner, and take appropriate measures to deal with them. The specific implementation is as follows:

[0123] Visualization interface design: In the results display module, design a user-friendly visualization interface that displays the target server's failure prediction results through charts, line graphs, pie charts, and other formats. The interface should be concise and clear, providing multiple interactive methods to enable administrators to easily view and analyze prediction results.

[0124] Failure Probability Display: Displays the failure probability of each component in a chart or pie chart. Administrators can clearly understand the likelihood of each component failing at a certain point in the future, allowing them to conduct targeted troubleshooting and resolution.

[0125] Real-time monitoring: This module also provides real-time monitoring functions and regularly updates fault prediction results to ensure that administrators obtain the latest server status information in a timely manner.

[0126] Alarm mechanism: If there is a high-risk failure situation in the prediction results, the module can automatically trigger the alarm mechanism and send an alarm notification to the administrator so that the administrator can take immediate measures to handle the failure.

[0127] Through the result display module, administrators can easily view the fault prediction results of the target server, understand the server status in real time, discover potential faults in time and take corresponding measures, thereby improving the stability and performance of the server and ensuring the reliable operation of the system.

[0128] 1. This application uses a transfer learning strategy for cross-service fault prediction. Currently, common server fault prediction methods rely on the server's own status data. However, when the target server has less data, the prediction accuracy is often low. Therefore, the transfer learning strategy solves the problem of overfitting or underfitting caused by too little sample data on the target server.

[0129] 2. This application uses a transfer learning mechanism based on a combination of deep domain obfuscation and fine-tuning. The MMD algorithm is used in the deep domain obfuscation mechanism, and the fine-tuning strategy uses a strategy that only fine-tunes the weight parameters of the last two layers. This effectively reduces the gap between the source server data domain and the target server data domain, accelerates the model's convergence speed, reduces the model's fine-tuning training time, and to a certain extent improves the generalization ability of the fault prediction model.

[0130] 3. The deep learning model based on LSTM and self-attention mechanism provided in this application can predict the probability of failure of different components of the server at a certain moment in the future based on the status data of the server at different moments.

[0131] It should be understood that, although each step in the flow chart of each embodiment of the present application is shown in sequence according to the indication of the arrow, these steps are not necessarily performed in sequence according to the order indicated by the arrow. Unless there is clear explanation in this article, the execution of these steps does not have strict order restriction, and these steps can be performed in other orders. Moreover, at least a portion of the steps in each embodiment may include a plurality of sub-steps or a plurality of stages, and these sub-steps or stages are not necessarily performed at the same time, but can be performed at different times, and the execution order of these sub-steps or stages is not necessarily performed in sequence, but can be performed in turn or alternately with at least a portion of other steps or sub-steps or stages of other steps.

[0132] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0133] The technical features of the above-mentioned embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above-mentioned embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0134] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present invention, all of which fall within the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be determined by the appended claims.

[0135] The above description is only a preferred embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent replacements and improvements made within the spirit and principles of the present application should be included in the scope of protection of the present application.

Claims

1. A cross-server fault prediction method based on transfer learning, characterized in that: The method comprises: Perform data collection, collect server data, pre-process the server data, and obtain pre-processed data; Construct a neural network model, construct a data set based on the preprocessed data, pre-train the neural network model based on the data set, and obtain a pre-trained model; Based on the deep domain obfuscation mechanism, the feature distributions between different server data domains are aligned to obtain a neural network model for deep domain obfuscation. Collect sample data, divide it into training data and verification data, adjust the neural network model based on the training data and verification data, and obtain a prediction neural network model; Perform data information testing on the server that needs to be predicted, obtain test results, and perform visualization based on the test results.

2. The cross-server fault prediction method based on transfer learning according to claim 1 is characterized in that: The step of collecting data, collecting server data, preprocessing the server data, and obtaining preprocessed data specifically includes: Collect data from multiple servers to obtain server data, wherein the server data includes at least operation data, log files, and fault information; Clean the server data, remove invalid data, and fill in missing data; Normalize the server data to the same data range and unit; Construct a data set and divide the normalized server data into a training set and a test set; Generate corresponding tags based on the fault information and the running status of the server.

3. The cross-server fault prediction method based on transfer learning according to claim 1 is characterized in that: The steps of constructing a neural network model, constructing a data set based on preprocessed data, pretraining the neural network model based on the data set, and obtaining a pretrained model specifically include: constructing a neural network model, the neural network model is based on LSTM and self-attention mechanism, freezing some parameters in the neural network model or all parameters of the LSTM layer, using the preprocessed data as training data, training the neural network model, updating parameters of the prediction layer in the neural network model, saving the neural network model after parameter update, and obtaining a pretrained model.

4. The cross-server fault prediction method based on transfer learning according to claim 1 is characterized in that: The step of aligning the feature distributions between different server data domains based on the deep domain obfuscation mechanism to obtain a neural network model of deep domain obfuscation specifically includes: The pre-trained model contains an adaptive layer that extracts feature representations from the data of the source and target servers; The adaptive layer calculates adaptive weights based on the distribution of data; The feature representation is adjusted based on the calculated weights.

5. The cross-server fault prediction method based on transfer learning according to claim 1 is characterized in that: The step of collecting sample data, dividing it into training data and verification data, adjusting the neural network model based on the training data and the verification data, and obtaining a prediction neural network model specifically includes: Collect sample data from the target server and use it as training data for fine-tuning and verification data for fine-tuning; Freeze some parameters in the neural network model of deep domain confusion; Adjust parameters of the adaptive layer and the fully connected layer based on the training data for fine-tuning and the validation data for fine-tuning; By performing repeated iterations and parameter adjustments, the neural network model for deep domain confusion is adjusted to obtain a predictive neural network model.

6. The cross-server fault prediction method based on transfer learning according to claim 1 is characterized in that: The step of performing data information testing on the server that needs to be predicted to obtain the test results specifically includes: constructing data input, where the data input comes from preprocessed data; importing the data input into the prediction neural network model, performing forward propagation calculations to obtain prediction results, and determining components with failure risks based on the prediction results.

7. The cross-server fault prediction method based on transfer learning according to claim 1 is characterized in that: The steps of performing visualization based on the test results specifically include building a visualization interface, generating charts based on the prediction results, displaying the charts on the visualization interface, and regularly updating the visualization interface.

8. The cross-server fault prediction method based on transfer learning according to claim 1, characterized in that: Based on the prediction results, it is determined whether an alarm is triggered. When an alarm is triggered, an alarm notification is sent to the administrator.

9. The cross-server fault prediction method based on transfer learning according to claim 4 is characterized in that: The following formula is used to calculate the adaptability weight: m t =β1·m t-1 +(1-β1)·g t Among them, m t is the momentum term at time step t, β1 is the momentum parameter, and g t is the gradient at the current time step.

10. A cross-server fault prediction system based on transfer learning, characterized in that: The system comprises: A data collection module, wherein the data collection module is used to collect data from each server; A data preprocessing module, wherein the data preprocessing module is used to preprocess the server data to obtain preprocessed data; A neural network model pre-training module, wherein the neural network model pre-training module is used to construct a neural network model, construct a data set based on pre-processed data, and pre-train the neural network model based on the data set to obtain a pre-trained model; A domain obfuscation module, wherein the domain obfuscation module is used to align feature distributions between different server data domains based on a deep domain obfuscation mechanism to obtain a neural network model of deep domain obfuscation; A model fine-tuning module, wherein the model fine-tuning module is used to collect sample data, divide the sample data into training data and verification data, and adjust the neural network model based on the training data and the verification data to obtain a prediction neural network model; A fault prediction module, wherein the fault prediction module is used to perform data information testing on the server that needs to be predicted to obtain a test result; A result display module is used to perform visual processing based on the test results.

Citation Information

Patent Citations

  • Dynamic field self-adaption method and device and computer readable storage medium

    CN110135510A

  • Unsupervised anomaly prediction method for two-stage cloud server

    CN111914873A

  • Medical image recognition method based on dynamic field adaptive learning

    CN113658110A

  • Cross-server fault prediction system and method based on transfer learning

    CN117950965A

  • Method and apparatus for detecting pores based on artificial neural network and visualizing the detected pores

    US11636600B1

Cited By

  • Multi-physical field regulation and control method and system in aluminum product processing

    CN120540253A

  • Grey fault processing method and device, medium and program product

    CN120631741A

  • Lightning arrester performance degradation index construction method based on improved neural network

    CN120744843A

  • Fan equipment fault prediction method based on meta learning and multi-expert model

    CN120868057A

  • Elevator fault intelligent maintenance processing method based on AI large model elevator industry data

    CN120875832A