Server Fault Prediction Method and System Based on Edge Computing of BMC Module

By deploying BMC modules at the server edge nodes, building an integrated learning fault prediction model and monitoring CPU occupancy, the problem of unreasonable resource utilization in the existing technology is solved, efficient fault prediction and resource scheduling is achieved, and the stability and accuracy of the server are improved.

CN118733317BActive Publication Date: 2025-07-11WUHAN UNIV OF TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202410826467.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-25
Publication Date
2025-07-11
Estimated Expiration
2044-06-25

AI Technical Summary

Technical Problem

Existing server failure prediction systems cannot make rational use of computing resources, lack continuous monitoring and analysis of server real-time running data, and it is difficult to adapt to complex and changeable server environments.

Method used

The BMC module is deployed at the edge node of the target server, and the historical operating status data is obtained through the BMC module, a fault prediction model based on integrated learning is built, and the CPU occupation is monitored for dynamic scheduling of computing resources is used. The rotation method of linear classifier-decision tree model-perceptual model-naive Bayesian model is used for training, and the resource allocation priority is adjusted according to the CPU occupancy rate.

Benefits of technology

Reduces server pressure, improves the accuracy of resource utilization and failure prediction, adapts to complex and changeable server environments, and ensures that critical services have sufficient computing resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118733317B_ABST
    Figure CN118733317B_ABST
Patent Text Reader

Abstract

The present invention discloses a server fault prediction method and system based on BMC module edge computing. The method includes: deploying a BMC module at the edge node of the target server, and obtaining the historical operation status data of the target server through the BMC module; performing format parsing and preprocessing on the historical operation status data of the target server to construct a sample set for fault prediction; constructing a fault prediction model based on ensemble learning, training the fault prediction model through the sample set, and deploying the trained fault prediction model to the BMC module; performing server fault prediction through the fault prediction model, and simultaneously monitoring the CPU occupancy of the BMC module, and dynamically scheduling the computing resources of the fault prediction model according to the CPU occupancy rate. The present invention reduces the pressure on the target server through an out-of-band prediction method, and at the same time can dynamically schedule the computing resources of the fault prediction model according to the CPU occupancy rate, improving the resource utilization rate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of server security, and specifically relates to a server fault prediction method and system based on edge computing of BMC modules. Background Art

[0002] With the rapid development of Internet services, servers, as the key infrastructure for carrying various services and data, have received unprecedented attention in terms of their stability and security. Server failures may lead to service interruptions, data loss, and business losses. Therefore, predicting and timely handling server failures are very important.

[0003] Generally, the server fault prediction and processing system is in-band prediction. The central server is under great pressure and cannot reasonably utilize computing resources. At the same time, it lacks continuous monitoring and analysis of the real-time operation data of the server, cannot adjust the prediction model in a timely manner, and is difficult to adapt to the complex and changeable server environment.

[0004] Therefore, a new server fault prediction method is needed to relieve the pressure on the central server and adapt to the complex and changeable server environment. Summary of the Invention

[0005] In view of this, the present invention proposes a server fault prediction method and system based on edge computing of BMC modules to solve the problem that the server fault prediction method cannot reasonably utilize computing resources.

[0006] In the first aspect of the present invention, a server fault prediction method based on edge computing of BMC modules is disclosed. The method includes:

[0007] Deploy a BMC module at the edge node of the target server, and obtain the historical operation status data of the target server through the BMC module;

[0008] Parse and preprocess the format of the historical operation status data of the target server to construct a sample set for fault prediction;

[0009] Construct a fault prediction model based on ensemble learning, train the fault prediction model through the sample set, and deploy the trained fault prediction model to the BMC module;

[0010] Perform server fault prediction through the fault prediction model, and at the same time monitor the CPU occupancy of the BMC module, and perform dynamic scheduling of the computing resources of the fault prediction model according to the CPU occupancy rate.

[0011] Based on the above technical solutions, preferably, the step of deploying a BMC module at the edge node of the target server and obtaining the historical operation status data of the target server through the BMC module specifically includes:

[0012] Obtain the resource URL of the Redfish service of the target server, send a GET request to the resource URL of the target server, and obtain the historical operation status data of the target server; the historical operation status data of the target server includes power status, thermal status, processor utilization rate, memory usage, and storage device status.

[0013] Based on the above technical solutions, preferably, the format parsing and preprocessing of the historical operation status data of the target server to construct a sample set for fault prediction specifically includes:

[0014] Perform format parsing and preprocessing on the historical operation status data of the target server, and the preprocessing includes missing value processing, outlier detection and elimination;

[0015] Obtain the fault types corresponding to the historical operation status data of the target server;

[0016] Using the preprocessed historical operation status data as sample attributes and the fault types as sample labels, construct training samples and perform data augmentation to obtain a sample set.

[0017] Based on the above technical solutions, preferably, the weak learners of the ensemble learning include linear classifiers, decision tree models, perceptron models, and naive Bayes models. Assign a weight to each weak learner, and perform weighted combination of different weak learners to form a strong classifier as the fault prediction model.

[0018] Based on the above technical solutions, preferably, the training of the fault prediction model through the sample set specifically includes:

[0019] Assign the same initial weight to each sample in the sample set to obtain a training set;

[0020] Train a weak learner through the training set to obtain a weak classifier;

[0021] Input each sample attribute in the training set into the weak classifier for prediction, and take the maximum value of the prediction probability as the prediction confidence C i , and take the prediction error of each sample as the classification difficulty D i ;

[0022] Establish a weight adjustment formula based on the prediction confidence and classification difficulty, and update the weight of each sample in the training set;

[0023] Train the next weak learner through the updated training set. In each iteration, use the prediction results of the previous weak learner to adjust the weight of each sample in the training set, and train a new weak learner until the preset number of iterations is reached.

[0024] Based on the above technical solutions, preferably, the weight adjustment formula is:

[0025]

[0026] ω i is the sample weight before adjustment, is the sample weight after adjustment, α is the weight of the weak learner, and C i is the prediction confidence, and D i is the classification difficulty, and y i 、 are the true value and the predicted value of the sample respectively.

[0027] Based on the above technical solutions, preferably, the dynamic scheduling of the computing resources of the fault prediction model according to the CPU occupancy rate specifically includes:

[0028] Calculating the normalized CPU occupancy rate deviation N according to the CPU occupancy rate:

[0029]

[0030] where L and H are the lower and upper limits of the preset CPU occupancy rate thresholds respectively, D = |U * 100 - 50|, and U is the currently obtained CPU occupancy rate;

[0031] Calculating the resource allocation coefficient RAC according to the normalized CPU occupancy rate deviation N:

[0032] RAC = 1 - N 2

[0033] Dynamically adjusting the priority of the computing resource allocation of the fault prediction model according to the value of the resource allocation coefficient RAC.

[0034] In the second aspect of the present invention, a server fault prediction system based on BMC module edge computing is disclosed, and the system includes:

[0035] Data acquisition module: used to deploy the BMC module at the edge node of the target server, and obtain the historical operation status data of the target server through the BMC module;

[0036] Model training module: used to parse and preprocess the format of the historical operation status data of the target server, construct a sample set for fault prediction; construct a fault prediction model based on ensemble learning, train the fault prediction model through the sample set, and deploy the trained fault prediction model to the BMC module;

[0037] Fault prediction module: used to perform server fault prediction through a fault prediction model, monitor the CPU occupancy of the BMC module at the same time, and perform dynamic scheduling of the computing resources of the fault prediction model according to the CPU occupancy rate.

[0038] In the third aspect of the present invention, an electronic device is disclosed, including: at least one processor, at least one memory, a communication interface, and a bus;

[0039] Wherein, the processor, the memory, and the communication interface complete mutual communication through the bus;

[0040] The memory stores program instructions executable by the processor, and the processor invokes the program instructions to implement the method described in the first aspect of the present invention.

[0041] In the fourth aspect of the present invention, a computer-readable storage medium is disclosed. The computer-readable storage medium stores computer instructions, and the computer instructions enable a computer to implement the method described in the first aspect of the present invention.

[0042] The present invention has the following beneficial effects compared with the prior art:

[0043] 1) In the present invention, a BMC module is deployed at the edge node of the target server. The historical operation status data and real-time operation status data of the target server are obtained through the BMC module. Based on the historical operation status data, an integrated learning method is used to train a fault prediction model. After the fault prediction model is deployed to the BMC module at the edge node, the real-time operation status data is used for fault prediction. This out-of-band prediction method reduces the pressure on the target server through edge computing and saves resources. In addition, while performing fault prediction, the present invention monitors the CPU occupancy of the BMC module and performs dynamic scheduling of the computing resources of the fault prediction model according to the CPU occupancy rate, improving the resource utilization rate and ensuring that key services have sufficient computing resources.

[0044] 2) The present invention uses a method of rotating weak learners such as a linear classifier - decision tree model - perceptron model - naive Bayes model for integrated learning training, assigns weights to each sample, and takes the maximum value of the prediction probabilities of the weak classifiers as the prediction confidence, and takes the prediction error of each sample as the classification difficulty; a weight adjustment formula is established based on the prediction confidence and classification difficulty to update the weights of each sample in the training set; the next weak learner is trained through the updated training set. In each iteration, the prediction results of the previous weak learner are used to adjust the weights of each sample in the training set until the training is completed. Compared with the conventional fault prediction method, the fault prediction model obtained by the present invention based on integrated learning training has higher prediction accuracy, and can improve the accuracy and generalization ability of the fault prediction model through continuous learning and optimization to adapt to the complex and changeable server environment.

[0045] 3) The present invention uses the polling API method to obtain the CPU occupancy of the BMC module in real time, calculates the normalized CPU occupancy deviation N based on the CPU occupancy rate, maps it into a computing resource allocation coefficient, and dynamically adjusts the computing resource allocation priority of the fault prediction model according to the computing resource allocation coefficient. When the BMC module faces performance bottlenecks or resource shortages, the priority of the fault prediction model is reduced, its computing resource allocation is decreased, the system burden is alleviated, and it is ensured that critical services obtain sufficient resource support. When the resources of the BMC module are sufficient, the priority of the fault prediction model is increased to ensure that it can obtain sufficient computing resources for efficient fault prediction. The dynamic scheduling method of computing resources improves resource utilization, guarantees the stability of the BMC module, and further alleviates the pressure on the target server. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0047] Figure 1 is a flow diagram of the server fault prediction method based on BMC module edge computing of the present invention;

[0048] Figure 2 is a training flow chart of the fault prediction model of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0049] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.

[0050] Please refer to Figure 1 , the present invention discloses a server fault prediction method based on BMC module edge computing, and the method includes:

[0051] S1. Deploy a BMC module at the edge node of the target server, and obtain the historical operation status data of the target server through the BMC module.

[0052] Step S1 specifically includes the following sub-steps:

[0053] S11. Obtain the Redfish service root URL of the target server and access the root URL of the Redfish service through a network communication tool; under the root URL of the Redfish service, find the link pointing to the Systems collection, which contains all system information about the target server.

[0054] S12. Send a GET request to the resource URL of the target server to obtain the historical operation status data of the target server.

[0055] By sending a GET request to the resource URL of the target server, detailed information about the operation status of the target server can be obtained, including power status, thermal status, processor utilization rate, memory usage, storage device status, etc., and can also include network status, storage device status, system load status, log information, and application status, etc.

[0056] The present invention constructs a sample set by obtaining the historical operation status data of the target server for training a fault prediction model.

[0057] S2. Perform format parsing and preprocessing on the historical operation status data of the target server to construct a sample set for fault prediction.

[0058] Step S2 specifically includes the following sub-steps:

[0059] S21. Perform format parsing and preprocessing on the historical operation status data of the target server

[0060] The data obtained in step S1 is usually in json format. Parse the json format data to obtain the data required in the subsequent steps. Then perform missing value processing, outlier detection and elimination on the parsed data. For the method of outlier detection and elimination, clustering algorithms can be used to detect outliers and then eliminate them.

[0061] S22. Obtain the fault types corresponding to the historical operation status data of the target server.

[0062] Obtain the fault types corresponding to the historical operation status data of the target server, such as power failures, including problems such as power outages and unstable power supplies that cause the server to fail to boot or work properly; heat dissipation failures, such as performance degradation or system crashes caused by overheating of the server; processor failures, such as problems such as overheating of the processor and processor damage that cause the server to work slowly or the system to be unstable; memory failures, such as memory damage or insufficient memory that cause the server to run slowly or the system to crash; network failures, such as problems with network connections that cause the server to be unable to communicate normally or access external networks; storage device failures, such as hard disk damage, RAID failures, etc. that cause data reading or storage failures, and even data loss, etc.; log anomalies, such as abnormal log information that may indicate potential problems or failures; application crashes, such as errors or crashes of applications running on the server.

[0063] S23. Use the preprocessed historical operation status data as sample attributes and the fault types as sample labels to construct training samples and perform data augmentation to obtain a sample set.

[0064] Specifically, new features can be constructed by means of feature extraction, data interpolation, data conversion, adding random noise, etc. to the preprocessed historical operation status data, perform data augmentation on the training samples to obtain a sample set, improve the diversity and quantity of the samples, and thus improve the training effect and generalization ability of the model.

[0065] S3. Construct a fault prediction model based on ensemble learning, train the fault prediction model through the sample set, and deploy the trained fault prediction model to the BMC module.

[0066] The present invention constructs a fault prediction model by using the ensemble learning method. The weak learners of the ensemble learning of the present invention include a linear classifier, a decision tree model, a perceptron model, and a naive Bayes model. A weight is assigned to each weak learner, and different weak learners are weighted and combined to form a strong classifier as the fault prediction model. In the ensemble learning of the present invention, the weak learners are set to rotate in the order of linear classifier - decision tree - perceptron - naive Bayes.

[0067] Figure 2 The following shows a schematic diagram of the training process of the fault prediction model of the present invention. The specific process of training the fault prediction model in step S3 includes:

[0068] S31. Assign the same initial weight to each sample in the sample set to obtain a training set.

[0069] The present invention selects the Boosting algorithm to train the model to improve the prediction accuracy and stability of the model, and assigns the same initial weight ω i to each sample in the sample set to obtain a training set.

[0070] S32. Train a weak learner using the training set to obtain a weak classifier.

[0071] Select a linear classifier as the first weak learner, and train this weak learner using the weak classifier obtained in step S31. After training, a weak classifier is obtained.

[0072] S33. Input the attributes of each sample in the training set into the weak classifier for prediction. Take the maximum value of the prediction probability as the prediction confidence, and take the prediction error of each sample as the classification difficulty.

[0073] Re - input the samples in the training set into the weak learner trained in step S32 for prediction, record the maximum value of the prediction probability, denoted as the prediction confidence C i ; record the prediction error of each sample, denoted as the classification difficulty D i .

[0074] S34. Establish a weight adjustment formula based on the prediction confidence and classification difficulty, and update the weight of each sample in the training set.

[0075] The basic idea of establishing the weight adjustment formula in the present invention is: for samples with high prediction confidence and small classification difficulty, reduce their weights; for samples with low prediction confidence and large classification difficulty, increase their weights to ensure that subsequent weak learners pay more attention to samples that are difficult to classify. Therefore, the established weight adjustment formula is:

[0076]

[0077] ω i is the sample weight before adjustment, is the sample weight after adjustment, α is the weight of the weak learner, C i is the prediction confidence, D i is the classification difficulty, y i 、 are the true value and predicted value of the sample respectively.

[0078] Calculate the new sample weights through the above weight adjustment formula, and update the weights of each sample in the training set.

[0079] S35. Return to step S32, and train the next weak learner using the updated training set. In each iteration, use the prediction results of the previous weak learner to adjust the weights of each sample in the training set, and train a new weak learner until the preset number of iterations is reached.

[0080] Since the weak learners are set to rotate among linear classifiers, decision trees, perceptrons, and naive Bayes, the next weak learner after the linear classifier is the decision tree model. Therefore, the decision tree model can be trained using the training set updated in step S34. The number of training times for each weak learner can be set to 100 times. After reaching the training times, the second weak classifier is obtained, and then step S32 is returned to continue training the next weak learner, and so on until the training is completed.

[0081] All the trained weak learners are combined with weights to generate a fault prediction model. The weights of each weak learner are used to perform performance testing on the validation set, and metrics such as accuracy, recall rate, and F1 score are used to evaluate and test the prediction performance of the fault prediction model. If the performance of the model does not meet the requirements, the parameters of the Boosting algorithm (such as the number of iterations, learning rate, etc.) can be adjusted, or different weak learners can be used for training to obtain the best prediction effect and get the final fault prediction model.

[0082] Finally, the fault prediction model is converted into a format that the BMC module can understand and execute, and deployed to the BMC module. The integrated fault prediction model is tested and verified in the test environment to ensure that the fault prediction model can correctly receive the data provided by the BMC module and generate accurate prediction results.

[0083] S4. Server fault prediction is performed through the fault prediction model, and at the same time, the CPU occupancy of the BMC module is monitored, and the computing resources of the fault prediction model are dynamically scheduled according to the CPU occupancy rate.

[0084] S41. Write a script program in the BMC module to monitor the CPU occupancy in real time.

[0085] This script program is written in the Python language and imports the necessary modules:

[0086] The requests module, which is used to send HTTP requests;

[0087] The json module, which is used to parse and process data in JSON format;

[0088] The time module, which is used to control the polling interval.

[0089] S42. Define a function get_cpu_usage() to send a GET request to the API of the BMC module and parse the returned JSON format data to obtain the CPU occupancy.

[0090] S43. Set the address api_url and authentication information headers of the API of the BMC module, and set the polling interval poll_interval.

[0091] The API address is usually a URL pointing to a specific endpoint on the BMC server, and the authentication information includes the username, password, API key, etc., depending on the security settings of the BMC system. The polling interval determines the time interval between each poll, and the time interval is in seconds.

[0092] S44. Continuously poll the API of the BMC module according to the set time interval. In each loop, the get_cpu_usage function is called to obtain and return the CPU occupancy of the BMC module in real time.

[0093] S45. Calculate the resource allocation coefficient according to the CPU occupancy rate. Specifically, let the currently obtained CPU occupancy rate be U, and calculate the CPU occupancy rate deviation according to the CPU occupancy rate U:

[0094] D = |U * 100 - 50|

[0095] This CPU occupancy rate deviation D is the deviation between the current CPU occupancy rate (U) and the ideal range (taking 50% as the middle value), indicating the tightness of the current system resources.

[0096] Since the high and low thresholds may be different, it is necessary to normalize the deviation to the range of [0, 1] for subsequent calculations. Thus, calculate the normalized CPU occupancy rate deviation N:

[0097]

[0098] where L and H are the lower and upper limits of the preset CPU occupancy rate thresholds respectively, and the values can be L = 30% and H = 70%.

[0099] Calculate the resource allocation coefficient RAC according to the normalized CPU occupancy rate deviation N:

[0100] RAC = 1 - N 2

[0101] S46. Dynamically adjust the priority of the computing resource allocation of the fault prediction model according to the value of the resource allocation coefficient RAC.

[0102] Set a lower threshold and an upper threshold for the resource allocation coefficient RAC. For example, assume that the lower threshold of the resource allocation coefficient RAC is 0.3 and the upper threshold is 0.7. Then when RAC ∈ [0 - 0.3), allocate a lower priority to the fault prediction model; when RAC ∈ [0.3 - 0.7), allocate a medium priority to the fault prediction model; when RAC ∈ [0.7 - 1), allocate a higher priority to the fault prediction model.

[0103] Real-time detect the CPU occupancy rate, and calculate the value of the Resource Allocation Coefficient (RAC) in real time. Calculate the dynamic resource allocation based on the value of the RAC. When the value of the RAC is lower than the lower limit of the preset threshold, it means that the system may face performance bottlenecks or resource shortages. At this time, lower the priority of the fault prediction model, reduce its computing resource allocation, relieve the system burden and ensure that critical services obtain sufficient resource support. When the value of the RAC exceeds the upper limit of the preset threshold, it means that there are more idle resources available in the system. At this time, increase the priority of the fault prediction model to ensure that it can obtain sufficient computing resources for efficient fault prediction.

[0104] Specifically, the present invention sets three priorities: low, medium, and high. The computing resources corresponding to each priority are as follows:

[0105] I. Low priority (RAC value range: 0 - 0.3)

[0106] Number of threads: 1 (or set according to the minimum number of threads in the system)

[0107] Memory allocation: 512MB (set according to the basic running requirements of the module)

[0108] CPU time slice: Set the weight of the CPU time slice to 10% to ensure that other high-priority tasks obtain CPU resources first.

[0109] II. Medium priority (RAC value range: 0.3 - 0.7)

[0110] Number of threads: 2

[0111] Memory allocation: 1024MB

[0112] CPU time slice: Set the weight of the CPU time slice to 50%.

[0113] III. High priority (RAC value range: 0.7 - 1)

[0114] Number of threads: 4

[0115] Memory allocation: 2048MB

[0116] CPU time slice: Set the weight of the CPU time slice to 80%.

[0117] To avoid system instability caused by sudden changes in priority or resource allocation, a gradual adjustment strategy can be adopted. When reducing or increasing the priority of the fault prediction model, set the step size of thread adjustment to 1, that is, increase or decrease 1 thread each time until the number of threads corresponding to the priority is reached. The step size of memory is 256MB, that is, increase or decrease 256MB of memory each time until the memory corresponding to the priority is reached. The step size of the CPU time slice is 10%, that is, increase or decrease the CPU time slice weight by 10% each time until the time slice weight corresponding to the priority is reached.

[0118] S5. Perform iteration and update of the fault prediction model.

[0119] The present invention uses a database to persistently store the predicted data for subsequent model updates and evaluations.

[0120] Every 1 week, retrieve the predicted data from the database, and combine it with the latest real-time data to perform incremental learning updates on the fault prediction model. During the update process, introduce a forgetting mechanism:

[0121] Set a memory bank with a size of 1,000,000. After each task training is completed, collect 5% of the samples of the task and store them in the memory bank. Set the memory bank update strategy. When the memory bank is full, after each training is completed, randomly select the same number of old samples and replace them with new samples. When learning a new task, in addition to using the data of the current task, also randomly extract a part of the samples of the old tasks from the memory bank for training. Set a replay ratio r, which represents the proportion of the samples extracted from the memory bank in the entire training batch when training a new task. Here, it is set to 10%. Finally, obtain a new fault prediction model.

[0122] Compare the accuracy and robustness of the new fault prediction model and the fault prediction model running on the current BMC. If the new fault prediction model shows better performance in the evaluation, aiming for higher accuracy, lower false alarm rate, etc., then perform the following steps:

[0123] 1. Backup the model: Before replacement, back up the fault prediction model on the current BMC module so that it can be rolled back if needed.

[0124] 2. Replace the model: Replace the fault prediction model on the BMC module with the new fault prediction model.

[0125] Obtain the running status data of the target server in real time, preprocess it and input it into the new fault prediction model for real-time fault prediction.

[0126] S6. Design a notification and response mechanism.

[0127] Step S6 specifically includes the following sub-steps:

[0128] S61. In the BMC module, when the fault prediction model predicts that a fault is about to occur, define the event through event configuration to trigger a notification. The notification template includes key information such as fault description, occurrence time, impact scope, and urgency.

[0129] S62. Specify which people or systems the fault notification should be sent to by configuring a specific user group.

[0130] S63. For some common and automatically processable fault types, design an automated response mechanism to reduce the need for manual intervention.

[0131] Use built-in automated operations such as restarting services, isolating faulty nodes, switching backup devices, and executing repair scripts. Write external scripts or programs to handle faults and call these scripts or programs through the API or integration framework of the BMC system. For some complex faults or faults that require manual intervention, rely on system administrators or technical support teams for manual response.

[0132] Based on the above method embodiments, the present invention also proposes a server fault prediction system based on edge computing of the BMC module. The system includes:

[0133] Data acquisition module: used to deploy the BMC module at the edge node of the target server and obtain the historical operation status data of the target server through the BMC module;

[0134] Model training module: used to perform format parsing and preprocessing on the historical operation status data of the target server, construct a sample set for fault prediction; construct a fault prediction model based on ensemble learning, train the fault prediction model through the sample set, and deploy the trained fault prediction model to the BMC module;

[0135] Fault prediction module: used to perform server fault prediction through the fault prediction model, and at the same time monitor the CPU occupancy of the BMC module, and perform dynamic scheduling of the computing resources of the fault prediction model according to the CPU occupancy rate.

[0136] The above system embodiments and method embodiments correspond one by one. For the brief description of the system embodiments, please refer to the method embodiments.

[0137] The present invention also discloses an electronic device, including: at least one processor, at least one memory, a communication interface, and a bus; wherein, the processor, memory, and communication interface complete mutual communication through the bus; the memory stores program instructions executable by the processor, and the processor calls the program instructions to implement the method described above in the present invention.

[0138] The present invention also discloses a computer-readable storage medium, which stores computer instructions that enable a computer to implement all or part of the steps of the method described in the embodiments of the present invention. The storage medium includes various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs.

[0139] The system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be distributed over multiple network units. Those of ordinary skill in the art can, without creative effort, select some or all of the modules according to actual needs to achieve the objectives of the solutions of this embodiment.

[0140] The foregoing is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A server fault prediction method based on edge computing of BMC module, characterized in that The method includes: Deploy a BMC module at the edge node of the target server, and obtain the historical operation status data of the target server through the BMC module; Perform format parsing and preprocessing on the historical operation status data of the target server to construct a sample set for fault prediction; Construct a fault prediction model based on ensemble learning, train the fault prediction model through the sample set, and deploy the trained fault prediction model to the BMC module; Perform server fault prediction through the fault prediction model, and at the same time monitor the CPU occupancy of the BMC module, and perform dynamic scheduling of the computing resources of the fault prediction model according to the CPU occupancy rate; The obtaining of the historical operation status data of the target server through the BMC module specifically includes: Obtain the resource URL of the Redfish service of the target server through the BMC module, send a GET request to the resource URL of the target server, and obtain the historical operation status data of the target server; The historical operation status data of the target server includes power status, thermal status, processor utilization rate, memory usage, and storage device status; The performing of format parsing and preprocessing on the historical operation status data of the target server to construct a sample set for fault prediction specifically includes: Perform format parsing and preprocessing on the historical operation status data of the target server, and the preprocessing includes missing value processing, outlier detection and elimination; Obtain the fault type corresponding to the historical operation status data of the target server; Use the preprocessed historical operation status data as sample attributes and the fault type as sample labels to construct training samples and perform data augmentation to obtain a sample set; The weak learners of ensemble learning include linear classifiers, decision tree models, perceptron models, and naive Bayes models. Assign a weight to each weak learner, and perform weighted combination of different weak learners to form a strong classifier as the fault prediction model; The training of the fault prediction model through the sample set specifically includes: Assign the same initial weight to each sample in the sample set to obtain a training set; Train a weak learner through the training set to obtain a weak classifier; Input each sample attribute in the training set into the weak classifier for prediction, use the maximum value of the prediction probability as the prediction confidence, and use the prediction error of each sample as the classification difficulty; Establish a weight adjustment formula based on the prediction confidence and classification difficulty, and update the weight of each sample in the training set; Train the next weak learner through the updated training set. In each iteration, use the prediction results of the previous weak learner to adjust the weight of each sample in the training set, and train a new weak learner until the preset number of iterations is reached; The weight adjustment formula is: ω i is the sample weight before adjustment, is the sample weight after adjustment, α is the weight of the weak learner, C i is the prediction confidence, D i is the classification difficulty, y i 、 are the true value and the predicted value of the sample, respectively; The performing of dynamic scheduling of the computing resources of the fault prediction model according to the CPU occupancy rate specifically includes: Calculate the normalized CPU occupancy rate deviation N according to the CPU occupancy rate: Wherein, L and H are respectively the lower and upper limits of the preset CPU occupancy rate threshold, D = |U * 100 - 50|, and U is the currently obtained CPU occupancy rate; Calculate the resource allocation coefficient RAC according to the normalized CPU occupancy rate deviation N: RAC = 1 - N 2 Dynamically adjust the priority of computing resource allocation for the fault prediction model according to the value of the Resource Allocation Coefficient (RAC).

2. A server fault prediction system based on edge computing of BMC modules, which is used to implement the server fault prediction method based on edge computing of BMC modules described in claim 1, and is characterized in that, The system includes: Data acquisition module: used to deploy the BMC module at the edge node of the target server, and obtain the historical operation status data of the target server through the BMC module; Model training module: used to perform format parsing and preprocessing on the historical operation status data of the target server, construct a sample set for fault prediction; construct a fault prediction model based on ensemble learning, train the fault prediction model through the sample set, and deploy the trained fault prediction model to the BMC module; Fault prediction module: used to perform server fault prediction through the fault prediction model, and at the same time monitor the CPU occupancy of the BMC module, and perform dynamic scheduling of the computing resources of the fault prediction model according to the CPU occupancy rate.

3. An electronic device, characterized in that, Including: At least one processor, at least one memory, a communication interface, and a bus; Wherein, the processor, memory, and communication interface complete mutual communication through the bus; The memory stores program instructions executable by the processor, and the processor invokes the program instructions to implement the method according to claim 1.

4. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions, and the computer instructions cause the computer to implement the method according to claim 1.

Citation Information

Patent Citations

  • Server fault detection method and device, computer equipment and storage medium

    CN113835962A