Distributed Data Center Operation and Maintenance Fault Prediction Method Based on Deep Learning
The deep learning-based method for distributed data center fault prediction addresses the issue of inadequate parameter analysis by monitoring multiple performance metrics, enabling accurate and proactive fault detection.
Patent Information
- Application Number
- CN202510032178.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-09
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-01-09
AI Technical Summary
In operation and maintenance fault detection of distributed data centers, single analysis parameters cause the analysis process to be insufficiently comprehensive and cannot achieve high-accurate predictions.
By monitoring the disk response time of distributed nodes, combining multiple parameters such as read and write rate, CPU usage and memory usage, multi-level and comprehensive monitoring and numerical analysis are carried out, and the packet loss rate is calculated using a network tester to generate an upcoming load signal and fault prediction.
It realizes accurate identification and failure prediction of distributed data center nodes, provides rich data support, reduces the risk of service interruption, and ensures the stable operation of the data center.
Smart Images

Figure CN120011118B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data centers, and specifically to a distributed data center operation and maintenance fault prediction method based on deep learning. Background Art
[0002] A distributed data center is a data center architecture that disperses data storage and processing functions across multiple geographical locations and physical facilities. Computational tasks are assigned to multiple computing nodes for parallel processing. Each node has a certain computing power and collaborates to complete complex computational tasks. Taking large-scale data analysis tasks as an example, data can be divided into multiple parts, analyzed and calculated on different nodes respectively, and finally the results are summarized, greatly shortening the processing time.
[0003] The application with the publication number CN115373879A discloses a disk fault prediction method for intelligent operation and maintenance of large-scale cloud data centers, including: first, performing information entropy feature processing on unbalanced data to select more important features; then dividing the processed unbalanced data to extract sample data of the minority class, that is, fault samples; then using the time progressive sampling method TPS to enhance the data of the fault samples to generate synthetic data, generating more fault sample data through TPS, so that the ratio between the number of healthy samples and the number of fault samples will reach a better balance; then merging the synthetic data with good generation effect and the original data to generate integrated data; finally, inputting the integrated data into the disk fault prediction model for training, and selecting a time window of 7 days to predict whether a fault will occur after 7 days and perform corresponding data marking.
[0004] When performing operation and maintenance fault detection on its distributed data center, it generally determines the response delay of the corresponding distributed node based on the data load and operation conditions of the corresponding data center, so as to comprehensively determine the specific fault conditions of the corresponding distributed center. However, in the actual analysis and processing process, there are multiple groups of parameters related to operation and maintenance faults for distributed nodes. If only analyzed individually, the analysis process will be incomplete and it is impossible to achieve better prediction accuracy. Summary of the Invention
[0005] Aiming at the deficiencies of the prior art, the present invention provides a distributed data center operation and maintenance fault prediction method based on deep learning, which solves the problem that there are multiple groups of parameters related to operation and maintenance faults for distributed nodes. If only analyzed individually, the analysis process will be incomplete and it is impossible to achieve better prediction accuracy.
[0006] To achieve the above objectives, the present invention is realized through the following technical solutions: A distributed data center operation and maintenance fault prediction method based on deep learning, including the following steps:
[0007] Step 1: Monitor the disk response times of several different distributed nodes in the distributed data center, and based on the real-time monitored disk response times, lock the nodes with abnormal responses. The specific method is as follows:
[0008] Monitor the disk response times of different distributed nodes and label them as X i , where i represents different distributed nodes;
[0009] Take the disk response time X associated with the corresponding distributed node i and compare it with the preset standard Y i . Here, Y i is the preset standard, and different distributed nodes have different preset standards: if X i ≤Y i , no processing is required and continuous monitoring can be carried out. If X i >Y i , then directly label this distributed node as an abnormal node;
[0010] Step 2: Monitor multiple parameters of the labeled abnormal nodes. The multiple parameters to be monitored include read / write rate parameters, CPU usage rate parameters, and memory occupancy rate parameters. Based on the real-time monitoring process of multiple parameters, select abnormal monitoring items and confirm the analysis period based on the selected abnormal monitoring items. The specific sub-steps are as follows:
[0011] S21: Monitor multiple parameters of the labeled abnormal nodes, confirm the monitoring parameters of different monitoring items in real time, and perform real-time verification on the monitoring parameters of different monitoring items: Identify the standard interval associated with the corresponding monitoring item. The standard interval is a preset interval, and different monitoring items correspond to different preset intervals. Label the monitoring parameters confirmed at the same time for different monitoring items as C k , where k represents different monitoring items. If C k ∉ the standard interval, then label this monitoring item as an abnormal monitoring item and label the current moment as an abnormal moment. If C k ∈ the standard interval, then continue monitoring;
[0012] S22: Take the determined abnormal moment as the initial moment and confirm a set of analysis periods. The duration of the analysis period is a preset duration;
[0013] S23: After the end of this analysis period, use the same method as in steps S21 - S22 to confirm the next set of analysis periods;
[0014] Step 3: Based on the determined analysis period, perform numerical analysis on multiple parameters monitored in real time during the analysis period. According to the change trends of different monitoring items and different monitoring parameters, confirm the prediction parameters associated with the corresponding monitoring item at the next moment, and based on the different prediction parameters confirmed for different monitoring items, perform load prediction, and confirm the upcoming load signal based on the prediction results. The specific sub-steps are as follows:
[0015] S31. Real-time confirm the different monitoring parameters associated with different monitoring items during the analysis period, and calibrate the different real-time confirmed monitoring parameters as CS k-t and CS k-(t-1) , where k represents different monitoring items, t represents the current moment, and (t - 1) represents the previous moment. Use: CZ k-t =CS k-t -CS k-(t-1) to confirm the change difference CZ of the monitoring parameters associated with the current moment k-t ;
[0016] S32. Confirm the multiple groups of change differences CZ k-t generated in real time for different monitoring items during the analysis period, lock CZ k- t min and CZ k-t max to confirm the difference interval. Based on the currently confirmed monitoring parameter CS k-t and the determined difference interval, predict the monitoring parameter interval associated with the next moment. Use: QS k min=CS k-t + the minimum value of the difference interval and QS k max=CS k-t + the maximum value of the difference interval. Based on the confirmed QS k min and QS k max, generate the monitoring parameter interval associated with the next moment for the corresponding monitoring item;
[0017] S33. Based on the monitoring parameter interval [QS k min, QS k max] predicted for different monitoring items regarding the next moment, use: FZ k min=QS k min×A k and FZ k max=QS k max×A k to confirm the load rate interval [FZ k min, FZ k max] associated with the current monitoring item at the next moment, where A k is the preset factor associated with the corresponding monitoring item;
[0018] S34. Sum the different load rate intervals predicted by different monitoring items to confirm the total load rate interval [FZmin, FZmax], and check the confirmed total load rate interval [FZmin, FZmax]:
[0019] Check whether there is a partial interval segment exceeding 90% within the total load rate interval. If so, confirm the specific occupancy ratio ZB of the partial interval segment in the total load rate interval. If ZB < Y1, where Y1 is a preset value, continue monitoring and confirm the specific occupancy ratio. If ZB ≥ Y1, generate an upcoming load signal and display the generated upcoming load signal. If not, continue monitoring and confirm the partial interval segment;
[0020] Step 4. Based on the confirmed upcoming load signal, perform a packet loss rate test on this response abnormal node, and perform a fault prediction on this response abnormal node based on the test result. The specific sub-steps for performing the fault prediction are as follows:
[0021] S41. Connect the port of the network tester to the interface of the response abnormal node to be tested, and set the number, size, and sending rate of the data packets to be sent in the network tester; start the test. The network tester sends data packets to the response abnormal node according to the set parameters, and simultaneously records the number of data packets sent SL1;
[0022] S42. The network tester receives the data packets returned from the response abnormal node or forwarded by the node, counts the number of received data packets SL2, and uses: DB = (SL2 - SL1) ÷ SL1 × 100% to confirm the packet loss rate DB of this response abnormal node. Compare the packet loss rate DB with the preset value Y2. If DB > Y2, it means that the packet loss rate of this response abnormal node is abnormal, and mark this response abnormal node as a faulty node and display the marked faulty node; if DB ≤ Y2, it means that the packet loss rate of this response abnormal node is not abnormal, and no processing is performed, and continue to perform the packet loss rate test on this response abnormal node.
[0023] The present invention provides a distributed data center operation and maintenance fault prediction method based on deep learning. Compared with the prior art, it has the following beneficial effects:
[0024] The present invention deeply monitors multiple key parameters including read-write rate, CPU usage rate, and memory occupancy rate for the marked abnormal nodes. According to the preset standard interval, these parameters are verified in real time, which can not only accurately identify abnormal monitoring items but also cleverly determine the analysis period. This all-round and multi-level monitoring and analysis mechanism can insight into the node operation state from multiple dimensions and provide rich and reliable data support for accurately judging the node load situation;
[0025] During a determined analysis period, through numerical analysis of multiple parameters, precise formulas are used to calculate the difference in parameter changes, predict parameter intervals, and load rate intervals. This data-driven prediction method fully considers the dynamic change characteristics of parameters, can predict the load situation of nodes in advance, generate upcoming load signals in a timely manner and display them, helping operation and maintenance personnel plan response strategies in advance, and effectively reducing the risk of service interruption caused by excessive load;
[0026] Based on the upcoming load signal, a packet loss rate test is carried out on nodes with abnormal responses. By scientifically setting parameters with a network tester, the packet loss rate is accurately calculated and compared with a preset value, so as to reliably determine whether the node is a faulty node. This process provides a direct and crucial basis for fault prediction, enabling operation and maintenance personnel to carry out hardware updates or parameter adjustments in a targeted manner, and ensuring the stable operation of the data center. Brief Description of the Drawings
[0027] Figure 1 It is a schematic flowchart of the method of the present invention;
[0028] Figure 2 It is a schematic flowchart of the packet loss rate test of the present invention. Detailed Embodiments
[0029] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0030] First Embodiment
[0031] Please refer to Figure 1 , this application provides a distributed data center operation and maintenance fault prediction method based on deep learning, including the following steps:
[0032] Step 1: Monitor the disk response time of several different distributed nodes in the distributed data center, and based on the real-time monitored disk response time, lock the nodes with abnormal responses. Specifically, the disk response times associated with different distributed nodes are relatively the same, and there are specific set standards. When the corresponding disk response time is too long, it means that the corresponding node has an abnormal response situation, and it is necessary to carry out relevant calibration of the nodes with abnormal responses;
[0033] Among them, the specific method of locking the nodes with abnormal responses is:
[0034] Monitor and calibrate the disk response time of different distributed nodes as X i , where i represents different distributed nodes;
[0035] Compare the disk response time X associated with the corresponding distributed node i with the preset standard Y i where Y i is the preset standard. Different distributed nodes have different preset standards, which are all formulated in advance by relevant personnel based on experience: If X i ≤Y i , then no processing is required and continuous monitoring can be carried out. If X i >Y i , then directly label this distributed node as an abnormal node (indicating that the response time does not meet the standard, and there may be abnormal related data inside).
[0036] Step 2: Monitor multiple parameters of the calibrated abnormal node. The multiple parameters to be monitored include read-write rate parameters, CPU usage rate parameters, and memory occupancy rate parameters. Based on the real-time monitoring process of multiple parameters, select abnormal monitoring items, and confirm the analysis period based on the selected abnormal monitoring items. The specific sub-steps for determining the analysis period are as follows:
[0037] S21: Monitor multiple parameters of the calibrated abnormal node, and confirm the monitoring parameters of different monitoring items in real time (since different monitoring items include read-write rate, CPU usage rate, and memory occupancy rate, three groups of monitoring change curves are generated in real time). Perform real-time verification on the monitoring parameters of different monitoring items: Identify the standard interval associated with the corresponding monitoring item. The standard interval is a preset interval, formulated by relevant operators based on experience. Different monitoring items correspond to different preset intervals. Calibrate the monitoring parameters confirmed at the same moment for different monitoring items as C k , where k represents different monitoring items. If C k ∈ standard interval, then continue monitoring. If C k ∉ standard interval, then label this monitoring item as an abnormal monitoring item and label the current moment as an abnormal moment;
[0038] S22: Take the determined abnormal moment as the initial moment and confirm a set of analysis periods. The duration of the analysis period is a preset duration, formulated by relevant personnel based on experience, generally taking a value of 20 min;
[0039] S23: After the end of this analysis period, use the same method as in steps S21 - S22 to confirm the next set of analysis periods;
[0040] Specifically, in the determined analysis period, perform numerical analysis on the relevant parameters of each different monitoring item, and based on the parameter trend changes of different monitoring items, evaluate whether this node will have a load situation and display relevant load signals in a timely manner;
[0041] Step 3: Based on the determined analysis period, perform numerical analysis on multiple parameters monitored in real time during the analysis period. According to the changing trends of different monitoring items and different monitoring parameters, confirm the predicted parameters associated with the corresponding monitoring items at the next moment. And based on the different predicted parameters confirmed for different monitoring items, perform load prediction, and confirm the upcoming load signal based on the prediction results. Specifically, in the actual analysis and processing process, if the parameters monitored by each monitoring item show significant changing trends, then the corresponding overloading situation will occur soon. Therefore, it is necessary to perform load prediction in advance, confirm the corresponding load signal in time and display it for relevant external operators to process in time and take relevant countermeasures in time;
[0042] Among them, the specific sub-steps for performing load prediction are as follows:
[0043] S31: Real-time confirm the different monitoring parameters associated with different monitoring items during the analysis period, and calibrate the different monitored parameters confirmed in real time as CS k-t and CS k-(t-1) , where k represents different monitoring items, t represents the current moment, and (t - 1) represents the previous moment. Use: CZ k-t =CS k-t -CS k-(t-1) to confirm the change difference CZ k-t of the monitoring parameter associated with the current moment;
[0044] S32: Confirm the multiple groups of change differences CZ k-t generated in real time for different monitoring items during the analysis period, lock CZ k- t min and CZ k-t max to confirm the difference interval (continuously update the difference interval in real time as the analysis period continues). Based on the currently confirmed monitoring parameter CS k-t and the determined difference interval, predict the monitoring parameter interval associated with the next moment. Use: QS k min=CS k-t + the minimum value of the difference interval and QS k max=CS k-t + the maximum value of the difference interval. Based on the confirmed QS k min and QS k max, generate the monitoring parameter interval associated with the next moment for the corresponding monitoring item;
[0045] S33: Based on the monitoring parameter intervals [QS k min, QS k max] predicted for different monitoring items regarding the next moment, use: FZ k min=QS k min×Ak and FZ k max = QS k max × A k Confirm the load rate interval [FZ k min, FZ k max] to which the current monitored item is associated at the next moment, where A k is the preset factor associated with the corresponding monitored item, and the preset factor is determined in advance by the operator according to experience;
[0046] S34. Sum up the different load rate intervals predicted by different monitored items, confirm the total load rate interval [FZmin, FZmax], and check the confirmed total load rate interval [FZmin, FZmax]:
[0047] Confirm whether there is a partial interval segment exceeding 90% within the total load rate interval. If so, confirm the specific occupancy ratio ZB of the partial interval segment in the total load rate interval. If ZB < Y1, where Y1 is a preset value and its specific value is determined by the operator according to experience, then continuously monitor and confirm the specific occupancy ratio. If ZB ≥ Y1, generate an upcoming load signal and display the generated upcoming load signal for external personnel to view;
[0048] If not, continuously monitor and confirm the partial interval segment.
[0049] By continuously confirming and calculating the monitoring parameters in real time, the change difference of each monitored item at different moments can be accurately obtained; for example, in S31, through a clear formula, the calculation result is precise; this precision helps to promptly detect abnormal changes in the monitoring parameters and provides a reliable basis for subsequent prediction;
[0050] Over time, the monitoring parameter interval will be dynamically adjusted according to the changes in real-time data and the difference interval. This dynamic adjustment makes the prediction more in line with the actual situation and improves the accuracy of the prediction;
[0051] By calculating the load rate interval, the load conditions of different monitored items at different moments can be effectively predicted; in S33, according to the preset factor and combined with the monitoring parameter interval, the load rate interval is obtained. This prediction method has high reliability and can provide a reference for actual decision-making;
[0052] By comprehensively considering factors such as the parameter changes of different monitored items, the monitoring parameter interval, and the load rate interval, the prediction result is more comprehensive; through the analysis and integration of each part, the operating state of the distributed data center can be grasped as a whole, providing strong support for operation and maintenance decision-making.
[0053] Second Embodiment
[0054] In the specific implementation process of this embodiment, compared with the above embodiment, this embodiment mainly focuses on the fault prediction of the corresponding response abnormal node, and the specific prediction steps include:
[0055] Step Four: Combine Figure 2 , based on the confirmed upcoming load signal, perform a packet loss rate test on this response abnormal node, and perform fault prediction on this response abnormal node based on the test results. The specific sub-steps for performing fault prediction are as follows:
[0056] S41. Connect the port of the network tester to the interface of the response abnormal node to be tested to ensure stable connection; if testing a node in the network link, it may be necessary to connect the tester in series to the link. Set parameters such as the number of data packets to be sent, size, and sending rate in the network tester; generally, different data packet sizes can be selected, such as 64 bytes, 1518 bytes, etc., to simulate different network load conditions. The set parameters are all preset parameters, which are set in advance by relevant operators. Start the test. The network tester sends data packets to the response abnormal node according to the set parameters, and at the same time records the number of data packets sent SL1;
[0057] S42. The network tester receives the data packets returned from the response abnormal node or forwarded through this node, counts the number of received data packets SL2, and uses: DB = (SL2 - SL1) ÷ SL1 × 100% to confirm the packet loss rate DB of this response abnormal node. Compare the packet loss rate DB with the preset value Y2, where the specific value of Y2 is determined by the operator according to experience. If DB > Y2, it means that the packet loss rate of this response abnormal node is abnormal, and this response abnormal node is marked as a faulty node, and the marked faulty node is displayed, and relevant personnel perform hardware updates or set relevant parameters to reduce the packet loss rate. If DB ≤ Y2, it means that the packet loss rate of this response abnormal node is not abnormal, and no processing is performed, and the packet loss rate test on this response abnormal node continues;
[0058] Specifically, during the packet loss rate test, there is an associated test cycle, which is a preset cycle, and the start time of the test cycle is the initial test time of the network tester, and the end time is extended by a certain duration after the end time of the network tester's test. The certain duration is a preset duration, which is determined in advance by relevant operators, generally within 1 minute.
[0059] Some of the data in the above formula are numerically calculated after removing their dimensions, and the content not described in detail in this specification belongs to the prior art well-known to those skilled in the art.
[0060] The above embodiments are only used to illustrate the technical method of the present invention and not to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical method of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical method of the present invention.
Claims
1. A method for predicting operation and maintenance faults in a distributed data center based on deep learning, characterized in that, It includes the following steps: Step 1: Monitor the disk response times of several different distributed nodes in the distributed data center, and calibrate the nodes with abnormal responses based on the real-time monitored disk response times; Step 2: Monitor multiple parameters of the calibrated abnormal nodes. The multiple parameters to be monitored include read / write rate parameters, CPU usage rate parameters, and memory occupancy rate parameters. Based on the real-time monitoring process of the multiple parameters, select the abnormal monitoring items, and confirm the analysis period based on the selected abnormal monitoring items; Step 3: Based on the determined analysis period, perform numerical analysis on the multiple parameters monitored in real time during the analysis period. According to the change trends of different monitoring parameters for different monitoring items, confirm the predicted parameters associated with the corresponding monitoring items at the next moment, and perform load prediction based on the different predicted parameters confirmed for different monitoring items. And based on the prediction results, confirm the upcoming load signal. The specific sub-steps are as follows: S31. Real-time confirm different monitoring parameters associated with different monitoring items within the analysis period, and calibrate the different monitoring parameters obtained from the real-time confirmation as CS k-t and CS k-(t-1) , where k represents different monitoring items, t represents the current moment, and (t - 1) represents the previous moment. Use: CZ k-t = CS k-t - CS k-(t-1) to confirm the change difference CZ of the monitoring parameters associated with the current moment k-t ; S32. Confirm multiple groups of change differences CZ generated in real time for different monitoring items during the analysis period, and lock CZ k-t . Confirm the difference range of CZ k-t min and CZ k-t max, and based on the currently confirmed monitoring parameter CS k-t and the determined difference range, predict the monitoring parameter range associated with the next moment, using: QS k min = CS k-t + the minimum value of the difference range and QS k max = CS k-t + the maximum value of the difference range. Based on the confirmed QS k min and QS k max, generate the monitoring parameter range associated with the next moment for the corresponding monitoring item; S33. Monitoring parameter intervals for the next moment predicted based on different monitoring items [QS k min, QS k max], adopt: FZ k min = QS k min × A k and FZ k max = QS k max × A k to confirm the load rate interval [FZ k min, FZ k max] associated with the current monitoring item at the next moment, where A k is the preset factor associated with the corresponding monitoring item; S34: Sum up the different load rate intervals predicted for different monitoring items, confirm the total load rate interval [FZmin, FZmax], and check the confirmed total load rate interval [FZmin, FZmax]: Check whether there is a partial interval segment exceeding 90% within the total load rate interval. If so, confirm the specific occupancy ratio ZB of the partial interval segment in the total load rate interval. If ZB < Y1, where Y1 is a preset value, continue to monitor and confirm the specific occupancy ratio. If ZB ≥ Y1, generate an upcoming load signal and display the generated upcoming load signal; Step 4: Based on the confirmed upcoming load signal, perform a packet loss rate test on the nodes with abnormal responses, and perform a fault prediction on the nodes with abnormal responses based on the test results.
2. The method for predicting operation and maintenance faults of a distributed data center based on deep learning according to claim 1, wherein In the above Step 1, the specific method for calibrating the nodes with abnormal responses is: Monitor the disk response time of different distributed nodes and calibrate it as X i , where i represents different distributed nodes; Compare the disk response time X associated with the corresponding distributed node i with the preset standard Y i where Y i is the preset standard, and different distributed nodes have different preset standards: if X i ≤Y i , no processing is required and continuous monitoring can be carried out. If X i >Y i , then directly mark this distributed node as an abnormal node.
3. The method for predicting operation and maintenance faults of a distributed data center based on deep learning according to claim 1, wherein In the above Step 2, the specific sub-steps for determining the analysis period are: S21. Monitor multiple parameters of the calibrated abnormal nodes, confirm the monitoring parameters of different monitoring items in real time, and perform real-time verification on the monitoring parameters of different monitoring items: identify the standard intervals associated with the corresponding monitoring items, where the standard intervals are preset intervals, and different monitoring items correspond to different preset intervals. Calibrate the monitoring parameters confirmed for different monitoring items at the same moment as C k , where k represents different monitoring items. If C k ∉ the standard interval, then calibrate this monitoring item as an abnormal monitoring item and mark the current moment as an abnormal moment; S22: Use the determined abnormal moment as the initial moment to confirm a set of analysis periods, and the duration of the analysis period is the preset duration; S23: After the end of this analysis period, use the same method as in Steps S21 - S22 to confirm the next set of analysis periods.
4. The method for predicting operation and maintenance faults of a distributed data center based on deep learning according to claim 3, wherein, In step S21, if C k ∈ the standard interval, continuous monitoring is carried out.
5. The method for predicting operation and maintenance faults of a distributed data center based on deep learning according to claim 1, wherein In the above Step S34, if there is no partial interval segment exceeding 90% within the total load rate interval, continue to monitor and confirm the partial interval segment.
6. The method for predicting operation and maintenance faults of a distributed data center based on deep learning according to claim 1, characterized in that, In the above Step 4, the specific sub-steps for performing the fault prediction are: S41: Connect the port of the network tester to the interface of the node with abnormal response to be tested, and set the number, size, and sending rate of the data packets to be sent in the network tester; Start the test. The network tester sends data packets to the node with abnormal response according to the set parameters, and simultaneously record the number of data packets sent SL1; S42. The network tester receives the data packets returned from the response abnormal node or forwarded by the node, counts the number SL2 of the received data packets, and uses: DB = (SL2 - SL1) ÷ SL1 × 100% to confirm the packet loss rate DB of this response abnormal node. The packet loss rate DB is compared with the preset value Y2. If DB > Y2, it means that the packet loss rate of this response abnormal node is abnormal, and this response abnormal node is calibrated as a faulty node, and the calibrated faulty node is displayed.
7. The method for predicting operation and maintenance faults of a distributed data center based on deep learning according to claim 6, characterized in that, In the step S42, if DB ≤ Y2, it means that the packet loss rate of this response abnormal node is not abnormal, no processing is performed, and the packet loss rate test for this response abnormal node continues.
Citation Information
Patent Citations
Disk fault prediction method for intelligent operation and maintenance of large-scale cloud data center
CN115373879A
Server load prediction method and device, electronic equipment and storage medium
CN113535530A
Cloud computing physical node load monitoring method and device, terminal and storage medium
CN113626282A