Distributed data center operation and maintenance fault prediction method based on deep learning
By adopting a deep learning-based approach in distributed data centers, monitoring and analyzing multiple key parameters, the problem that a single analysis cannot fully understand the parameters related to distributed nodes and operation and maintenance failures is solved, and higher prediction accuracy and lower risk of service interruption is achieved.
Patent Information
- Application Number
- CN202510032178.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-09
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-01-09
AI Technical Summary
When detecting operation and maintenance faults in distributed data centers, single analysis cannot fully understand the multiple sets of parameters related to distributed nodes and operation and maintenance faults, resulting in low prediction accuracy.
Deep learning-based method is used to monitor the disk response time of different nodes in a distributed data center, calibrate abnormal nodes, and conduct real-time monitoring of multiple parameters (read and write rate, CPU usage, memory usage) on these nodes. Through numerical analysis and load prediction, an impending load signal is generated and fault prediction is performed through packet loss rate testing.
It realizes all-round and multi-level monitoring and analysis of distributed data center operation and maintenance failures, improves prediction accuracy, generates load signals in a timely manner, and reduces the risk of service interruption caused by excessive load.
Smart Images

Figure CN120011118A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data center technology, and specifically to a distributed data center operation and maintenance fault prediction method based on deep learning. Background Art
[0002] A distributed data center is a data center architecture that distributes data storage and processing functions across multiple geographic locations and physical facilities. Computing tasks are assigned to multiple computing nodes for parallel processing. Each node has a certain computing power and works together to complete complex computing tasks. For example, in large-scale data analysis tasks, data can be divided into multiple parts, analyzed and calculated on different nodes, and the results are finally aggregated, which greatly shortens the processing time.
[0003] The application with publication number CN115373879A discloses a disk failure prediction method for intelligent operation and maintenance of large-scale cloud data centers, including: first, processing the unbalanced data with information entropy features to select more important features; then dividing the processed unbalanced data to extract sample data of a small number of categories, namely, fault samples; then using the time progressive sampling method TPS to perform data enhancement on the fault samples to generate synthetic data, and generating more fault sample data through TPS so that the ratio between the number of healthy samples and the number of fault samples will achieve a better balance; then merging the synthetic data with good generation effect with the original data to generate integrated data; finally, inputting the integrated data into the disk failure prediction model for training, and selecting a time window of 7 days to predict whether a failure will occur after 7 days, and performing corresponding data marking.
[0004] When performing operation and maintenance fault detection, its distributed data center generally determines the response delay of the corresponding distributed node based on the data load and computing conditions of the corresponding data center, so as to comprehensively determine the specific fault situation of the corresponding distributed center. However, in the actual analysis and processing process, there are multiple groups of parameters related to distributed nodes and operation and maintenance faults. If only a single analysis is performed, the analysis process will be incomplete and unable to achieve better prediction accuracy. Summary of the invention
[0005] In view of the shortcomings of the prior art, the present invention provides a distributed data center operation and maintenance fault prediction method based on deep learning, which solves the problem that there are multiple groups of parameters related to distributed nodes and operation and maintenance faults. If only a single analysis is performed, the analysis process will be not comprehensive enough and better prediction accuracy cannot be achieved.
[0006] To achieve the above objectives, the present invention is implemented through the following technical solutions: A distributed data center operation and maintenance fault prediction method based on deep learning, comprising the following steps:
[0007] Step 1: Monitor the disk response time of several different distributed nodes in the distributed data center, and lock the abnormal response node based on the real-time monitored disk response time. The specific method is as follows:
[0008] Monitor the disk response time of different distributed nodes and mark it as X i , where i represents different distributed nodes;
[0009] The disk response time associated with the corresponding distributed node is X i With the preset standard Y i Check, where Y i is the preset standard. Different distributed nodes have different preset standards: if X i ≤Y i , no processing is performed, and continuous monitoring is sufficient. If X i >Y i , then this distributed node is directly marked as an abnormal node;
[0010] Step 2: Monitor multiple parameters of the calibrated abnormal nodes. The monitored parameters include read and write rate parameters, CPU usage parameters, and memory occupancy parameters. Based on the real-time monitoring process of multiple parameters, select abnormal monitoring items, and confirm the analysis period based on the selected abnormal monitoring items. The specific sub-steps are:
[0011] S21, perform multiple parameter monitoring on the calibrated abnormal node, confirm the monitoring parameters of different monitoring items in real time, and perform real-time verification on the monitoring parameters of different monitoring items: identify the standard interval associated with the corresponding monitoring item, the standard interval is the preset interval, different monitoring items correspond to different preset intervals, and calibrate the monitoring parameters confirmed by different monitoring items at the same time as C k , where k represents different monitoring items. If C k standard interval, then this monitoring item is marked as an abnormal monitoring item, and the current moment is marked as an abnormal moment. If C k ∈ standard interval, then continue monitoring;
[0012] S22, using the determined abnormal time as the initial time, confirming a set of analysis time periods, wherein the duration of the analysis time periods is the preset duration;
[0013] S23, after the analysis period ends, the next set of analysis periods is confirmed in the same manner as steps S21-S22;
[0014] Step 3: Based on the determined analysis period, perform numerical analysis on multiple parameters monitored in real time during the analysis period, confirm the prediction parameters associated with the corresponding monitoring items at the next moment according to the change trends of different monitoring parameters of different monitoring items, and perform load prediction based on different prediction parameters confirmed by different monitoring items, and confirm the upcoming load signal based on the prediction results. The specific sub-steps are:
[0015] S31, confirm the different monitoring parameters associated with different monitoring items during the analysis period in real time, and mark the different monitoring parameters confirmed in real time as CS k-t and CS k-(t-1) , k represents different monitoring items, where t represents the current time, (t-1) represents the previous time, using: CZ k-t =CS k-t -CS k-(t-1) Confirm the change difference CZ of the monitoring parameter associated with the current moment k-t ;
[0016] S32, multiple groups of change differences CZ generated in real time for different monitoring items during the analysis period k-t Confirm and lock CZ k- t min and CZ k-t max confirmed difference interval, based on the currently confirmed monitoring parameter CS k-t And the determined difference interval, predict the monitoring parameter interval associated with the next moment, using: QS k min=CS k-t + Minimum difference interval and QS k max=CS k-t + Maximum value of the difference interval, based on the confirmed QS k min and QS k max generates the monitoring parameter interval associated with the corresponding monitoring item at the next moment;
[0017] S33, based on the different monitoring items predicted monitoring parameter interval at the next moment [QS k min, QS k max], using: FZ k min=QS k min×C k and FZ k max=QS k max×C k Confirm the load rate interval [FZ k min, FZ k max], where C k It is the preset factor associated with the corresponding monitoring item;
[0018] S34, summing up the different load rate intervals predicted by different monitoring items, confirming the total load rate interval [FZmin, FZmax], and checking the confirmed total load rate interval [FZmin, FZmax]:
[0019] Confirm whether there is a partial interval exceeding 90% in the total load rate interval. If so, confirm the specific proportion value ZB of the partial interval in the total load rate interval. If ZB<Y1, where Y1 is a preset value, continue to monitor and confirm the specific proportion value. If ZB≥Y1, generate a load signal and display the generated load signal. If not, continue to monitor and confirm the partial interval.
[0020] Step 4: Based on the confirmed upcoming load signal, a packet loss rate test is performed on the abnormal response node, and a fault prediction is performed on the abnormal response node based on the test result. The specific sub-steps of the fault prediction are:
[0021] S41, connect the port of the network tester to the interface of the abnormal response node to be tested, set the number, size and sending rate of the data packets to be sent in the network tester; start the test, the network tester sends data packets to the abnormal response node according to the set parameters, and records the number of data packets sent SL1 at the same time;
[0022] S42. The network tester receives data packets returned from the response abnormal node or forwarded through the node, counts the number of received data packets SL2, and uses: DB = (SL2-SL1) ÷ SL1 × 100% to confirm the packet loss rate DB of this response abnormal node, and compares the packet loss rate DB with the preset value Y2. If DB≤Y2, it means that the packet loss rate of this response abnormal node is abnormal, and this response abnormal node is marked as a faulty node, and the marked faulty node is displayed; if DB>Y2, it means that the packet loss rate of this response abnormal node is not abnormal, and no processing is performed, and the packet loss rate test of this response abnormal node is continuously performed.
[0023] The present invention provides a distributed data center operation and maintenance fault prediction method based on deep learning. Compared with the prior art, it has the following beneficial effects:
[0024] The present invention targets calibrated abnormal nodes and deeply monitors multiple key parameters including read and write rates, CPU usage, and memory occupancy. These parameters are verified in real time according to the preset standard interval, which can not only accurately identify abnormal monitoring items, but also cleverly determine the analysis period. This all-round, multi-level monitoring and analysis mechanism can provide insights into the node operation status from multiple dimensions, providing rich and reliable data support for accurately judging the node load situation;
[0025] During a certain analysis period, through numerical analysis of multiple parameters, accurate formulas are used to calculate parameter change differences, predict parameter intervals, and load rate intervals. This data-driven prediction method fully considers the dynamic change characteristics of parameters, can predict the load of nodes in advance, and timely generate and display upcoming load signals, helping operation and maintenance personnel to plan response strategies in advance and effectively reduce the risk of service interruption caused by excessive load;
[0026] Based on the upcoming load signal, the packet loss rate test is carried out on the node that responds abnormally. With the help of the network tester, the parameters are scientifically set, the packet loss rate is accurately calculated, and compared with the preset value, so as to reliably determine whether the node is a faulty node. This process provides a direct and critical basis for fault prediction, allowing operation and maintenance personnel to carry out hardware updates or parameter adjustments in a targeted manner to ensure the stable operation of the data center. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 It is a schematic diagram of the process of the present invention;
[0028] Figure 2 The figure is a flow chart of the packet loss rate test of the present invention. DETAILED DESCRIPTION
[0029] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0030] First embodiment
[0031] See also Figure 1 , this application provides a distributed data center operation and maintenance fault prediction method based on deep learning, comprising the following steps:
[0032] Step 1: Monitor the disk response time of several different distributed nodes in the distributed data center, and lock the abnormal response node based on the real-time monitored disk response time. Specifically, the disk response time associated with different distributed nodes is relatively the same, and there are specific standards set. When the corresponding disk response time is too long, it means that the corresponding node has an abnormal response, and it is necessary to perform relevant calibration of the abnormal response node;
[0033] Among them, the specific method of locking the response abnormal node is:
[0034] Monitor the disk response time of different distributed nodes and mark it as X i , where i represents different distributed nodes;
[0035] The disk response time associated with the corresponding distributed node is X i With the preset standard Y i Check, where Y i Different distributed nodes have different preset standards, which are all prepared in advance by relevant personnel based on experience: i ≤Y i , no processing is performed, and continuous monitoring is sufficient. If X i >Y i , then this distributed node is directly marked as an abnormal node (indicating that the response time does not meet the standard and there may be abnormal related data inside it).
[0036] Step 2: Monitor multiple parameters of the calibrated abnormal node, including read and write rate parameters, CPU usage parameters, and memory occupancy parameters. Based on the real-time monitoring process of multiple parameters, select abnormal monitoring items, and confirm the analysis period based on the selected abnormal monitoring items. The specific sub-steps for determining the analysis period are:
[0037] S21. Perform multiple parameter monitoring on the calibrated abnormal node, confirm the monitoring parameters of different monitoring items in real time (because different monitoring items include read and write rate, CPU usage rate and memory occupancy rate, three sets of monitoring change curves are generated in real time), and perform real-time verification on the monitoring parameters of different monitoring items: identify the standard interval associated with the corresponding monitoring item, and the standard interval is a preset interval, which is formulated by relevant operators based on experience. Different monitoring items correspond to different preset intervals, and the monitoring parameters confirmed by different monitoring items at the same time are calibrated as C k , where k represents different monitoring items. If C k ∈ standard interval, then continue monitoring, if C k If the monitoring item is not within the standard interval, the monitoring item is marked as an abnormal monitoring item, and the current moment is marked as an abnormal moment;
[0038] S22, using the determined abnormal time as the initial time, confirming a set of analysis time periods, where the duration of the analysis time period is a preset duration, which is determined by relevant personnel based on experience, and is generally 20 minutes;
[0039] S23, after the analysis period ends, the next set of analysis periods is confirmed in the same manner as steps S21-S22;
[0040] Specifically, during the determined analysis period, numerical analysis is performed on the relevant parameters of each different monitoring item, and based on the parameter trend changes of different monitoring items, it is evaluated whether the node will have a load condition, and the relevant display of the load signal is performed in a timely manner;
[0041] Step 3: Based on the determined analysis period, numerical analysis is performed on multiple parameters monitored in real time during the analysis period. According to the change trends of different monitoring parameters of different monitoring items, the prediction parameters associated with the corresponding monitoring items at the next moment are confirmed. Based on the different prediction parameters confirmed by different monitoring items, load prediction is performed, and the upcoming load signal is confirmed based on the prediction results. Specifically, in the actual analysis and processing process, if the parameters monitored by each monitoring item show a large change trend, then the corresponding overload situation will soon appear, so it is necessary to perform load prediction in advance, confirm the corresponding load signal in time and display it for timely processing by external relevant operators, and make relevant response measures in time;
[0042] The specific sub-steps of load forecasting are:
[0043] S31, confirm the different monitoring parameters associated with different monitoring items during the analysis period in real time, and mark the different monitoring parameters confirmed in real time as CS k-t and CS k-(t-1) , k represents different monitoring items, where t represents the current time, (t-1) represents the previous time, using: CZ k-t =CS k-t -CS k-(t-1) Confirm the change difference CZ of the monitoring parameter associated with the current moment k-t ;
[0044] S32, multiple groups of change differences CZ generated in real time for different monitoring items during the analysis period k-t Confirm and lock CZ k- t min and CZ k-t max confirms the difference interval (the difference interval is updated in real time as the analysis period continues), based on the currently confirmed monitoring parameter CS k-t And the determined difference interval, predict the monitoring parameter interval associated with the next moment, using: QS k min=CS k-t + Minimum difference interval and QS k max=CS k-t + Maximum value of the difference interval, based on the confirmed QS k min and QS k max generates the monitoring parameter interval associated with the corresponding monitoring item at the next moment;
[0045] S33, based on the different monitoring items predicted monitoring parameter interval at the next moment [QS k min, QS k max], using: FZ k min=QSk min×C k and FZ k max=QS k max×C k Confirm the load rate interval [FZ k min, FZ k max], where C k It is the preset factor associated with the corresponding monitoring item, and its preset factor is prepared in advance by the operator based on experience;
[0046] S34, summing up the different load rate intervals predicted by different monitoring items, confirming the total load rate interval [FZmin, FZmax], and checking the confirmed total load rate interval [FZmin, FZmax]:
[0047] Confirm whether there is a partial interval exceeding 90% in the total load rate interval. If so, confirm the specific proportion value ZB of the partial interval in the total load rate interval. If ZB<Y1, where Y1 is a preset value, and its specific value is determined by the operator based on experience, then continue to monitor and confirm the specific proportion value. If ZB≥Y1, generate a load signal, and display the generated load signal for external personnel to view;
[0048] If it does not exist, continue monitoring and confirm the partial interval segments.
[0049] By real-time confirmation and calculation of monitoring parameters, the difference of changes of each monitoring item at different times can be accurately obtained; for example, in S31, the calculation result is accurate through a clear formula; this accuracy helps to timely discover abnormal changes in monitoring parameters and provide a reliable basis for subsequent predictions;
[0050] As time goes by, the monitoring parameter interval will be dynamically adjusted according to the changes in real-time data and difference intervals. This dynamic adjustment makes the prediction more in line with the actual situation and improves the accuracy of the prediction;
[0051] By calculating the load rate interval, the load conditions of different monitoring items at different times can be effectively predicted; in S33, the load rate interval is obtained based on the preset factors and the monitoring parameter interval. This prediction method has high reliability and can provide a reference for actual decision-making;
[0052] The parameter changes of different monitoring items, monitoring parameter ranges, load rate ranges and other factors are comprehensively considered to make the prediction results more comprehensive; through the analysis and integration of each part, the operating status of the distributed data center can be grasped as a whole, providing strong support for operation and maintenance decisions.
[0053] Second embodiment
[0054] In the specific implementation process of this embodiment, compared with the above embodiment, this embodiment is mainly aimed at the fault prediction of the corresponding response abnormal node, and its specific prediction steps include:
[0055] Step 4: Combination Figure 2 Based on the confirmed upcoming load signal, a packet loss rate test is performed on the abnormal response node, and a fault prediction is performed on the abnormal response node based on the test result, wherein the specific sub-steps of the fault prediction are:
[0056] S41. Connect the port of the network tester to the interface of the abnormal response node to be tested to ensure a stable connection; if the node in the network link is to be tested, the tester may need to be connected in series to the link, and the number, size, and transmission rate of the data packets to be sent can be set in the network tester; generally, different data packet sizes can be selected, such as 64 bytes, 1518 bytes, etc., to simulate different network load conditions. The set parameters are all preset parameters, which are set in advance by the relevant operators. Start the test, and the network tester sends data packets to the abnormal response node according to the set parameters, and records the number of data packets sent SL1 at the same time;
[0057] S42, the network tester receives data packets returned from the response abnormal node or forwarded through the node, counts the number of received data packets SL2, uses: DB = (SL2-SL1) ÷ SL1 × 100% to confirm the packet loss rate DB of this response abnormal node, and checks the packet loss rate DB with the preset value Y2, wherein the specific value of Y2 is formulated by the operator based on experience. If DB≤Y2, it means that the packet loss rate of this response abnormal node is abnormal, and this response abnormal node is marked as a faulty node, and the marked faulty node is displayed, and the relevant personnel perform hardware updates or set relevant parameters to reduce the packet loss rate. If DB>Y2, it means that the packet loss rate of this response abnormal node is not abnormal, and no processing is performed, and the packet loss rate test of this response abnormal node is continuously performed;
[0058] Specifically, during the packet loss rate test, there is an associated test cycle, which is a preset cycle, and the starting time of the test cycle is the initial test time of the network tester, and the end time is the end time of the network tester test and then extended for a certain period of time. The certain period of time is a preset period of time, which is prepared in advance by relevant operators and is generally within 1 minute.
[0059] Some of the data in the above formulas are dimensionless and numerically calculated. Meanwhile, the contents not described in detail in this specification belong to the prior art known to those skilled in the art.
[0060] The above embodiments are only used to illustrate the technical method of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical method of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical method of the present invention.
Claims
1. A distributed data center operation and maintenance fault prediction method based on deep learning, characterized in that: The following steps are involved: Step 1: Monitor the disk response time of several different distributed nodes in the distributed data center, and lock the abnormal response node based on the real-time monitored disk response time; Step 2: Monitor multiple parameters of the calibrated abnormal node, including read / write rate parameters, CPU usage parameters, and memory occupancy parameters. Based on the real-time monitoring process of multiple parameters, select abnormal monitoring items, and confirm the analysis period based on the selected abnormal monitoring items; Step 3: Based on the determined analysis period, numerical analysis is performed on multiple parameters monitored in real time during the analysis period, and according to the change trends of different monitoring parameters of different monitoring items, the prediction parameters associated with the corresponding monitoring items at the next moment are confirmed, and load prediction is performed based on the different prediction parameters confirmed by different monitoring items, and the upcoming load signal is confirmed based on the prediction results; Step 4: Based on the confirmed upcoming load signal, a packet loss rate test is performed on the abnormal response node, and a fault prediction is performed on the abnormal response node based on the test result.
2. The distributed data center operation and maintenance fault prediction method based on deep learning according to claim 1 is characterized in that: In the step 1, the specific method of locking the node responding to abnormality is: Monitor the disk response time of different distributed nodes and mark it as X i , where i represents different distributed nodes; The disk response time associated with the corresponding distributed node is X i With the preset standard Y i Check, where Y i is the preset standard. Different distributed nodes have different preset standards: if X i ≤Y i , no processing is performed, and continuous monitoring is sufficient. If X i >Y i , then this distributed node is directly marked as an abnormal node.
3. The distributed data center operation and maintenance fault prediction method based on deep learning according to claim 1 is characterized in that: The specific sub-steps of step 2, determining the analysis period, are: S21, perform multiple parameter monitoring on the calibrated abnormal node, confirm the monitoring parameters of different monitoring items in real time, and perform real-time verification on the monitoring parameters of different monitoring items: identify the standard interval associated with the corresponding monitoring item, the standard interval is the preset interval, different monitoring items correspond to different preset intervals, and calibrate the monitoring parameters confirmed by different monitoring items at the same time as C k , where k represents different monitoring items. If the monitoring item is not within the standard interval, the monitoring item is marked as an abnormal monitoring item, and the current moment is marked as an abnormal moment; S22, using the determined abnormal time as the initial time, confirming a set of analysis time periods, wherein the duration of the analysis time periods is the preset duration; S23. After the analysis period ends, the next set of analysis periods is confirmed in the same manner as steps S21-S22.
4. The distributed data center operation and maintenance fault prediction method based on deep learning according to claim 3 is characterized in that: In step S21, if C k ∈ standard interval, then continue monitoring.
5. The distributed data center operation and maintenance fault prediction method based on deep learning according to claim 1 is characterized in that: In step 3, the specific sub-steps of performing load prediction are: S31, confirm the different monitoring parameters associated with different monitoring items during the analysis period in real time, and mark the different monitoring parameters confirmed in real time as CS k-t and CS k-(t-1) , k represents different monitoring items, where t represents the current time, (t-1) represents the previous time, using: CZ k-t =CS k-t -CS k-(t-1) Confirm the change difference CZ of the monitoring parameter associated with the current moment k-t ; S32, multiple groups of change differences CZ generated in real time for different monitoring items during the analysis period k-t Confirm and lock CZ k-t min and CZ k-t max confirmed difference interval, based on the currently confirmed monitoring parameter CS k-t And the determined difference interval, predict the monitoring parameter interval associated with the next moment, using: QS k min=CS k-t + Minimum difference interval and QS k max=CS k-t + Maximum value of the difference interval, based on the confirmed QS k min and QS k max generates the monitoring parameter interval associated with the corresponding monitoring item at the next moment; S33, based on the different monitoring items predicted monitoring parameter interval at the next moment [QS k min, QS k max], using: FZ k min=QS k min×C k and FZ k max=QS k max×C k Confirm the load rate interval [FZ k min, FZ k max], where C k The preset factor associated with the corresponding monitoring item; S34, summing up the different load rate intervals predicted by different monitoring items, confirming the total load rate interval [FZmin, FZmax], and checking the confirmed total load rate interval [FZmin, FZmax]: Confirm whether there is a partial interval exceeding 90% in the total load rate interval. If so, confirm the specific proportion value ZB of the partial interval in the total load rate interval. If ZB<Y1, where Y1 is a preset value, continue to monitor and confirm the specific proportion value. If ZB≥Y1, generate an impending load signal and display the generated impending load signal.
6. The distributed data center operation and maintenance fault prediction method based on deep learning according to claim 5 is characterized in that: In the step S34, if there is no partial interval exceeding 90% in the total load rate interval, the partial interval is continuously monitored and confirmed.
7. The distributed data center operation and maintenance fault prediction method based on deep learning according to claim 1 is characterized in that: In step 4, the specific sub-steps of performing fault prediction are: S41, connecting the port of the network tester to the interface of the abnormal response node to be tested, and setting the number, size, and transmission rate of the sent data packets in the network tester; Start the test, the network tester sends data packets to the response abnormal node according to the set parameters, and records the number of data packets sent SL1 at the same time; S42. The network tester receives data packets returned from the response abnormal node or forwarded through the node, counts the number of received data packets SL2, and uses: DB = (SL2-SL1) ÷ SL1 × 100% to confirm the packet loss rate DB of this response abnormal node, and compares the packet loss rate DB with the preset value Y2. If DB≤Y2, it means that the packet loss rate of this response abnormal node is abnormal, and this response abnormal node is marked as a faulty node, and the marked faulty node is displayed.
8. The distributed data center operation and maintenance fault prediction method based on deep learning according to claim 1 is characterized in that: In the step S42, if DB>Y2, it means that the packet loss rate of the abnormal response node is not abnormal, and no processing is performed, and the packet loss rate test of the abnormal response node is continuously performed.
Citation Information
Patent Citations
Disk fault prediction method for intelligent operation and maintenance of large-scale cloud data center
CN115373879A
Server load prediction method and device, electronic equipment and storage medium
CN113535530A
Cloud computing physical node load monitoring method and device, terminal and storage medium
CN113626282A
Fault detection method and device and storage medium
CN116668274A
Comprehensive monitoring SaaS service platform based on cloud computing architecture
CN118093244A
Cited By
Solid state disk fault intelligent prediction system
CN120687276A