Server exception detection and early warning method and device, electronic equipment and readable medium

By creating a load metric relationship model and a health status joint prediction model, the data collection frequency and threshold of the server are dynamically adjusted, which solves the problems of false alarms and missed alarms in server anomaly detection, realizes early identification and warning of potential risks, and reduces resource waste.

CN121542091BActive Publication Date: 2026-04-17GUANGZHOU CLOUDSINO INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
GUANGZHOU CLOUDSINO INFORMATION TECH CO LTD
Filing Date
2026-01-15
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

In existing technologies, server anomaly detection and early warning methods rely on fixed temperature and power consumption thresholds, which cannot adapt to different workloads and environmental conditions. This leads to frequent false alarms or missed alarms, and the inability to identify potential risks in advance, resulting in wasted early warning resources and unidentified potential risks.

Method used

By acquiring historical monitoring indicator data sequences of the server, a load indicator relationship model is created to perform adaptive dynamic monitoring, dynamically adjust the data collection frequency and threshold, and combine it with a health status joint prediction model to achieve real-time anomaly detection and early warning.

Benefits of technology

It reduces the waste of early warning resources, improves the ability to identify potential risks to servers, reduces false alarms and missed alarms, achieves early warning, and ensures the stable operation of servers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121542091B_ABST
    Figure CN121542091B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure disclose a server anomaly detection and early warning method, device, electronic equipment and readable medium. A specific embodiment of the method comprises: obtaining a historical monitoring index data sequence corresponding to a preset to-be-monitored server; creating a load index relationship model; performing adaptive dynamic monitoring processing on the preset to-be-monitored server within a preset monitoring time period, collecting monitoring index data according to a dynamic collection frequency, and performing real-time anomaly detection and early warning processing on the collected monitoring index data; sorting each monitoring index data collected in the preset monitoring time period, training a preset health state joint prediction model, and obtaining a health state joint prediction model; inputting a to-be-predicted monitoring index data sequence into the health state joint prediction model to obtain a predicted monitoring index data sequence; and performing future potential risk early warning processing on the preset to-be-monitored server. The embodiment reduces the waste of early warning resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of this disclosure relate to the field of computer technology, and more specifically to server anomaly detection and early warning methods, apparatuses, electronic devices, and readable media. Background Technology

[0002] The stable operation of servers experiencing sudden high loads (e.g., servers handling real-time interactive services during e-commerce promotions, interactive servers for live streaming and short video platforms, and financial securities trading servers) is crucial for ensuring the smooth operation of critical businesses such as big data processing, intelligent manufacturing, and financial services. Server anomaly detection and early warning is a technology for detecting and issuing warnings about server anomalies. Currently, the common method for server anomaly detection and early warning is to collect server temperature and power consumption data in real time and then perform anomaly detection and warning based on fixed temperature / power consumption thresholds.

[0003] However, when using the above methods for server anomaly detection and early warning, the following technical problems often arise:

[0004] Fixed temperature and power consumption thresholds may not be suitable for different workloads and environmental conditions. As the load changes, the normal operating temperature and power consumption of servers may fluctuate due to instantaneous high loads. Using static thresholds may lead to frequent false alarms or missed alarms, resulting in a waste of early warning resources. Moreover, real-time collection of server temperature and power consumption, and anomaly detection and early warning based on fixed temperature / power consumption thresholds, cannot identify potential risks in advance and provide early warnings.

[0005] The information disclosed in this background section is only intended to enhance the understanding of the background of the inventive concept, and therefore may contain information that does not form prior art known to those skilled in the art. Summary of the Invention

[0006] The summary portion of this disclosure is intended to provide a brief overview of the concepts, which will be described in detail in the detailed description portion. This summary portion is not intended to identify key or essential features of the claimed technical solutions, nor is it intended to limit the scope of the claimed technical solutions.

[0007] Some embodiments of this disclosure provide server anomaly detection and early warning methods, apparatuses, electronic devices, and computer-readable media to address one or more of the technical problems mentioned in the background section above.

[0008] In a first aspect, some embodiments of this disclosure provide a method for server anomaly detection and early warning. The method includes: acquiring a historical monitoring indicator data sequence corresponding to a preset server to be monitored; creating a load indicator relationship model based on the historical monitoring indicator data sequence; performing adaptive dynamic monitoring processing on the preset server to be monitored within a preset monitoring time period to collect monitoring indicator data according to a dynamic collection frequency; and performing real-time anomaly detection and early warning processing on the collected monitoring indicator data based on the load indicator relationship model; sorting the monitoring indicator data collected within the preset monitoring time period to obtain a monitoring indicator data sequence; training a preset joint health status prediction model based on the monitoring indicator data sequence to obtain a joint health status prediction model; truncating the monitoring indicator data sequence to obtain a monitoring indicator data sequence to be predicted; inputting the monitoring indicator data sequence to be predicted into the joint health status prediction model to obtain a predicted monitoring indicator data sequence; and performing future potential risk early warning processing on the preset server to be monitored based on the predicted monitoring indicator data sequence and the load indicator relationship model.

[0009] Secondly, some embodiments of this disclosure provide a server anomaly detection and early warning device, comprising: an acquisition unit configured to acquire a historical monitoring indicator data sequence corresponding to a preset server to be monitored; a creation unit configured to create a load indicator relationship model based on the aforementioned historical monitoring indicator data sequence; an adaptive dynamic monitoring unit configured to perform adaptive dynamic monitoring processing on the preset server to be monitored within a preset monitoring time period, to collect monitoring indicator data according to a dynamic acquisition frequency, and to perform real-time anomaly detection and early warning processing on the collected monitoring indicator data based on the aforementioned load indicator relationship model; a training unit configured to train a preset health status joint prediction model based on the aforementioned monitoring indicator data sequence to obtain a health status joint prediction model; an interception unit configured to intercept the aforementioned monitoring indicator data sequence to obtain a monitoring indicator data sequence to be predicted; an input unit configured to input the monitoring indicator data sequence to be predicted into the aforementioned health status joint prediction model to obtain a predicted monitoring indicator data sequence; and a future potential risk early warning unit configured to perform future potential risk early warning processing on the preset server to be monitored based on the predicted monitoring indicator data sequence and the aforementioned load indicator relationship model.

[0010] Thirdly, some embodiments of this disclosure provide an electronic device, including: one or more processors; and a storage device having one or more programs stored thereon, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any implementation of the first aspect above.

[0011] Fourthly, some embodiments of this disclosure provide a computer-readable medium having a computer program stored thereon, wherein the program, when executed by a processor, implements the method described in any of the implementations of the first aspect above.

[0012] The above embodiments of this disclosure have the following beneficial effects: the server anomaly detection and early warning methods of some embodiments of this disclosure reduce the waste of early warning resources and achieve early identification of potential risks and early warning. Specifically, the waste of early warning resources is due to the fact that fixed temperature and power consumption thresholds may not be suitable for different workloads and environmental conditions. As the load changes, the normal operating temperature and power consumption of servers with instantaneous high loads will also fluctuate. Using static thresholds may lead to frequent false alarms or missed alarms, resulting in a waste of early warning resources. Moreover, real-time collection of server temperature and power consumption, and anomaly detection and early warning based on fixed temperature / power consumption thresholds, cannot identify potential risks and provide early warning. Based on this, the server anomaly detection and early warning methods of some embodiments of this disclosure first obtain a historical monitoring indicator data sequence corresponding to a preset server to be monitored. Thus, a load indicator relationship model for creating a load indicator relationship model can be obtained. Then, based on the above historical monitoring indicator data sequence, a load indicator relationship model is created. Thus, a load indicator relationship model including the correspondence between indicator ranges (e.g., temperature range, power consumption range) under different loads can be created. Next, within the preset monitoring period, adaptive dynamic monitoring is performed on the preset servers to be monitored. Monitoring indicator data is collected at a dynamic acquisition frequency, and based on the aforementioned load indicator relationship model, real-time anomaly detection and early warning processing are performed on the collected monitoring indicator data. The dynamic acquisition frequency adjusts the data collection frequency in real time according to changes in server load. The acquisition frequency is increased when load changes drastically to capture indicator changes more promptly, while the acquisition frequency is reduced when the load is stable to minimize unnecessary resource consumption. Simultaneously, real-time anomaly detection using the load indicator relationship model dynamically adjusts the monitoring thresholds within the normal range based on real-time load conditions. When server load changes, the model dynamically adjusts the judgment criteria for the normal range of indicators based on the current load, thus adapting to different workloads and environmental conditions. This avoids potential misjudgments that might occur with fixed thresholds under different loads, thereby reducing false alarms or missed alarms and minimizing the waste of early warning resources. Next, the monitoring indicator data collected within the preset monitoring period is sorted to obtain a monitoring indicator data sequence. This yields the monitoring indicator data sequence for the preset monitoring period. Then, based on the aforementioned monitoring indicator data sequence, a preset joint health status prediction model is trained to obtain the joint health status prediction model. Therefore, the pre-defined health status joint prediction model can be trained to learn the changes in server indicators at different time points. Next, the monitoring indicator data sequence is truncated to obtain the monitoring indicator data sequence to be predicted. This yields the monitoring indicator data sequence reflecting the current state and recent trends of the server. Then, the monitoring indicator data sequence to be predicted is input into the aforementioned health status joint prediction model to obtain the predicted monitoring indicator data sequence.Therefore, a pre-trained joint health status prediction model can be used to predict the data sequence of monitoring indicators for the server over a future period, based on the data sequence of the monitoring indicators to be predicted. Finally, based on the predicted monitoring indicator data sequence and the aforementioned load indicator relationship model, potential risk warnings are issued for the pre-defined servers to be monitored. This achieves the early identification and warning of potential risks in server operation. Attached Figure Description

[0013] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and elements are not necessarily drawn to scale.

[0014] Figure 1 This is a flowchart of some embodiments of the server anomaly detection and early warning method according to this disclosure;

[0015] Figure 2 This is a schematic diagram of the structure of some embodiments of the server anomaly detection and early warning device according to the present disclosure;

[0016] Figure 3 This is a schematic diagram of the structure of an electronic device suitable for implementing some embodiments of the present disclosure. Detailed Implementation

[0017] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0018] It should also be noted that, for ease of description, only the parts relevant to the invention are shown in the accompanying drawings. Unless otherwise specified, the embodiments and features described in this disclosure can be combined with each other.

[0019] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.

[0020] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0021] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.

[0022] This disclosure will now be described in detail with reference to the accompanying drawings and embodiments.

[0023] Figure 1 A flow 100 of some embodiments of the server anomaly detection and early warning method according to the present disclosure is shown. The server anomaly detection and early warning method includes the following steps:

[0024] Step 101: Obtain the historical monitoring indicator data sequence corresponding to the preset server to be monitored.

[0025] In some embodiments, the execution entity (e.g., a computing device) of the server anomaly detection and early warning method can obtain a historical monitoring indicator data sequence corresponding to a preset server to be monitored. The preset server to be monitored can be a server experiencing instantaneous high load (e.g., a server responsible for real-time interactive services during e-commerce promotions, a live streaming and short video platform interactive server, a financial securities trading server, etc.). The execution entity can obtain the historical monitoring indicator data sequence corresponding to the preset server to be monitored from a preset database. The historical monitoring indicator data sequence can be data reflecting the temperature, power consumption, load, and status of the preset server to be monitored over a past period. Each historical monitoring indicator data in the historical monitoring indicator data sequence represents the historical monitoring indicator data of the preset server to be monitored at a past instant, including historical temperature data, historical power consumption data, historical load data, and historical status data. The historical temperature data can represent the instantaneous temperature of the preset server to be monitored at the aforementioned past instant, and the instantaneous temperature can be, but is not limited to, one of the following: CPU temperature, motherboard temperature, memory temperature, hard disk temperature, and chassis inlet or outlet temperature. The historical power consumption data mentioned above can represent the instantaneous power consumption of the server under monitoring, which can be, but is not limited to, one of the following: total system power consumption (e.g., 200W) or component (e.g., CPU, GPU) power consumption (e.g., CPU power consumption: 30W). The historical load data mentioned above can represent the instantaneous load value (e.g., CPU utilization) of the server under monitoring at a specific point in the past. The historical status data mentioned above can reflect the physical status of the server and historical event records. The status data includes fan speed and status logs. The status logs can be logs recording historical event data, which can be log information related to temperature and power consumption.

[0026] Step 102: Create a load indicator relationship model based on historical monitoring indicator data sequences.

[0027] In some embodiments, the aforementioned execution entity may create a load indicator relationship model based on the aforementioned historical monitoring indicator data sequence. This load indicator relationship model may be a relationship model that includes the correspondence between indicator ranges (e.g., temperature range, power consumption range) under different loads.

[0028] In some optional implementations of certain embodiments, the aforementioned execution entity may create a load indicator relationship model based on the aforementioned historical monitoring indicator data sequence through the following steps:

[0029] The first step is to sort the historical load data within the historical monitoring indicator data sequence to obtain the historical load data sequence. Each of these historical monitoring indicator data corresponds to a collection time. In practice, the executing entity can sort the historical load data from front to back according to the collection time corresponding to each historical load data point to obtain the historical load data sequence.

[0030] The second step involves quantile processing of the aforementioned historical load data sequence to obtain a quantile historical load range information sequence and a historical load data group sequence. Each quantile historical load range information in the quantile historical load range information sequence corresponds to a historical load data group in the historical load data group sequence. The aforementioned historical load data group can be at least one historical load data point in the historical load data sequence that falls within the range represented by the quantile historical load range information. In practice, the executing entity can use a quartile grouping algorithm to quantify the aforementioned historical load data sequence to obtain the quantile historical load range information sequence and the historical load data group sequence. As an example, the aforementioned historical load data sequence can be {10, 20, 15, 30, 25, 40, 35, 50, 45, 60}. The historical load data sequence is then sorted in ascending order: {10, 15, 20, 25, 30, 35, 40, 45, 50, 60}. Calculate the quartiles of the sorted sequence: Q1 (25th quartile): Position = 0.25 × (10 - 1) + 1 = 3.25, taking the interpolation of the 3rd and 4th numbers, i.e., Q1 = 20 + 0.25 × (25 - 20) = 21.25. Q2 (50th quartile, median): Position = 0.5 × (10 - 1) + 1 = 5.5, taking the average of the 5th and 6th numbers, i.e., Q2 = (30 + 35) / 2 = 32.5. Q3 (75th quartile): Position = 0.75 × (10 - 1) + 1 = 7.75, taking the interpolation of the 7th and 8th numbers, i.e., Q3 = 45 + 0.75 × (50 - 45) = 48.75. Values ​​less than or equal to Q1, i.e., less than or equal to 21.25, are identified as the first historical load range information in the historical load range information sequence. Values ​​greater than Q1 and less than or equal to Q2 (i.e., greater than 21.25 and less than or equal to 32.5) are identified as the second historical load range in the historical load range information sequence. Values ​​greater than Q2 and less than or equal to Q3 (i.e., greater than 32.5 and less than or equal to 48.75) are identified as the third historical load range in the historical load range information sequence. Values ​​greater than or equal to Q3 (i.e., less than or equal to 48.75) are identified as the last historical load range in the historical load range information sequence. The historical load data group sequence can be {[10, 15, 20], [25, 30], [35, 40, 45], [50, 60]}.

[0031] The third step is to create a load indicator relationship model based on the historical load data set sequence and the aforementioned historical monitoring indicator data sequence.

[0032] In some optional implementations of certain embodiments, the aforementioned execution entity may create a load metric relationship model based on the historical load data group sequence and the aforementioned historical monitoring metric data sequence through the following steps:

[0033] The first step is to perform the following steps for each historical load data group in the above historical load data group sequence:

[0034] The first sub-step involves identifying the historical monitoring indicator data containing the historical load data within each of the aforementioned historical load data groups as the target historical monitoring indicator data for each historical load data set.

[0035] The second sub-step involves identifying the historical temperature data, which includes the historical monitoring indicators for each target, as a historical temperature data group.

[0036] The third sub-step involves defining the historical power consumption data, which includes the historical monitoring indicators of each target, as a historical power consumption data group.

[0037] The fourth sub-step is to determine the quantile historical load range information corresponding to the historical load data group in the above quantile historical load range information sequence as the target quantile historical load range information.

[0038] The fifth sub-step involves defining the historical temperature data set within a normal range to obtain historical normal temperature range information. In practice, the executing entity can use a quartile grouping algorithm to group the historical temperature data set into quartiles. These quartiles can be represented by QT1 (the 25th quartile of the historical temperature data set), QT2 (the 50th quartile of the historical temperature data set), and QT3 (the 75th quartile of the historical temperature data set). The executing entity can then define the range represented by QT1 and less than or equal to QT3 as the historical normal temperature range information.

[0039] The sixth sub-step involves defining the historical power consumption data set within a normal range to obtain historical normal power consumption range information. In practice, the executing entity can use a quartile grouping algorithm to group the historical power consumption data set into quartiles. These quartiles can be represented by QE1 (the 25th quartile of the historical power consumption data set), QE2 (the 50th quartile of the historical power consumption data set), and QE3 (the 75th quartile of the historical power consumption data set). The executing entity can then define the range represented by QE1 and less than or equal to QE3 as the historical normal power consumption range information.

[0040] The seventh sub-step involves determining the aforementioned historical normal temperature range information and the aforementioned historical normal power consumption range information as a load index relationship sub-model corresponding to the aforementioned target quantile historical load range information.

[0041] The second step is to define the identified sub-models of the load index relationships as the load index relationship model.

[0042] Step 103: Within the preset monitoring time period, perform adaptive dynamic monitoring on the preset servers to be monitored, so as to collect monitoring indicator data according to the dynamic collection frequency, and perform real-time anomaly detection and early warning processing on the collected monitoring indicator data based on the load indicator relationship model.

[0043] In some embodiments, the aforementioned execution entity may perform adaptive dynamic monitoring on a preset server to be monitored within a preset monitoring time period, so as to collect monitoring indicator data according to the dynamic collection frequency, and perform real-time anomaly detection and early warning processing on the collected monitoring indicator data based on the aforementioned load indicator relationship model.

[0044] In addressing the technical problems mentioned above by adopting technical solutions, the application scenario of detecting and issuing early warnings for servers experiencing sudden high loads often presents the following challenges: During periods of high server load or busy business operations, hardware conditions (such as temperature and power consumption) change more rapidly and drastically. If monitoring data is collected at a fixed, low frequency (e.g., every 5 minutes), abnormal conditions (such as rapid temperature increases) may not be detected and addressed in a timely manner, leading to delayed or missed warnings. This increases the likelihood of server damage and service interruptions due to these delays and missed warnings. Therefore, this application scenario requires the following characteristics: It should be suitable for detecting and issuing early warnings for servers experiencing sudden high loads. Faced with these technical problems, we have decided to adopt the following solution:

[0045] In some optional implementations of certain embodiments, the aforementioned execution entity may perform adaptive dynamic monitoring of a preset server to be monitored within a preset monitoring time period through the following steps: collecting monitoring indicator data according to a dynamic collection frequency; and performing real-time anomaly detection and early warning processing on the collected monitoring indicator data based on the aforementioned load indicator relationship model.

[0046] The first step is to obtain the system time as the current time.

[0047] The second step, based on the obtained current time, is to perform the following real-time anomaly detection and early warning processing:

[0048] The first sub-step involves collecting the operational status information of a preset server to be monitored. This operational status information can be the server's CPU utilization or memory utilization. In practice, the executing entity can directly connect to the preset server to be monitored via SSH or other remote command execution protocols and execute commands to obtain the server's operational status information.

[0049] The second sub-step involves determining the sampling frequency as the dynamic sampling frequency based on the operational status information. In practice, when the operational status information meets preset conditions, the aforementioned execution entity can increase the preset frequency by a preset factor, and determine the frequency after increasing the preset factor as the sampling frequency as the dynamic sampling frequency. The preset frequency can be a pre-set sampling frequency; for example, if the preset frequency is once every 5 minutes and the preset factor is 5, then the frequency after increasing the preset factor can be once every 1 minute.

[0050] The third sub-step involves dynamically collecting monitoring indicator data in real time within a preset time period following the current time. This monitoring indicator data includes temperature data, power consumption data, load data, and status data.

[0051] The fourth sub-step involves performing anomaly detection and early warning processing on the collected monitoring indicator data, based on the aforementioned load indicator relationship model.

[0052] The third step is to obtain the system time again as the current time after a preset time period following the current time, and in response to determining that the current time is within the preset monitoring time period, to perform the above-mentioned real-time anomaly detection and early warning processing again based on the obtained current time.

[0053] The above-described technical solution and its related content, as an inventive point of this disclosure, solve the technical problem of "increasing the possibility of server damage and service interruption due to early warning delays and missed reports." Factors that increase the possibility of server damage and service interruption due to early warning delays and missed reports often include: during periods of high server load or busy business, the hardware status (such as temperature and power consumption) changes more rapidly and drastically. If a fixed, low frequency (such as every 5 minutes) is used to collect monitoring indicator data, abnormal states (such as rapid temperature rise) may not be detected and responded to in a timely manner, leading to early warning delays and missed reports, thus increasing the possibility of server damage and service interruption due to early warning delays and missed reports. Solving these factors can reduce the possibility of server damage and service interruption due to early warning delays and missed reports. To achieve this effect, firstly, the system time is obtained as the current time. Then, based on the obtained current time, the following real-time anomaly detection and early warning processing is performed: collecting the operating status information of a preset server to be monitored. Thus, the operating status information of the preset server to be monitored can be collected to determine whether the server is under high load or busy business. Then, based on the operating status information, the collection frequency is determined as the dynamic collection frequency. Therefore, the data collection frequency can be automatically adjusted based on the preset operating status of the server to be monitored. During periods of high load, the collection frequency can be increased to ensure timely detection of abnormal changes. Then, monitoring indicator data is collected in real-time within a preset time period following the current time at a dynamic collection frequency. This allows for continuous and real-time collection of monitoring indicator data at the adjusted dynamic frequency. Next, based on the aforementioned load indicator relationship model, anomaly detection and early warning processing are performed on the collected monitoring indicator data. Thus, anomaly detection and early warning processing can be performed based on the real-time collected monitoring indicator data using the load indicator relationship model to identify potential anomalies and issue warnings. After a preset time period following the current time, the system time is obtained again as the current time, and in response to determining that the current time is within the preset monitoring time period, the aforementioned real-time anomaly detection and early warning processing is executed again based on the obtained current time. Because it adopts the method of adjusting the collection frequency in real time according to the operating status information of the preset monitored server, the collection frequency can be automatically increased during periods of high load or busy business, thereby obtaining key monitoring indicator data more frequently. The high-frequency data collection makes it easier to detect abnormal situations (such as rapid temperature rise) in a timely manner, reducing the possibility of damage to the preset monitored server and service interruption due to delays or missed reports.

[0054] In addressing the technical challenges of the aforementioned background technologies, the application scenario—anomaly detection and early warning for financial securities trading servers with extremely high business continuity requirements—often presents the following technical problems: traditional detection relies solely on single indicator values, failing to incorporate correlated data such as fan speed and status logs. This results in a lack of specificity in early warning information, necessitating manual repair responses after anomalies occur, leading to delays and increasing the risk of server hardware damage or service interruptions. This application scenario requires the following characteristics: anomaly detection and early warning for servers with extremely high business continuity requirements; rapid response after anomalies occur; and automatic fault mitigation. Faced with these technical challenges, we have decided to adopt the following solution:

[0055] In some optional implementations of certain embodiments, the aforementioned execution entity can perform anomaly detection and early warning processing on the collected monitoring indicator data based on the aforementioned load indicator relationship model through the following steps: First, determine the load data included in the monitoring indicator data as the load data to be matched. Each monitoring indicator data in the aforementioned monitoring indicator data sequence includes temperature data, power consumption data, load data, and status data. The status data includes fan speed and status logs. Each load indicator relationship sub-model in the aforementioned load indicator relationship model corresponds to a quantile historical load range information in the aforementioned quantile historical load range information sequence. The aforementioned load indicator relationship sub-model includes historical normal temperature range information and historical normal power consumption range information.

[0056] The second step is to determine the quantile historical load range information corresponding to the quantile historical load range where the load data to be matched is located in the quantile historical load range information sequence as the matching load range information.

[0057] The third step is to determine the load index relationship sub-model that corresponds to the above-mentioned matching load range information in each load index relationship sub-model as the target load index relationship sub-model.

[0058] The fourth step involves determining that the temperature data included in the aforementioned monitoring indicator data is not within the range represented by the historical normal temperature range information included in the target load indicator relationship sub-model. This involves retrieving temperature-related log information from the status logs included in the status data of the monitoring indicator data. In practice, the aforementioned executing entity can analyze the status logs using log analysis techniques to obtain temperature-related log information. For example, log analysis techniques can be used to process the logs and filter out log information containing a first preset keyword, such as "temperature" or "Temperature," as temperature-related log information.

[0059] The fifth step involves generating an early warning message based on the fan speed data included in the temperature-related log information and monitoring metrics. In practice, the filtered temperature-related log information can be correlated with the fan speed data, and an early warning message can be generated according to preset early warning rules. These preset early warning rules can be a series of pre-defined judgment conditions and corresponding processing criteria for monitoring and analyzing the operating status of preset servers. For example, the preset early warning rule could be: "When log information indicating an abnormally high temperature is detected, check the fan speed within the corresponding time period. When the fan speed is lower than 80% of the preset normal speed, an early warning message containing the name of the preset server to be monitored (e.g., server name), preset alarm content (e.g., excessively high temperature and excessively low fan speed), preset alarm reason (e.g., poor heat dissipation, fan failure, etc.), and preset mitigation task information (e.g., an instruction to start a backup fan) will be identified as the early warning message."

[0060] The sixth step involves determining that the power consumption data included in the aforementioned monitoring indicator data is not within the range represented by the historical normal power consumption range information included in the target load indicator relationship sub-model. This involves retrieving load-related log information from the status logs included in the status data of the monitoring indicator data. In practice, the aforementioned executing entity can process the status logs using log analysis technology to filter out log information containing a second preset keyword such as "CPU utilization rate" or "CPU-user%" as temperature-related log information. The aforementioned load-related log information may include CPU utilization rate.

[0061] Step 7: Based on the aforementioned load-related log information, generate early warning information. In practice, in response to determining that the CPU utilization rate included in the load-related log information is greater than a preset CPU utilization threshold, the aforementioned execution entity can determine the early warning information as including the preset name of the server to be monitored (e.g., server name), preset alarm content (e.g., overload), and preset mitigation task information (e.g., instructions indicating the triggering of load migration).

[0062] Step 8: Send the aforementioned warning information to the maintenance terminal. The maintenance terminal can be a terminal device (such as a mobile phone, computer, etc.) used by the maintenance personnel of the server to be monitored.

[0063] Step nine: Based on the aforementioned warning information, execute automatic mitigation tasks, which include at least one of the following: starting the backup fan in the server, triggering load migration, and adjusting the fan speed. In practice, the executing entity can determine the preset mitigation task information included in the warning information as the mitigation task information to be executed. Then, the executing entity can send the mitigation task information to the preset monitored server, so that the preset monitored server can execute the automatic mitigation task corresponding to the mitigation task information to be executed.

[0064] The above-described technical solution and its related content, as an inventive point of this disclosure, solve the technical problem of "easily causing server hardware damage or service interruption". Factors leading to server hardware damage or service interruption often include: traditional detection relies solely on a single indicator value, without combining it with related data such as fan speed and status logs, resulting in a lack of targeted early warning information. This leads to reliance on manual repair responses after anomalies occur, causing delays in handling and easily causing server hardware damage or service interruption. Solving these factors can mitigate faults and reduce server hardware damage or service interruption. To achieve this effect, firstly, the load data included in the monitoring indicator data is identified as the load data to be matched. Then, in the quantile historical load range information sequence, the quantile historical load range information corresponding to the quantile historical load range where the aforementioned load data to be matched is located is identified as the matching load range information. Next, the load indicator relationship sub-models corresponding to the aforementioned matching load range information in each load indicator relationship sub-model are identified as the target load indicator relationship sub-models. Thus, a target load indicator relationship sub-model representing the normal range of temperature and power consumption under the current load level can be obtained. Next, in response to determining that the temperature data included in the aforementioned monitoring indicator data is not within the range represented by the historical normal temperature range information included in the target load indicator relationship sub-model, log information corresponding to the temperature is obtained from the status log included in the status data of the monitoring indicator data as temperature-related log information. Thus, log information corresponding to the temperature can be obtained as temperature-related log information. Next, based on the aforementioned temperature-related log information and the fan speed included in the status data of the monitoring indicator data, an early warning message is generated. Thus, by combining the temperature-related log information and fan speed, an early warning message for executing automatic mitigation tasks can be generated in the event of abnormal temperature. Next, in response to determining that the power consumption data included in the aforementioned monitoring indicator data is not within the range represented by the historical normal power consumption range information included in the target load indicator relationship sub-model, log information corresponding to the load is obtained from the status log included in the status data of the monitoring indicator data as load-related log information. Then, based on the aforementioned load-related log information, an early warning message is generated. Thus, in the event of abnormal power consumption, an early warning message for executing automatic mitigation tasks can be generated. Finally, the aforementioned early warning message is sent to the maintenance terminal. Finally, based on the aforementioned warning information, an automatic mitigation task is executed. This automatic mitigation task includes at least one of the following: activating the backup fan in the server, triggering load migration, and adjusting the fan speed. Therefore, after a warning is issued to the maintenance terminal, at least one of these actions—activating the backup fan in the server, triggering load migration, and adjusting the fan speed—can be performed based on the generated warning information to mitigate the fault. This reduces the risk of server hardware damage or service interruption caused by delays in handling responses that rely solely on manual maintenance.

[0065] Step 104 involves sorting the monitoring indicator data collected during the preset monitoring time period to obtain a monitoring indicator data sequence.

[0066] In some embodiments, the executing entity can sort the monitoring indicator data collected within a preset monitoring time period to obtain a monitoring indicator data sequence. In practice, the executing entity can arrange the monitoring indicator data from front to back according to their collection time to obtain a monitoring indicator data sequence. The collection time interval between every two monitoring indicator data in the above monitoring indicator data sequence is a preset time interval (e.g., 5 minutes).

[0067] Step 105: Based on the monitoring indicator data sequence, train the preset joint prediction model of health status to obtain the joint prediction model of health status.

[0068] In some embodiments, the aforementioned executing entity may train a preset joint prediction model of health status based on the aforementioned monitoring indicator data sequence to obtain a joint prediction model of health status.

[0069] In some optional implementations of certain embodiments, the aforementioned execution entity may train a preset joint prediction model of health status based on the aforementioned monitoring indicator data sequence through the following steps to obtain the joint prediction model of health status:

[0070] The first step is to group the above monitoring indicator data sequence to obtain a monitoring indicator data group sequence. In practice, the executing entity can group the monitoring indicator data sequence according to a preset length to obtain the monitoring indicator data group sequence. As an example, the monitoring indicator data sequence can be {monitoring indicator data 1, monitoring indicator data 2, monitoring indicator data 3, monitoring indicator data 4, monitoring indicator data 5, monitoring indicator data 6}. If the preset length is 2, then the monitoring indicator data group sequence can be {[monitoring indicator data 1, monitoring indicator data 2], [monitoring indicator data 3, monitoring indicator data 4], [monitoring indicator data 5, monitoring indicator data 6]}.

[0071] The second step involves performing the following steps for every two consecutive monitoring indicator data groups in the above monitoring indicator data group sequence:

[0072] The first sub-step involves identifying the monitoring indicator data group that precedes the other monitoring indicator data group from two consecutive monitoring indicator data groups as the sample monitoring indicator data group, and identifying the other monitoring indicator data group as the sample target monitoring indicator data group corresponding to the sample monitoring indicator data group. For example, [Monitoring indicator data 1, Monitoring indicator data 2] is the sample monitoring indicator data group, and [Monitoring indicator data 3, Monitoring indicator data 4] is the sample target monitoring indicator data group corresponding to [Monitoring indicator data 1, Monitoring indicator data 2]. Similarly, [Monitoring indicator data 3, Monitoring indicator data 4] is the sample monitoring indicator data group, and [Monitoring indicator data 5, Monitoring indicator data 6] is the sample target monitoring indicator data group corresponding to [Monitoring indicator data 3, Monitoring indicator data 4].

[0073] The second sub-step involves determining the sample monitoring indicator data set and the corresponding sample target monitoring indicator data set as training samples. For example, the training samples can be the sample monitoring indicator data set: [monitoring indicator data 1, monitoring indicator data 2] and the corresponding sample target monitoring indicator data set: [monitoring indicator data 3, monitoring indicator data 4].

[0074] The third step is to define each of the identified training samples as the training sample set.

[0075] The fourth step is to train the preset joint prediction model of health status based on the above training sample set to obtain the joint prediction model of health status.

[0076] In some optional implementations of certain embodiments, the aforementioned execution entity may train a preset joint prediction model of health status based on the aforementioned training sample set through the following steps to obtain a joint prediction model of health status:

[0077] The first step is to perform the following training steps based on the training sample set:

[0078] Sub-step one involves inputting the sample monitoring index data set of at least one training sample from the training sample set into the initial neural network to obtain the prediction monitoring index data set corresponding to each of the at least one training sample. The initial neural network can be a Long Short-Term Memory (LSTM) network.

[0079] Sub-step two involves comparing the predicted monitoring indicator data set corresponding to each training sample in the at least one training sample with the corresponding target monitoring indicator data set. In practice, the execution entity can use the cross-entropy loss function to compare and determine the difference between the predicted monitoring indicator data set corresponding to each training sample in the at least one training sample and the corresponding target monitoring indicator data set.

[0080] Sub-step three involves determining whether the initial neural network has achieved the preset optimization objective based on the comparison results. This optimization objective can be that the difference between the predicted monitoring indicator data set corresponding to the training samples and the corresponding target monitoring indicator data set is less than or equal to a preset threshold.

[0081] Sub-step four: In response to determining that the initial neural network has achieved the above optimization objective, the initial neural network is used as the trained joint prediction model for health status.

[0082] The second step involves adjusting the network parameters of the initial neural network in response to the determination that it has not achieved the aforementioned optimization objective. This is done by removing previously used training samples from the training set to update the training set. The adjusted initial neural network is then used as the new initial neural network, and the training steps described above are executed again based on the updated training set. As an example, the back propagation algorithm (BP algorithm) and gradient descent methods (such as mini-batch gradient descent) can be used to adjust the network parameters of the initial neural network.

[0083] Step 106: Extract the monitoring indicator data sequence to obtain the monitoring indicator data sequence to be predicted.

[0084] In some embodiments, the aforementioned execution entity may perform truncation processing on the aforementioned monitoring indicator data sequence to obtain a monitoring indicator data sequence to be predicted.

[0085] In some optional implementations of certain embodiments, the aforementioned execution entity may perform truncation processing on the aforementioned monitoring indicator data sequence through the following steps to obtain the monitoring indicator data sequence to be predicted:

[0086] The first step is to extract the last monitoring indicator data and a predetermined number of monitoring indicator data before it from the monitoring indicator data sequence, resulting in a extracted monitoring indicator data sequence. The difference between the predetermined number and the predetermined length is 1.

[0087] The second step is to determine the above-mentioned intercepted monitoring indicator data sequence as the monitoring indicator data sequence to be predicted.

[0088] Step 107: Input the data sequence of the monitoring indicators to be predicted into the joint prediction model of health status to obtain the data sequence of the predicted monitoring indicators.

[0089] In some embodiments, the execution entity may input the data sequence of the monitoring indicators to be predicted into the joint health status prediction model to obtain a predicted monitoring indicator data sequence. The predicted monitoring indicator data sequence may be server indicator data (including data sequences of temperature, power consumption, load, and status) corresponding to a predicted future time period (e.g., the next ten minutes).

[0090] Step 108: Based on the predictive monitoring indicator data sequence and load indicator relationship model, perform future potential risk warning processing on the preset servers to be monitored.

[0091] In some embodiments, the aforementioned execution entity may perform future potential risk warning processing on the aforementioned preset servers to be monitored based on the predicted monitoring indicator data sequence and the aforementioned load indicator relationship model.

[0092] In some optional implementations of certain embodiments, the aforementioned execution entity can perform future potential risk warning processing on the aforementioned preset server to be monitored based on the predicted monitoring indicator data sequence and the aforementioned load indicator relationship model through the following steps:

[0093] The first step is to perform the following steps for each predicted monitoring indicator data in the above predicted monitoring indicator data sequence:

[0094] The first sub-step involves determining the load index relationship sub-model corresponding to the predicted monitoring index data from the aforementioned load index relationship model as the future load index relationship sub-model. The predicted monitoring index data includes temperature data, power consumption data, load data, and status data. In practice, the executing entity can determine the load data included in the predicted monitoring index data as the predicted load data. Then, in the quantile historical load range information sequence, the quantile historical load range information corresponding to the quantile historical load range where the predicted load data is located is determined as the matching load range information. Next, the load index relationship sub-model corresponding to the matching load range information among the various load index relationship sub-models included in the load index relationship model is determined as the future load index relationship sub-model.

[0095] The second sub-step, in response to determining that the temperature data included in the aforementioned predictive monitoring index data is not within the range represented by the historical normal temperature range information included in the future load index relationship sub-model, generates a temperature deviation value based on the aforementioned temperature data and the aforementioned historical normal temperature range information.

[0096] The third sub-step involves, in response to determining that the temperature deviation is less than or equal to a preset value, identifying the preset information indicating temperature anomalies and the preset repair information as potential future risk warning information, and sending the potential future risk warning information to the maintenance terminal. The preset information indicating temperature anomalies can be text information (e.g., a preset message that the temperature of the server to be monitored is too high). The preset repair information can be a text description of the repair strategy corresponding to the preset information indicating temperature anomalies (e.g., checking and cleaning the CPU cooler fan).

[0097] The fourth sub-step involves, in response to determining that the temperature deviation exceeds a preset value, shutting down the aforementioned preset monitored server and sending a preset temperature anomaly emergency message to the maintenance terminal. This preset temperature anomaly emergency message can be a preset warning message indicating a temperature anomaly (e.g., thermal runaway has been forcibly shut down, requiring immediate on-site handling).

[0098] The above-mentioned technical solution and related content, as an inventive point of this disclosure, solve the technical problem of "increased frequency of server hardware failures". Factors leading to an increased frequency of server hardware failures are often as follows: When handling potential future risks, traditional solutions do not differentiate the severity of anomalies when facing future temperature anomalies, adopting a uniform warning or handling mode. If minor anomalies are only given a simple warning without providing targeted repair guidance, it may lead to delayed operation and maintenance response, easily causing server hardware failure (e.g., after receiving a warning, operation and maintenance personnel need to spend extra time investigating the anomaly, extending the operation and maintenance response time). If severe anomalies are not addressed with timely forced loss mitigation measures, operation and maintenance personnel may miss the optimal repair opportunity, leading to server hardware failure and an increased frequency of server hardware failures. Solving these factors can reduce the frequency of server hardware failures. To achieve this effect, firstly, for each predicted monitoring indicator data in the above predicted monitoring indicator data sequence, the following steps are performed: Step 1, determine the load indicator relationship sub-model corresponding to the above predicted monitoring indicator data from the above load indicator relationship model as the future load indicator relationship sub-model. Thus, a future load indicator relationship sub-model containing historical normal temperature range information corresponding to future load scenarios can be obtained. The second step involves generating a temperature deviation value based on the temperature data and historical normal temperature range information included in the future load indicator relationship sub-model, in response to the determination that the temperature data is outside the range represented by the historical normal temperature range information. This generates a temperature deviation value that characterizes the severity of future temperature anomalies. The third step involves identifying the temperature deviation value as less than or equal to a preset value, defining the preset information and preset repair information representing the temperature anomaly as potential future risk warning information, and sending this warning to the maintenance terminal. This allows for the sending of potential future risk warning information, including preset information and preset repair information, for minor future temperature anomalies. Maintenance personnel can directly perform repair work without additional investigation, significantly shortening response time and effectively avoiding hardware damage caused by delayed response. The fourth step involves shutting down the preset monitored server and sending the preset temperature anomaly emergency information to the maintenance terminal, in response to the determination that the temperature deviation value is greater than the preset value. Therefore, in response to severe future temperature anomalies, the system can directly prevent further damage to the hardware by immediately shutting down the server as a stopgap measure. Simultaneously, it sends emergency messages to alert maintenance personnel to intervene, avoiding missing the optimal repair window and reducing the risk of server hardware failure. Furthermore, because it matches appropriate warnings and responses to different levels of anomalies, it effectively solves the problems of delayed maintenance response and untimely handling of severe anomalies in the traditional unified model, reducing the likelihood and frequency of server hardware failure.

[0099] The above embodiments of this disclosure have the following beneficial effects: the server anomaly detection and early warning methods of some embodiments of this disclosure reduce the waste of early warning resources and achieve early identification of potential risks and early warning. Specifically, the waste of early warning resources is due to the fact that fixed temperature and power consumption thresholds may not be suitable for different workloads and environmental conditions. As the load changes, the normal operating temperature and power consumption of servers with instantaneous high loads will also fluctuate. Using static thresholds may lead to frequent false alarms or missed alarms, resulting in a waste of early warning resources. Moreover, real-time collection of server temperature and power consumption, and anomaly detection and early warning based on fixed temperature / power consumption thresholds, cannot identify potential risks and provide early warning. Based on this, the server anomaly detection and early warning methods of some embodiments of this disclosure first obtain a historical monitoring indicator data sequence corresponding to a preset server to be monitored. Thus, a load indicator relationship model for creating a load indicator relationship model can be obtained. Then, based on the above historical monitoring indicator data sequence, a load indicator relationship model is created. Thus, a load indicator relationship model including the correspondence between indicator ranges (e.g., temperature range, power consumption range) under different loads can be created. Next, within the preset monitoring period, adaptive dynamic monitoring is performed on the preset servers to be monitored. Monitoring indicator data is collected at a dynamic acquisition frequency, and based on the aforementioned load indicator relationship model, real-time anomaly detection and early warning processing are performed on the collected monitoring indicator data. The dynamic acquisition frequency adjusts the data collection frequency in real time according to changes in server load. The acquisition frequency is increased when load changes drastically to capture indicator changes more promptly, while the acquisition frequency is reduced when the load is stable to minimize unnecessary resource consumption. Simultaneously, real-time anomaly detection using the load indicator relationship model dynamically adjusts the monitoring thresholds within the normal range based on real-time load conditions. When server load changes, the model dynamically adjusts the judgment criteria for the normal range of indicators based on the current load, thus adapting to different workloads and environmental conditions. This avoids potential misjudgments that might occur with fixed thresholds under different loads, thereby reducing false alarms or missed alarms and minimizing the waste of early warning resources. Next, the monitoring indicator data collected within the preset monitoring period is sorted to obtain a monitoring indicator data sequence. This yields the monitoring indicator data sequence for the preset monitoring period. Then, based on the aforementioned monitoring indicator data sequence, a preset joint health status prediction model is trained to obtain the joint health status prediction model. Therefore, the pre-defined health status joint prediction model can be trained to learn the changes in server indicators at different time points. Next, the monitoring indicator data sequence is truncated to obtain the monitoring indicator data sequence to be predicted. This yields the monitoring indicator data sequence reflecting the current state and recent trends of the server. Then, the monitoring indicator data sequence to be predicted is input into the aforementioned health status joint prediction model to obtain the predicted monitoring indicator data sequence.Therefore, a pre-trained joint health status prediction model can be used to predict the data sequence of monitoring indicators for the server over a future period, based on the data sequence of the monitoring indicators to be predicted. Finally, based on the predicted monitoring indicator data sequence and the aforementioned load indicator relationship model, potential risk warnings are issued for the pre-defined servers to be monitored. This achieves the early identification and warning of potential risks in server operation.

[0100] Further reference Figure 2 As an implementation of the methods shown in the figures, this disclosure provides some embodiments of a server anomaly detection and early warning device, which are similar to... Figure 1 Corresponding to the method embodiments shown, the device can be specifically applied to various electronic devices.

[0101] like Figure 2 As shown, a server anomaly detection and early warning device 200 in some embodiments includes: an acquisition unit 201, a creation unit 202, an adaptive dynamic monitoring unit 203, a sorting unit 204, a training unit 205, an interception unit 206, an input unit 207, and a future potential risk early warning unit 208. The acquisition unit 201 is configured to acquire a sequence of historical monitoring indicator data corresponding to a preset server to be monitored; the creation unit 202 is configured to create a load indicator relationship model based on the aforementioned historical monitoring indicator data sequence; the adaptive dynamic monitoring unit 203 is configured to perform adaptive dynamic monitoring processing on the preset server to be monitored within a preset monitoring time period, to collect monitoring indicator data according to a dynamic collection frequency, and to perform real-time anomaly detection and early warning processing on the collected monitoring indicator data based on the aforementioned load indicator relationship model; the sorting unit 204 is configured to sort the various monitoring indicator data collected within the preset monitoring time period to obtain monitoring indicators. The data sequence; the training unit 205 is configured to train the preset joint prediction model of health status based on the above monitoring indicator data sequence to obtain the joint prediction model of health status; the interception unit 206 is configured to intercept the above monitoring indicator data sequence to obtain the monitoring indicator data sequence to be predicted; the input unit 207 is configured to input the above monitoring indicator data sequence to be predicted into the above joint prediction model of health status to obtain the predicted monitoring indicator data sequence; the future potential risk warning unit 208 is configured to perform future potential risk warning processing on the above preset server to be monitored based on the predicted monitoring indicator data sequence and the above load indicator relationship model.

[0102] It is understandable that the units described in the device 200 are related to the reference. Figure 1 The steps in the method described above correspond to each other. Therefore, the operations, features, and beneficial effects described above for the method also apply to the device 200 and the units contained therein, and will not be repeated here.

[0103] The following is for reference. Figure 3 It shows a schematic diagram of the structure of an electronic device 300 suitable for implementing some embodiments of the present disclosure. Figure 3 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments of this disclosure.

[0104] like Figure 3 As shown, the electronic device 300 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 301, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 302 or a program loaded from a storage device 308 into a random access memory (RAM) 303. The RAM 303 also stores various programs and data required for the operation of the electronic device 300. The processing unit 301, ROM 302, and RAM 303 are interconnected via a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.

[0105] Typically, the following devices can be connected to I / O interface 305: input devices 306 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 307 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 308 including, for example, magnetic tapes, hard disks, etc.; and communication devices 309. Communication device 309 allows electronic device 300 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 3 An electronic device 300 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively. Figure 3 Each box shown can represent a device or multiple devices as needed.

[0106] In particular, according to some embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, some embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication device 309, or installed from storage device 308, or installed from ROM 302. When the computer program is executed by processing device 301, it performs the functions defined in the methods of some embodiments of this disclosure.

[0107] It should be noted that, in some embodiments of this disclosure, the computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium may be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In some embodiments of this disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In some embodiments of this disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0108] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.

[0109] Computer-readable media may be contained within an electronic device or exist independently, not assembled into the electronic device. The computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to: acquire a historical monitoring indicator data sequence corresponding to a preset server to be monitored; create a load indicator relationship model based on the aforementioned historical monitoring indicator data sequence; perform adaptive dynamic monitoring processing on the preset server to be monitored within a preset monitoring time period, to collect monitoring indicator data according to a dynamic acquisition frequency, and perform real-time anomaly detection and early warning processing on the collected monitoring indicator data based on the aforementioned load indicator relationship model; sort the monitoring indicator data collected within the preset monitoring time period to obtain a monitoring indicator data sequence; train a preset joint health status prediction model based on the aforementioned monitoring indicator data sequence to obtain a joint health status prediction model; truncate the aforementioned monitoring indicator data sequence to obtain a monitoring indicator data sequence to be predicted; input the aforementioned monitoring indicator data sequence to be predicted into the aforementioned joint health status prediction model to obtain a predicted monitoring indicator data sequence; and perform future potential risk early warning processing on the aforementioned preset server to be monitored based on the predicted monitoring indicator data sequence and the aforementioned load indicator relationship model.

[0110] Computer program code for performing operations of some embodiments of this disclosure can be written in one or more programming languages ​​or a combination thereof. Programming languages ​​include object-oriented programming languages—such as Java, Smalltalk, and C++—and conventional procedural programming languages—such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0111] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0112] The units described in some embodiments of this disclosure can be implemented in software or hardware. The described units can also be housed in a processor; for example, a processor may be described as including an acquisition unit, a creation unit, an adaptive dynamic monitoring unit, a training unit, an interception unit, an input unit, and a future potential risk warning unit. The names of these units do not necessarily limit the unit itself; for example, the acquisition unit may also be described as "a unit that acquires a sequence of historical monitoring indicator data corresponding to a preset server to be monitored."

[0113] The functions described above in this document can be performed at least in part by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), and so on.

[0114] The above description is merely a selection of preferred embodiments of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of the invention involved in the embodiments of this disclosure is not limited to technical solutions formed by specific combinations of technical features, but should also cover other technical solutions formed by arbitrary combinations of technical features or their equivalents without departing from the inventive concept. For example, technical solutions formed by substituting features with (but not limited to) technical features with similar functions disclosed in the embodiments of this disclosure.

Claims

1. A method for server anomaly detection and early warning, comprising: Obtain the historical monitoring indicator data sequence corresponding to the preset server to be monitored; Based on the historical monitoring indicator data sequence, a load indicator relationship model is created, wherein each historical monitoring indicator data in the historical monitoring indicator data sequence includes historical temperature data, historical power consumption data, historical load data, and historical status data, and each historical monitoring indicator data corresponds to a collection time. The load indicator relationship model created based on the historical monitoring indicator data sequence includes: The historical load data sequence is obtained by sorting the various historical load data included in the historical monitoring indicator data sequence. The historical load data sequence is quantized to obtain a quantized historical load range information sequence and a historical load data group sequence, wherein each quantized historical load range information in the quantized historical load range information sequence corresponds to a historical load data group in the historical load data group sequence. A load indicator relationship model is created based on the historical load data group sequence and the historical monitoring indicator data sequence. Within a preset monitoring period, adaptive dynamic monitoring is performed on a preset server to be monitored, so as to collect monitoring index data according to the dynamic collection frequency, and to perform real-time anomaly detection and early warning processing on the collected monitoring index data based on the load index relationship model. The monitoring indicator data collected during the preset monitoring period will be sorted to obtain a monitoring indicator data sequence; Based on the monitoring indicator data sequence, a preset joint prediction model for health status is trained to obtain the joint prediction model for health status, including: The monitoring indicator data sequence is grouped to obtain a monitoring indicator data group sequence; For every two consecutive monitoring indicator data groups in the monitoring indicator data group sequence, perform the following steps: The monitoring indicator data group that precedes the other monitoring indicator data group in two consecutive monitoring indicator data groups is defined as the sample monitoring indicator data group, and the other monitoring indicator data group is defined as the sample target monitoring indicator data group corresponding to the sample monitoring indicator data group. The sample monitoring indicator data set and the sample target monitoring indicator data set corresponding to the sample monitoring indicator data set are determined as training samples; Each of the identified training samples is designated as the training sample set; Based on the training sample set, the preset joint prediction model of health status is trained to obtain the joint prediction model of health status. The monitoring indicator data sequence is truncated to obtain the monitoring indicator data sequence to be predicted; The data sequence of the monitoring indicators to be predicted is input into the joint prediction model of health status to obtain the data sequence of the predicted monitoring indicators. Based on the predicted monitoring indicator data sequence and the load indicator relationship model, the preset servers to be monitored are subjected to future potential risk warning processing.

2. The method of claim 1, wherein, The step of creating a load indicator relationship model based on the historical load data group sequence and the historical monitoring indicator data sequence includes: For each historical load data group in the historical load data group sequence, perform the following steps: For each historical load data in the historical load data group, the historical monitoring indicator data containing the historical load data in the historical monitoring indicator data sequence is determined as the target historical monitoring indicator data. The historical temperature data included in the historical monitoring index data of each target is defined as a historical temperature data group. The historical power consumption data included in the historical monitoring index data of each target are defined as historical power consumption data groups. The quantile historical load range information corresponding to the historical load data group in the quantile historical load range information sequence is determined as the target quantile historical load range information. The historical temperature data set is subjected to normal range definition processing to obtain historical normal temperature range information; The historical power consumption data group is subjected to normal range definition processing to obtain historical normal power consumption range information; The historical normal temperature range information and the historical normal power consumption range information are determined as a load index relationship sub-model corresponding to the target quantile historical load range information; The determined sub-models of the load index relationships are defined as the load index relationship model.

3. The method according to claim 1, wherein, The step of truncating the monitoring indicator data sequence to obtain the monitoring indicator data sequence to be predicted includes: The last monitoring indicator data and a predetermined number of monitoring indicator data before the last monitoring indicator data are extracted from the monitoring indicator data sequence to obtain the extracted monitoring indicator data sequence. The extracted monitoring indicator data sequence is determined as the monitoring indicator data sequence to be predicted.

4. The method of claim 1, wherein, The step of training a preset joint prediction model of health status based on the training sample set to obtain the joint prediction model of health status includes: The following training steps are performed based on the training sample set: Input the sample monitoring index data set of at least one training sample in the training sample set into the initial neural network to obtain the prediction monitoring index data set corresponding to each training sample in the at least one training sample. Compare the predicted monitoring index data set corresponding to each training sample in the at least one training sample with the corresponding sample target monitoring index data set. Based on the comparison results, determine whether the initial neural network has achieved the preset optimization objective; In response to determining that the initial neural network has achieved the optimization objective, the initial neural network is used as a trained joint prediction model for health status. In response to determining that the initial neural network has not reached the optimization objective, the network parameters of the initial neural network are adjusted, and the training samples used in the training sample set are deleted to update the training sample set. The adjusted initial neural network is then used as the initial neural network, and the training steps are performed again based on the updated training sample set.

5. A server anomaly detection and early warning device, comprising: The acquisition unit is configured to acquire a sequence of historical monitoring indicator data corresponding to a preset server to be monitored. A creation unit is configured to create a load indicator relationship model based on the historical monitoring indicator data sequence. Each historical monitoring indicator data in the historical monitoring indicator data sequence includes historical temperature data, historical power consumption data, historical load data, and historical status data. Each historical monitoring indicator data corresponds to a collection time. The creation of the load indicator relationship model based on the historical monitoring indicator data sequence includes: sorting the historical load data included in the historical monitoring indicator data sequence to obtain a historical load data sequence; performing quantile processing on the historical load data sequence to obtain a quantile historical load range information sequence and a historical load data group sequence, wherein each quantile historical load range information in the quantile historical load range information sequence corresponds to a historical load data group in the historical load data group sequence; and creating the load indicator relationship model based on the historical load data group sequence and the historical monitoring indicator data sequence. The adaptive dynamic monitoring unit is configured to perform adaptive dynamic monitoring on a preset server to be monitored within a preset monitoring time period, so as to collect monitoring index data according to the dynamic acquisition frequency, and to perform real-time anomaly detection and early warning processing on the collected monitoring index data based on the load index relationship model. The sorting unit is configured to sort the monitoring indicator data collected during a preset monitoring time period to obtain a monitoring indicator data sequence. The training unit is configured to train a preset joint prediction model for health status based on the monitoring indicator data sequence to obtain a joint prediction model for health status. The training includes: grouping the monitoring indicator data sequence to obtain a monitoring indicator data group sequence; for every two consecutive monitoring indicator data groups in the monitoring indicator data group sequence, performing the following steps: determining the monitoring indicator data group preceding the other monitoring indicator data group in the sequence as a sample monitoring indicator data group, and determining the other monitoring indicator data group as a sample target monitoring indicator data group corresponding to the sample monitoring indicator data group; determining the sample monitoring indicator data group and the sample target monitoring indicator data group corresponding to the sample monitoring indicator data group as training samples; determining each determined training sample as a training sample set; and training the preset joint prediction model for health status based on the training sample set to obtain the joint prediction model for health status. The interception unit is configured to intercept the monitoring indicator data sequence to obtain the monitoring indicator data sequence to be predicted. The input unit is configured to input the data sequence of the monitoring indicator to be predicted into the joint prediction model of health status to obtain the predicted monitoring indicator data sequence. The potential risk warning unit is configured to perform potential risk warning processing on the preset server to be monitored based on the predictive monitoring indicator data sequence and the load indicator relationship model.

6. An electronic device, comprising: One or more processors; A storage device on which one or more programs are stored; When the one or more programs are executed by the one or more processors, the one or more processors implement a method as claimed in any of claims 1 to 4.

7. A computer readable medium having stored thereon a computer program, wherein, The program, which when executed by a processor, implements a method as claimed in any of claims 1 to 4.

Citation Information

Patent Citations

  • Server health diagnosis method and system, electronic equipment and storage medium

    CN117194188A

  • Server health examination method, apparatus and device, and medium

    CN121116778A