Server failure prediction method and apparatus, electronic device, and storage medium

By extracting the timing multi-dimensional characteristics of the server running status timing data, the problem of the inability to predict server failures in the prior art is solved, and a higher fault prediction accuracy and recall rate are achieved.

WO2025133748A1PCT designated stage expired Publication Date: 2025-06-26CLOUD INTELLIGENCE ASSETS HOLDING (SINGAPORE) PTE LTD

Patent Information

Application Number
PCT/IB2024/061529
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-22
Filing Date
2024-11-19
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

The prior art is difficult to predict server failures in advance, resulting in the inability to actively carry out data backup and service migration.

Method used

By obtaining the server's running status timing data, cutting and stacking data, extracting timing multi-dimensional features, and using these features for failure prediction.

Benefits of technology

It realizes early prediction of server failures, improves the accuracy and recall of fault prediction, and enables data backup and service migration in advance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IB2024061529_26062025_PF_FP_ABST
    Figure IB2024061529_26062025_PF_FP_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure relate to the technical field of cloud computing, and provide a server failure prediction method and apparatus, an electronic device, and a storage medium. The method comprises: first, acquiring time series data of the operating state of a server; cutting the time series data of the operating state according to a period, and stacking a plurality of pieces of sub-time series data obtained by cutting; extracting time series multi-dimensional features from the stacked data; and performing failure prediction on the server by using the time series multi-dimensional features. In the embodiments, the time series multi-dimensional features are extracted on the basis of the time series data of the operating state for failure prediction, so that a server failure can be predicted in advance. Additionally, the time series data of the operating state are cut according to the period and stacked, and then the time series multi-dimensional features are extracted, so that related information on the fluctuation trend of data over time and different performances of the data before and after the period can be extracted, thereby improving the accuracy and recall rate of failure prediction.
Need to check novelty before this filing date? Find Prior Art

Description

[0001]Server Failure Prediction Method, Apparatus, Electronic Device, and Storage Medium Cross-Reference This disclosure claims priority to Chinese patent application No. 202311793866.8, filed with the Patent Office of China on December 22, 2023, entitled "Server Failure Prediction Method, Apparatus, Electronic Device, and Storage Medium," the entire contents of which are incorporated herein by reference. Technical Field This disclosure relates to the field of cloud computing, and more particularly to a server failure prediction method, apparatus, electronic device, and storage medium. Background: Servers are the foundation of cloud computing, and their stability directly determines the reliability of cloud computing services. Server stability largely depends on the stability of their components. For example, failures in components such as hard drives, memory, central processing units (CPUs), and graphics processing units (GPUs) can cause fluctuations in upper-layer applications at best, or even data loss and server downtime at worst. Currently, most response solutions rely on the domain knowledge of operations and maintenance experts and hardware specialists to develop a series of rules for detection and perform reactive operations and maintenance after a failure occurs. Analyzing and predicting possible future failures using current monitoring data, and proactively performing operations such as data backup and service migration, has become a key concern in proactive operations and maintenance. SUMMARY OF THE INVENTION Embodiments of the present disclosure provide a server failure prediction method, apparatus, electronic device, and storage medium to address the inability to predict server failures in advance. In a first aspect, embodiments of the present disclosure provide a server failure prediction method, comprising: obtaining server operating status time series data; segmenting the operating status time series data into periods and stacking the resulting multiple sub-series data; extracting time series multidimensional features from the stacked data; and using the time series multidimensional features to predict server failures. In a second aspect, embodiments of the present disclosure provide a server fault prediction device, comprising: a data acquisition component configured to acquire server operating status time series data; a data stacking component configured to segment the operating status time series data into periods and stack the resulting multiple sub-time series data; a feature extraction component configured to extract multi-dimensional time series features from the stacked data; and a fault prediction component configured to use the multi-dimensional time series features to predict server faults. In a third aspect, embodiments of the present disclosure provide an electronic device, comprising a memory, a processor, and a computer program stored in the memory, wherein the processor implements any of the aforementioned methods when executing the computer program.In a fourth aspect, embodiments of the present disclosure provide a computer-readable storage medium storing a computer program. When executed by a processor, the computer program implements any of the aforementioned methods. In a fifth aspect, embodiments of the present disclosure further provide a computer program product, including a computer program. When executed by a processor, the computer program implements any of the aforementioned server failure prediction methods. In a sixth aspect, embodiments of the present disclosure further provide a computer program product, including a non-volatile computer-readable storage medium storing a computer program. When executed by a processor, the computer program implements any of the aforementioned server failure prediction methods. In a seventh aspect, embodiments of the present disclosure further provide a computer program. When executed by a processor, the computer program implements any of the aforementioned server failure prediction methods. Compared to the prior art, the embodiments of the present disclosure have the following advantages: The embodiments of the present disclosure provide a server fault prediction method, apparatus, electronic device, and storage medium. First, server operating status time series data is acquired; the operating status time series data is segmented by period, and the resulting multiple sub-series data are stacked; time series multidimensional features are extracted from the stacked data; and server fault prediction is performed using the time series multidimensional features. In this embodiment, fault prediction based on the time series multidimensional features extracted from the operating status time series data can achieve early prediction of server faults. Furthermore, segmenting and stacking the operating status time series data by period, followed by extraction of the time series multidimensional features, can extract information related to the data's fluctuation trends and the different performance of the data in previous and subsequent periods, thereby improving the accuracy and recall of fault prediction. The above description is merely an overview of the technical solutions of the embodiments of the present disclosure. To better understand the technical solutions of the embodiments of the present disclosure, implementation can be carried out in accordance with the present specification. To further enhance the understanding of the above and other purposes, features, and advantages of the embodiments of the present disclosure, the following describes specific implementations of the embodiments of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS In the accompanying drawings, unless otherwise specified, identical reference numerals throughout the multiple figures denote identical or similar components or elements. The drawings are not necessarily drawn to scale. It should be understood that these drawings merely depict some implementations according to the embodiments of the present disclosure and should not be construed as limiting the scope of the embodiments of the present disclosure. FIG1 is a schematic diagram of an application scenario of the server failure prediction method provided in an embodiment of the present disclosure. FIG2 is a flowchart of the server failure prediction method in an embodiment of the present disclosure. FIG3 is a schematic diagram of the time series multidimensional feature construction in an embodiment of the present disclosure. FIG4 is a flowchart of the server failure prediction method in an embodiment of the present disclosure. FIG5 is a flowchart of the server failure prediction method in an embodiment of the present disclosure.Figure 6 is a block diagram of a server failure prediction device according to an embodiment of the present disclosure. Figure 7 is a block diagram of an electronic device used to implement the embodiments of the present disclosure. The following briefly describes certain exemplary embodiments. As those skilled in the art will appreciate, the described embodiments may be modified in various ways without departing from the spirit or scope of the embodiments of the present disclosure. Therefore, the drawings and descriptions are to be considered illustrative in nature, rather than restrictive. To facilitate understanding of the technical solutions of the embodiments of the present disclosure, the following describes related technologies of the embodiments of the present disclosure. The following related technologies are optional solutions that can be combined with the technical solutions of the embodiments of the present disclosure in any manner and fall within the scope of protection of the embodiments of the present disclosure. It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, storage, and display, etc.) involved in the embodiments of the present disclosure are all information and data authorized by the user or fully authorized by all parties. The collection, use, and processing of the relevant data must comply with the relevant laws, regulations, and standards of the relevant countries and regions, and corresponding operation portals are provided for the user to choose to authorize or deny. Figure 1 is a schematic diagram of an application scenario of the server failure prediction method provided by the embodiments of the present disclosure. As shown in Figure 1, this process specifically includes the following steps: data preprocessing, which obtains server operating status time series data. For example, hard drive SMART time series data is obtained from the Self-Monitoring Analysis and Reporting Technology (SMART) log (see time series data 1 in Figure 1); key performance indicators (KPIs) of the server's component monitoring time series data (see time series data 2 in Figure 1), such as CPU utilization, memory utilization, power consumption, hard drive write rate, and hard drive write speed; and static data, namely, server-specific attribute data, such as machine model and component model data. Data cleaning is performed on each of these data, and missing data is then supplemented. Feature generation, which constructs multidimensional and discrete features from the time series.Specifically, features are generated for the preprocessed data. SMART time series data and KPI data are segmented into sub-series data of the same period length, and the resulting sub-series data of the same period length are stacked. A multidimensional tensor is constructed based on the acquisition time, data attributes, and stacked data of the SMART time series data and KPI data. A convolution operation is performed on the multidimensional tensor to obtain time series multidimensional features. Discrete features are constructed using static data, and the time series multidimensional features and discrete features are then concatenated for subsequent model training. During model training, a training sample set is constructed using the concatenated features. The training sample set includes multiple training sample pairs, each of which includes sample data and a sample label. The concatenated features are used as the sample data, and a sample label is assigned. The sample label can indicate whether a fault has occurred and the faulty component. A machine learning model is trained using a training sample set. The machine learning model can be a classification model or anomaly detection model, such as a decision tree model, a neural network model, or a variational autoencoder (VAE). A prediction output is generated, and the trained machine learning model is used to perform a fault prediction task and output a prediction result. Specifically, time series data and static data to be predicted are obtained, and splicing features are constructed according to the above method. The splicing features are input into the trained machine learning model, and the machine learning model outputs a fault prediction result for the server component. In this embodiment, fault prediction is performed by extracting multidimensional time series features based on the operating status time series data, which can achieve early prediction of server failures. Furthermore, by segmenting and stacking the operating status time series data according to cycles and then extracting multidimensional time series features, information about the fluctuation trends of the data and the different performance of the data in the previous and next cycles can be extracted, thereby improving the precision and recall of fault prediction. The disclosed embodiments provide a server fault prediction method. This method can be applied to computing devices, such as servers and user terminals. FIG2 is a flowchart of a server fault prediction method according to an embodiment of the present disclosure, comprising: Step S201: Acquiring server operating status time series data; Step S202: Slicing the operating status time series data into periods and stacking the resulting multiple sub-time series data; Step S203: Extracting time series multidimensional features from the stacked data; and Step S204: Using the time series multidimensional features, performing server fault prediction.The operating status time series data includes data representing the operating status of server components that changes over time as the server runs, such as hard drive SMART time series data and KPIs for server component monitoring time series data. KPIs include, but are not limited to, CPU usage, memory usage, power consumption, hard drive write rate, and hard drive write speed. The operating status time series data can be single-dimensional, such as hard drive time series data, or multi-dimensional, such as hard drive, memory, and CPU time series data. The server fault prediction method provided in the disclosed embodiments first obtains server operating status time series data; then segments the operating status time series data into periods and stacks the resulting multiple sub-series data; extracts multi-dimensional time series features from the stacked data; and uses these multi-dimensional time series features to predict server faults. In this embodiment, fault prediction based on the extraction of multi-dimensional time series features from the operating status time series data can achieve early prediction of server faults. Furthermore, by segmenting and stacking the operating status time series data by period and then extracting multidimensional time series features, information about the data's fluctuation trends and the different performance of the data in the previous and next periods can be extracted, thereby improving the precision and recall of fault prediction. The specific implementation process of each of the above steps is described below through multiple implementations: In one implementation, step S202, segmenting the operating status time series data by period and stacking the multiple sub-time series data obtained by segmentation, includes segmenting the operating status time series data by period length and stacking the sub-time series data with the same period length. In practical applications, data segmentation can be performed by the same period length. For example, seven days of operating status data for a hard drive can be obtained by segmenting it by period length of one day, resulting in seven sub-time series data with a period length of one day. These seven sub-time series data are then stacked. In this embodiment, stacking data with the same period length and then extracting features for prediction can extract the different performance of the time series data in the previous and next periods, thereby improving the precision and recall of the prediction results. In one implementation, step S202, cutting the operating status time series data by period and stacking the multiple sub-time series data obtained by the cutting, includes cutting the operating status time series data by different period lengths and stacking the sub-time series data of different period lengths obtained by the cutting. In practical applications, data can be cut by different period lengths and stacked according to the respective period lengths to obtain tensors corresponding to the respective period lengths.For example, 60 days of hard drive runtime data is obtained and segmented into periods of 1 day, 7 days, and 30 days. The resulting sub-time series data with periods of 1 day, 7 days, and 30 days are stacked according to their respective period lengths. Convolution processing is then performed on the tensors corresponding to 1 day, 7 days, and 30 days to extract their corresponding features. The features corresponding to 1 day, 7 days, and 30 days are then concatenated to obtain multidimensional time series features, which are then input into a machine learning model for fault prediction. In this embodiment, stacking data with the same period length and then extracting features for prediction can extract the fluctuation trends of the time series data at different period lengths, thereby improving the precision and recall of the prediction results. In one implementation, step S203 extracts time series multidimensional features from the stacked data, including: step S2031 constructing a multidimensional tensor based on the stacked data; and step S2032 performing a convolution operation on the multidimensional tensor to obtain time series multidimensional features. Data cut according to different cycle lengths is stacked according to their respective cycle lengths, and then the tensors corresponding to the different cycle lengths are convolved to extract features corresponding to each tensor. The features corresponding to the different cycle lengths are then concatenated to obtain time series multidimensional features, which are then input into a machine learning model for fault prediction. In practical applications, when constructing a multidimensional tensor, the stacked data is used as the data for one dimension of the multidimensional tensor. Data for other dimensions of the multidimensional tensor are then added as needed to obtain the multidimensional tensor. A convolution operation is then performed on the multidimensional tensor. Multiple multidimensional convolution kernels of different sizes, steps, and weights are set as needed. After performing the convolution operation using the multidimensional convolution kernels, a multidimensional time series feature is obtained. The dimension of the multidimensional time series feature is related to the size, step number, and weight of the convolution kernel. In one implementation, step S2031 constructs a multidimensional tensor based on the stacked data. This includes constructing a multidimensional tensor based on the acquisition time, data attributes, and stacked data of the operating status time series data. The data attributes indicate the dimension to which the data belongs. For example, when predicting hard drive failures on a server, hard drive time series data is used. Data attributes include hard drive write rate, hard drive write speed, and so on.In one example, time series multidimensional features are constructed by stacking the operating status time series data into a multidimensional tensor based on the size of the periodicity and extracting the time series multidimensional features using a multidimensional convolution kernel, as shown in Figure 3. The specific process is as follows: Period detection: Using periodicity detection methods (e.g., spectral methods such as Fourier transform, autocorrelation function methods, trend decomposition algorithms, etc.), the operating status time series data (e.g., multi-dimensional data such as SMART data and KPI data) is detected for a period length T, such as period 1 shown in Figure 3. If there is no periodicity, a default period length Td is set. Period cutting: Based on the detected period length, the operating status time series data is cut into multiple sub-time series data segments of length T. Period stacking: The multiple sub-time series data segments of length T are stacked into a multidimensional tensor. Figure 3 shows a 3D tensor, comprising dimensions (i.e., data attributes), time (within a cycle) (i.e., acquisition time), and cycle (i.e., stacked data). This can also be expanded to more than 3 dimensions based on actual scenario requirements. For example, the tensor constructed using this method for the memory graph used in memory fault prediction scenarios can be 4-dimensional or greater. Multidimensional convolution: Multidimensional convolution kernels of varying sizes are used to extract temporal multidimensional features from the constructed multidimensional tensor. The size of the convolution kernel can be selected based on specific needs and is not limited in this disclosure. For example, as shown in the grayscale portion of Figure 3, a 2 x 2 x 2 convolution kernel can be selected for the convolution operation. In this embodiment, multi-dimensional time series data of operating status across multiple periods and dimensions is constructed into a multidimensional tensor based on its data dimension, time, and periodicity. Convolution is then performed using convolution kernels of various sizes to extract features. This allows for the extraction of at least three aspects of information: correlations between time series data of different dimensions, fluctuation trends in the data, and the different performance of the data in previous and subsequent periods. Finally, the multidimensional time series features are input into a machine learning model for training and prediction, which can improve the precision and recall of fault prediction. In one implementation, after extracting the multidimensional time series features from the stacked data, the method further includes: obtaining inherent attribute data of the server, the inherent attribute data including at least one of the server model and component model; constructing discrete features based on the inherent attribute data; and performing server fault prediction using the multidimensional time series features in step S204, including: concatenating the discrete features with the multidimensional time series features to obtain concatenated features; and performing server fault prediction using the concatenated features. Discrete features are features with finite and discontinuous values. The value of at least one of the server model and component model may be mapped to a label or category to obtain a discrete feature.There are many specific ways to combine discrete features and time series multidimensional features. For example, a 5-dimensional discrete feature and a 100-dimensional time series multidimensional feature can be directly combined horizontally to obtain a 105-dimensional combined feature. In this embodiment, inherent attribute data is added to the time series data. The combined features obtained by combining discrete features and time series multidimensional features are used to predict server faults. This can predict the model or specific component that may fail in the future, making the prediction more accurate. In one implementation, step S204, using the time series multidimensional features to predict server faults, includes: inputting the time series multidimensional features into a machine learning model, and using the output of the machine learning model as the fault prediction result. The machine learning model may include a classification model or a fault detection model. The machine learning model is pre-trained. The machine learning model can be a classification model or anomaly detection model, such as a decision tree model, a neural network model, or a VAE. The machine learning model is trained in the following manner: First, a training sample set is constructed. Specifically, SMART time series data and KPI data are segmented according to the same cycle length, and the resulting sub-time series data of the same cycle length are stacked. A multidimensional tensor is constructed based on the acquisition time, data attributes, and stacked data of the SMART time series data and KPI data. A convolution operation is performed on the multidimensional tensor to obtain time series multidimensional features. A training sample set is constructed using the time series multidimensional features. The training sample set includes multiple training sample pairs, each of which includes sample data and a sample label. The time series multidimensional features are used as sample data, and sample labels are assigned. The sample labels can indicate whether a fault will occur, for example. A machine learning model is trained using the training sample set until a preset training end condition is met, thereby obtaining a trained machine learning model. The trained machine learning model is then used to perform a fault prediction task. The data to be predicted is converted into time series multidimensional features according to the above feature construction method, and the features are input into the trained machine learning model. The machine learning model then outputs a fault prediction result. The disclosed embodiments provide a server fault prediction method. The method in this embodiment can be applied to computing devices, such as servers and user terminals. FIG4 is a flowchart of a server fault prediction method according to an embodiment of the present disclosure, comprising: Step S401: Acquiring server operating status time series data; Step S402: Slicing the operating status time series data into segments of different cycle lengths and stacking the resulting sub-series data of different cycle lengths; Step S403: Constructing a multidimensional tensor based on the acquisition time, data attributes, and stacked data of the operating status time series data.In step S404, a convolution operation is performed on the multidimensional tensor to obtain time series multidimensional features. In step S405, the time series multidimensional features are input into a machine learning model, and the output of the machine learning model is used as the fault prediction result. The machine learning model may include a classification model or a fault detection model. The specific implementation process of each of the above steps is detailed in the specific implementation methods of the above embodiments and will not be repeated here. The present disclosure provides a server fault prediction method. The method in this embodiment can be applied to computing devices, which may include servers, user terminals, etc. FIG5 is a flowchart of the server fault prediction method according to one embodiment of the present disclosure, comprising: Step S501: Obtaining server operating status time series data. Step S502: Slicing the operating status time series data into segments of equal cycle length and stacking the obtained sub-series data of equal cycle length. Step S503: Constructing a multidimensional tensor based on the acquisition time, data attributes, and stacked data of the operating status time series data. Step S504: Convolution operation is performed on the multidimensional tensor to obtain time series multidimensional features. In step S505, inherent attribute data of the server is obtained, including at least one of the server model and component model. Discrete features are constructed based on the inherent attribute data. In step S506, the discrete features and the time series multidimensional features are concatenated to obtain concatenated features. In step S507, fault prediction for the server is performed using the concatenated features. The specific implementation of each of the above steps is detailed in the specific implementation methods in the above embodiments and will not be repeated here. Corresponding to the application scenarios and methods of the methods provided in the embodiments of the present disclosure, the embodiments of the present disclosure also provide a server fault prediction device. Figure 6 shows a block diagram of a server fault prediction device according to an embodiment of the present disclosure. The device includes: a data acquisition component 601 configured to acquire server operating status time series data; a data stacking component 602 configured to segment the operating status time series data into periods and stack the resulting sub-series data; a feature extraction component 603 configured to extract multidimensional time series features from the stacked data; and a fault prediction component 604 configured to use the multidimensional time series features to predict server faults. The server fault prediction device provided in an embodiment of the present disclosure first acquires server operating status time series data; segments the operating status time series data into periods and stacks the resulting sub-series data; extracts multidimensional time series features from the stacked data; and uses the multidimensional time series features to predict server faults.In this embodiment, fault prediction is performed by extracting multidimensional time series features from operating status time series data, enabling early prediction of server failures. Furthermore, by segmenting and stacking the operating status time series data by period and then extracting multidimensional time series features, information about the data's fluctuation trends and the different performances of the data in previous and subsequent periods can be extracted, improving the accuracy and recall of fault prediction. In one implementation, the data stacking component 602 is configured to segment the operating status time series data into sub-periods of the same length and stack the resulting sub-periods of the same length. In another implementation, the data stacking component 602 is configured to segment the operating status time series data into sub-periods of different lengths and stack the resulting sub-periods of different lengths. In another implementation, the feature extraction component 603 is configured to construct a multidimensional tensor based on the stacked data and perform a convolution operation on the multidimensional tensor to obtain the multidimensional time series features. In one implementation, when constructing a multidimensional tensor based on the stacked data, the feature extraction component 603 is configured to construct the multidimensional tensor based on the collection time, data attributes, and stacked data of the operating status time series data. In one implementation, the device is further configured to, after extracting the time series multidimensional features from the stacked data, obtain inherent attribute data of the server, the inherent attribute data including at least one of the server model and component model; construct discrete features based on the inherent attribute data; and the fault prediction component 604 is configured to concatenate the discrete features with the time series multidimensional features to obtain concatenated features; and use the concatenated features to predict server faults. In one implementation, the fault prediction component 604 is configured to input the time series multidimensional features into a machine learning model and use the output of the machine learning model as a fault prediction result; the machine learning model may include a classification model or a fault detection model. The functions of each component in the embodiments of the present disclosure can be found in the corresponding description of the above-mentioned method, and they have corresponding beneficial effects, and are not further described here. Figure 7 is a block diagram of an electronic device used to implement the embodiments of the present disclosure. As shown in Figure 7, the electronic device includes a memory 710 and a processor 720. The memory 710 stores a computer program executable on the processor 720. When the processor 720 executes the computer program, the method described in the above embodiment is implemented. The number of the memory 710 and the processor 720 can be one or more. The electronic device also includes a communication interface 730 configured to communicate with external devices and exchange data.If the memory 710, processor 720, and communication interface 730 are implemented independently, they can be interconnected via a bus and communicate with each other. This bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus. This bus can be divided into address buses, data buses, control buses, etc. For ease of illustration, FIG7 shows only one thick line, but this does not mean that there is only one bus or only one type of bus. Optionally, in a specific implementation, if the memory 710, processor 720, and communication interface 730 are integrated on a single chip, the memory 710, processor 720, and communication interface 730 can communicate with each other via an internal interface. The present embodiment provides a computer-readable storage medium storing a computer program. When executed by a processor, the program implements the method provided in the present embodiment. The present disclosure also provides a chip, including a processor configured to retrieve and execute instructions stored in a memory, so that a communication device equipped with the chip performs the methods provided in the present disclosure. The present disclosure also provides a chip, including an input interface, an output interface, a processor, and a memory. The input interface, the output interface, the processor, and the memory are connected via an internal connection path. The processor is configured to execute code in the memory. When the code is executed, the processor is configured to perform the methods provided in the present disclosure. It should be understood that the processor may be a CPU, other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.It is worth noting that the processor may be a processor supporting the Advanced RISC Machines (ARM) architecture. Furthermore, optionally, the aforementioned memory may include read-only memory and random access memory. The memory may be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. Non-volatile memory may include read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory may include random access memory (RAM), which serves as an external cache. By way of example and not limitation, many forms of RAM may be used. For example, static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronized link dynamic random access memory (SLDRAM) and direct rambus random access memory (DR RAM). oAccording to another aspect of the present disclosure, a computer program product is provided, including a computer program. Optionally, when executed by a processor, the computer program implements any of the aforementioned server failure prediction methods. According to another aspect of the present disclosure, a computer program product is provided, including a non-volatile computer-readable storage medium. Optionally, the non-volatile computer-readable storage medium stores a computer program, which, when executed by a processor, implements any of the aforementioned server failure prediction methods. According to another aspect of the present disclosure, a computer program is provided. Optionally, when executed by a processor, the computer program implements any of the aforementioned server failure prediction methods. In the above embodiments, all or part of the embodiments may be implemented using software, hardware, firmware, or any combination thereof. When implemented using software, all or part of the embodiments may be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions according to the embodiments of the present disclosure are generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. Computer instructions can be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another. Throughout this specification, references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples" indicate that the specific features, structures, materials, or characteristics described in conjunction with that embodiment or example are included in at least one embodiment or example of the disclosed embodiments. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples. Furthermore, those skilled in the art may combine and integrate the different embodiments or examples described in this specification, as well as features from different embodiments or examples, unless otherwise specified. Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed to indicate or imply relative importance or implicitly specify the number of technical features indicated. Therefore, features specified as "first" or "second" may explicitly or implicitly include at least one of those features. In the description of the disclosed embodiments, "plurality" means two or more, unless otherwise specifically defined. Any process or method described in the flowchart or otherwise described herein may be understood to represent a component, segment, or portion of code that includes one or more executable instructions for implementing the specific logical function or process steps.Furthermore, the scope of preferred embodiments of the disclosed embodiments includes alternative implementations in which functions may be performed out of the order shown or discussed, including performing functions substantially simultaneously or in reverse order depending on the functions involved. The logic and / or steps described in the flowcharts or otherwise herein may, for example, be considered a sequenced list of executable instructions for implementing the logical functions and may be embodied in any computer-readable medium for use by an instruction execution system, apparatus, or device (e.g., a computer-based system, a system including a processor, or other system that can fetch and execute instructions from an instruction execution system, apparatus, or device), or in conjunction with such an instruction execution system, apparatus, or device. It should be understood that various aspects of the disclosed embodiments may be implemented using hardware, software, firmware, or a combination thereof. In the aforementioned embodiments, multiple steps or methods may be implemented using software or firmware stored in memory and executed by a suitable instruction execution system. All or part of the steps of the methods of the aforementioned embodiments may be performed by a program that instructs the relevant hardware. The program may be stored in a computer-readable storage medium. When executed, the program includes one or a combination of the steps of the method embodiments. Furthermore, the functional components in the various embodiments of the present disclosure may be integrated into a single processing component, each component may exist physically separately, or two or more components may be integrated into a single component. These integrated components may be implemented in either hardware or software functional components. If these integrated components are implemented as software functional components and sold or used as standalone products, they may also be stored in a computer-readable storage medium. This storage medium may be a read-only memory, a magnetic disk, or an optical disk. The foregoing is merely an exemplary implementation of the embodiments of the present disclosure, but the scope of protection of the embodiments of the present disclosure is not limited thereto. Any person skilled in the art can readily conceive of various variations or substitutions within the technical scope of the embodiments of the present disclosure, and such variations or substitutions are intended to fall within the scope of protection of the embodiments of the present disclosure. Therefore, the scope of protection of the embodiments of the present disclosure shall be subject to the scope of protection of the claims. Industrial Applicability: The solution provided by the embodiments of the present disclosure first obtains server operating status time series data; then segments the operating status time series data into periods and stacks the resulting sub-series data; then extracts multidimensional time series features from the stacked data; and uses these multidimensional time series features to predict server failures. In this embodiment, fault prediction based on the multidimensional time series features extracted from the operating status time series data can achieve early prediction of server failures.Furthermore, by segmenting and stacking the operating status time series data by period and then extracting multidimensional time series features, we can extract information about the data's fluctuation trends and the different performance of the data in the previous and next periods, thereby improving the accuracy and recall of fault prediction. This achieves the technical effect of predicting server failures in advance, solving the technical problem of being unable to predict server failures in advance.

Claims

Claims 1. A server failure prediction method, comprising: Get the server's running status time series data; The operation status time series data is cut according to periods, and a plurality of sub-time series data obtained by cutting are stacked; time series multi-dimensional features are extracted from the stacked data; and fault prediction is performed on the server using the time series multi-dimensional features.

2. The method according to claim 1, wherein: The step of cutting the running status time series data according to cycles and stacking the multiple sub-time series data obtained by cutting includes: cutting the running status time series data according to the same cycle length and stacking the sub-time series data of the same cycle length obtained by cutting.

3. The method according to claim 1, wherein: The step of cutting the running status time series data according to cycles and stacking the multiple sub-time series data obtained by cutting includes: cutting the running status time series data according to different cycle lengths and stacking the sub-time series data with different cycle lengths obtained by cutting.

4. The method according to claim 2 or 3, wherein: The extracting the time series multidimensional features from the stacked data includes: constructing a multidimensional tensor based on the stacked data; and performing a convolution operation on the multidimensional tensor to obtain the time series multidimensional features.

5. The method according to claim 4, wherein: The constructing a multidimensional tensor based on the stacked data includes: constructing a multidimensional tensor based on the collection time and data attributes of the running status time series data and the stacked data.

6. The method according to claim 4, wherein: The step of performing a convolution operation on the multidimensional tensor to obtain a time series multidimensional feature includes: performing a convolution operation on the multidimensional tensor to obtain a feature corresponding to each tensor in the multidimensional tensor; and concatenating the features corresponding to each tensor to obtain the time series multidimensional feature.

7. The method according to claim 1, wherein: After extracting the time series multidimensional features from the stacked data, the method further includes: acquiring inherent attribute data of the server, wherein the inherent attribute data includes at least one of a machine model and a component model of the server; and constructing discrete features based on the inherent attribute data.

8. The method according to claim 7, wherein: The using the time series multidimensional feature to predict the fault of the server includes: splicing the discrete feature and the time series multidimensional feature to obtain a spliced ​​feature; and using the spliced ​​feature to predict the fault of the server.

9. The method according to any one of claims 1 to 8, wherein: The using the time series multi-dimensional features to predict faults of the server includes: inputting the time series multi-dimensional features into a machine learning model, and using the output of the machine learning model as a fault prediction result; the machine learning model includes a classification model or a fault detection model.

10. A server failure prediction device, comprising: A data acquisition component, configured to acquire the server's running status time series data; a data stacking component, configured to cut the running status time series data into periods, and stack the multiple sub-time series data obtained by cutting; A feature extraction component is configured to extract time series multi-dimensional features from the stacked data; A fault prediction component is configured to use the time series multi-dimensional features to perform fault prediction on the server.

11. A computer program product comprising: A computer program, wherein when the computer program is executed by a processor, the method according to any one of claims 1 to 9 is implemented.

12. A computer program product comprising: A non-volatile computer-readable storage medium storing a computer program, wherein the computer program implements the method according to any one of claims 1 to 9 when executed by a processor.

13. A computer program, wherein When the computer program is executed by a processor, the method according to any one of claims 1 to 9 is implemented.

14. An electronic device, characterized in that: The invention comprises a memory, a processor and a computer program stored in the memory, wherein the processor implements the method according to any one of claims 1 to 9 when executing the computer program.

15. A computer-readable storage medium, wherein a computer program is stored in the computer-readable storage medium, and when the computer program is executed by a processor, the method according to any one of claims 1 to 9 is implemented.

Citation Information

Patent Citations

  • Abnormal data detection method and device, electronic equipment and storage medium

    CN114722937A

  • Fault detection method and device for database instance

    CN115509784A

  • Computing system fault prediction method and system based on component call analysis

    CN115509789A

  • Method and device for extracting characteristic value of time series data

    CN116802616A

  • Stateful detection of anomalous events in virtual machines

    US20160371170A1

Cited By

  • Hardware fault information management method and device, electronic equipment and storage medium

    CN120540925A

  • Intelligent kitchen electrical equipment data analysis method and system based on Internet of Things cloud platform

    CN121144763A

  • Turbine controller edge side data cleaning and anomaly detection method for industrial internet

    CN122064074A