Fault feature extraction method, electronic equipment and storage medium

By obtaining the current running data of the server, calculating the importance score of the feature value and dynamically adjusting the fault feature set, the problem of failure characteristics cannot be captured in traditional server management is solved, and the accuracy and reliability of fault prediction are improved.

CN120336990AActive Publication Date: 2025-07-18INSPUR SUZHOU INTELLIGENT TECH CO LTD

Patent Information

Application Number
CN202510832049.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-20
Publication Date
2025-07-18
Estimated Expiration
2045-06-20

AI Technical Summary

Technical Problem

Traditional server management methods cannot fully capture the feature information related to failure, resulting in limited prediction capabilities of the fault prediction model.

Method used

By obtaining the server's current running data, extracting the fault feature set, and calculating the importance score of the feature values in the current running data, dynamically adjusting the feature vectors in the fault feature set to form the updated fault feature set for failure detection.

Benefits of technology

It significantly enhances the accuracy and reliability of fault prediction, can capture critical information related to faults in real time, and improves the performance of fault prediction models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336990A_ABST
    Figure CN120336990A_ABST
Patent Text Reader

Abstract

The invention discloses a fault feature extraction method, electronic equipment and a storage medium, and relates to the technical field of servers, and the method comprises the steps: obtaining the current operation data of a server; the current operation data comprises a multi-dimensional current characteristic value; extracting a fault feature set from the current operation data, and calculating an importance score of a current feature value of one dimension in the current operation data; the fault feature set comprises a plurality of first feature vectors; adjusting a single first feature vector in the fault feature set according to an importance score of a current feature value of one dimension in the current operation data to obtain an updated fault feature set; and the updated fault feature set is used for fault detection, so that the problem that the prediction capability of a fault prediction model is limited due to the fact that feature information related to faults cannot be fully captured in related technologies is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of server management, and particularly to a method for extracting fault features, an electronic device, and a storage medium. Background Art

[0002] Server management methods refer to a series of technologies and processes for monitoring, maintaining, and optimizing server performance, aiming to ensure the high availability, stability, and security of the server. In the process of feature extraction by traditional server management methods, it is impossible to fully capture the feature information related to faults, resulting in limited prediction ability of the fault prediction model. Summary of the Invention

[0003] This application provides a method for extracting fault features, an electronic device, and a storage medium, so as to at least solve the problem in the related art that the feature information related to faults cannot be fully captured, resulting in limited prediction ability of the fault prediction model.

[0004] This application provides a method for extracting fault features, including: obtaining the current operation data of the server; the current operation data includes multi-dimensional current feature values; extracting a fault feature set from the current operation data, and calculating the importance score of the current feature value of one dimension in the current operation data; the fault feature set includes multiple first feature vectors; adjusting the single first feature vector in the fault feature set according to the importance score of the current feature value of one dimension in the current operation data to obtain the updated fault feature set; the updated fault feature set is used for fault detection.

[0005] This application also provides a device for extracting fault features, including: a collection module, configured to obtain the current operation data of the server; the current operation data includes multi-dimensional current feature values; a feature extraction module, configured to extract a fault feature set from the current operation data, and calculate the importance score of the current feature value of one dimension in the current operation data; the fault feature set includes multiple first feature vectors; a feature update module, configured to adjust the single first feature vector in the fault feature set according to the importance score of the current feature value of one dimension in the current operation data to obtain the updated fault feature set; the updated fault feature set is used for fault detection.

[0006] This application also provides an electronic device, including: a memory, configured to store a computer program; a processor, configured to implement the steps of any of the above methods for extracting fault features when executing the computer program.

[0007] This application also provides a computer-readable storage medium, in which a computer program is stored, and the computer program, when executed by a processor, implements the steps of any of the above methods for extracting fault features.

[0008] The present application also provides a computer program product, including a computer program which, when executed by a processor, implements the steps of any of the above-mentioned fault feature extraction methods.

[0009] Through the present application, a fault feature set is extracted from the currently running data of the server obtained in real time, and the importance score of the current feature value of a dimension in the currently running data is also calculated. The importance score reflects the actual influence of the feature on fault prediction at a specific time point. Different from the previous static feature weights, the importance score can be adaptively adjusted according to different running states and scenarios, ensuring the dynamics and flexibility of the feature set. According to the importance score of the current feature value of a dimension in the currently running data, the single first feature vectors in the fault feature set are adjusted in real time and intelligently to obtain an updated fault feature set. Among them, the updated fault feature set is optimized according to its importance in the current data state. Therefore, when the fault prediction model performs fault detection based on the updated fault feature set, it can capture the key information related to faults in real time, significantly enhancing the accuracy and reliability of fault prediction, and solving the problem in the related art that the feature information related to faults cannot be fully captured, resulting in limited prediction ability of the fault prediction model. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] In order to more clearly illustrate the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0011] Figure 1 FIG. is an application scenario diagram of a fault feature extraction method provided by an embodiment of the present application.

[0012] Figure 2 FIG. is a flowchart of a fault feature extraction method provided by an embodiment of the present application.

[0013] Figure 3 FIG. is a flowchart of a server fault handling method provided by an embodiment of the present application.

[0014] Figure 4 FIG. is a structural diagram of a server management system provided by an embodiment of the present application.

[0015] Figure 5 FIG. is a structural diagram of a fault feature extraction device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0016] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without making creative efforts belong to the protection scope of the present application.

[0017] It should be noted that in the description of the present application, the terms "including", "comprising" or any other variation thereof are intended to cover a non-exclusive inclusion, such that a process, method, article or device including a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present application are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0018] In order to enable those skilled in the art of this technology to better understand the solution of the present application, the present application will be further described in detail below in conjunction with the accompanying drawings and specific embodiments.

[0019] According to one aspect of the embodiments of the present application, a method for extracting fault features is provided. Optionally, in this embodiment, the above method for extracting fault features can be but is not limited to being applied to a hardware environment including a terminal device 102 and a server 104 as Figure 1 shown. The server 104 can be connected to the terminal device 102 through a network, and can be used to provide services (such as application services, etc.) for the terminal device 102 or the client installed on the terminal device 102. A database can be set on the server 104 or independently of the server 104 to provide data storage services for the server 104.

[0020] The above network can include but is not limited to at least one of the following: a wired network, a wireless network. The above wired network can include but is not limited to at least one of the following: a wide area network, a metropolitan area network, a local area network. The above wireless network can include but is not limited to at least one of the following: Wireless Fidelity (WIFI for short), Bluetooth. The terminal device 102 can be but is not limited to a personal computer (PC for short), a mobile phone, a tablet computer, etc. The server 104 can be but is not limited to a cloud server, a server or other server types.

[0021] The fault feature extraction method of the embodiments of the present application can be executed by the server 104, or by the terminal device 102, or jointly by the server 104 and the terminal device 102. Among them, when the terminal device 102 executes the fault feature extraction method of the embodiments of the present application, it can also be executed by the client installed thereon.

[0022] Taking the execution of the fault feature extraction method in this embodiment by the server 104 as an example, Figure 2 is a schematic flowchart of an optional fault feature extraction method according to the embodiments of the present application, as Figure 2 shown, the process of this method can include the following steps:

[0023] Step S202, obtain the current running data of the server; the current running data includes multi-dimensional current feature values.

[0024] The fault feature extraction method provided in this embodiment can be applied in the field of electronic device fault detection. Among them, the electronic device can be a terminal device, a server, or other electronic devices.

[0025] Among them, the current running data refers to the real-time data set generated by the server at a specific time point and reflecting its running state. For example, the current running data includes, but is not limited to, multi-dimensional performance indicators such as the usage rate of the central processing unit (CPU), memory occupancy, disk input / output (I / O), and network traffic.

[0026] The multi-dimensional current feature values refer to the specific values of the performance indicators in the current running data. Each indicator corresponds to a dimension, and the feature value is the actual observed value in that dimension. For example, the CPU usage rate is 75%, and the memory occupancy is 3GB. These are the feature values of the current running data in different dimensions.

[0027] Step S204, extract a fault feature set from the current running data, and calculate the importance score of the current feature value of one dimension in the current running data.

[0028] Among them, the fault feature set refers to the data features or index combinations extracted from the current running data of the server and closely related to the possibility of a fault occurring. The fault feature set is identified and selected through statistical methods and machine learning algorithms for constructing a fault prediction model.

[0029] The current eigenvalue of a dimension refers to the real-time observed value of a specific performance metric (such as CPU usage rate) in the current running data, which is a specific slice in multi-dimensional data. The importance score of the current eigenvalue of a dimension represents the contribution degree of the eigenvalue of a specific dimension in the current running data to the output result of the fault prediction model. The higher the importance score, the more significant the role of this feature in fault prediction. The importance score of the current eigenvalue of a dimension in the current running data can be obtained through a machine learning model or by measuring the correlation between features through multivariate statistical analysis. For example, assume that a machine learning model (such as a neural network) has been trained for server fault prediction. During the server fault prediction training phase, record the number of times the current eigenvalue of a dimension in the current running data is used to update the weights and the change trend of the gradient magnitude generated during the gradient descent process. Calculate the total number of times the current eigenvalue of a dimension in the current running data is cited throughout the training process, which reflects the activity degree of the current eigenvalue of a dimension in the current running data in the model decision-making process. Analyze the magnitude of the gradient update accompanied by the citation of the current eigenvalue of a dimension in the current running data. A larger gradient update means that the current eigenvalue of a dimension in the current running data has a more significant impact on the change of the server fault prediction output, so it has a higher importance score. For example, a multivariate statistical correlation matrix can also be constructed based on the server running data to reflect the linear and non-linear correlation degrees between different eigenvalues; use various statistical metrics, such as Pearson correlation coefficient, Spearman rank correlation coefficient, etc., to analyze the correlation matrix and identify the feature set highly correlated with the fault event; use statistical tools such as conditional entropy, standard error in regression analysis, etc., to quantify the association strength between the key factors and the fault event; based on the above analysis results, establish a scoring system to quantify the importance of the current eigenvalue of a dimension in the current running data.

[0030] Step S206: Adjust the single first feature vector in the fault feature set according to the importance score of the current eigenvalue of a dimension in the current running data to obtain an updated fault feature set; the updated fault feature set is used for fault detection.

[0031] Among them, the process of adjusting the single first feature vector in the fault feature set refers to the process of modifying or optimizing a specific feature in the fault feature set based on the importance score. For example, a fusion strategy can be designed, such as using the importance score of the current feature value of a dimension in the current running data as a weight, and fusing the current feature value of a dimension in the current running data with other relevant features according to this weight to update the fault feature set. For example, a scoring threshold can also be set to reduce the dimension of the single first feature vector in the fault feature set whose importance score is less than or equal to the scoring threshold. For example, a complex system log feature can be simplified to the count of key events, or multiple low-scoring features can be merged into one comprehensive feature; perform feature enhancement processing on the single first feature vector in the fault feature set whose importance score is greater than the scoring threshold, and at the same time fuse the low-scoring features after dimensionality reduction and the high-scoring features after enhancement to obtain the reconstructed high-scoring features, and determine the reconstructed high-scoring features and the low-scoring features after dimensionality reduction as the updated fault feature set. The updated fault feature set more accurately reflects the key information of the server running state, is used to optimize the fault prediction model, so that it performs better in fault prediction and reduces false alarms and missed alarms.

[0032] Through the embodiments of the present application, a fault feature set is extracted from the current running data of the server obtained in real time, and the importance score of the current feature value of a dimension in the current running data is also calculated. The importance score reflects the actual influence of the feature on fault prediction at a specific time point. Different from the previous static feature weights, the importance score can be adaptively adjusted according to different running states and scenarios, ensuring the dynamics and flexibility of the feature set; according to the importance score of the current feature value of a dimension in the current running data, the single first feature vector in the fault feature set is adjusted in real time and intelligently to obtain the updated fault feature set. Among them, the updated fault feature set is optimized according to its importance in the current data state. Therefore, when the fault prediction model performs fault detection based on the updated fault feature set, it can capture the key information related to faults in real time, significantly enhancing the accuracy and reliability of fault prediction, and solving the problem that the fault prediction model has limited prediction ability due to the inability to fully capture the feature information related to faults in the related art.

[0033] In an exemplary embodiment, in the field of server management methods, related server management often relies on data collection methods such as timed polling or manual triggering, which is not only inefficient but also difficult to achieve real-time monitoring. Therefore, to solve the above problems, in this embodiment, the current running data of the server is obtained, including: collecting the current running data through a collection component pre-deployed in the server; the configuration parameters of the collection component are configured according to the number of nodes and network environment of the server cluster where the server is located.

[0034] Among them, the server can be one of the servers in a server cluster. The server cluster can be a Kafka cluster. The collection component refers to a software or hardware module pre-deployed in the server, and its main function is to collect the running data of the server in real time. For example, the collection component can be a Kafka producer. A Kafka producer is a client program responsible for publishing data to a specific topic in the Kafka cluster. Its main task is to collect the running data of the server in real time (such as CPU usage, memory occupancy, disk I / O, and network traffic), and then send the data to the specified Topic (a topic is a unit for classifying or storing messages) in the Kafka cluster. The configuration parameters of the collection component refer to the setting values used to customize the running characteristics of the collection component, and are mainly adjusted according to the number of nodes and network environment of the server cluster where the server is located. For example, the configuration parameters of the collection component include the number of brokers (message brokers) in the Kafka cluster and the number of Topic partitions.

[0035] Through this embodiment, by pre-deploying the collection component on the server, the running data of the server can be continuously collected, rather than relying on periodic task scheduling or manual triggering, which improves the collection efficiency and real-time performance; the configuration parameters of the collection component will be automatically adjusted according to the number of cluster nodes and network environment where the server is located, and the configuration parameters of the collection component can be flexibly configured to adapt to server clusters of different scales and maintain high-efficiency data transmission performance in a high-concurrency environment, which not only improves the reliability and real-time performance of data collection, but also provides a solid foundation for subsequent data processing and analysis.

[0036] In an exemplary embodiment, before extracting the fault feature set from the current running data, the method further includes:

[0037] First, identify, remove, and fill in the first abnormal data points in the current running data to obtain the cleaned current running data; the first abnormal data points refer to the data points whose feature values are not within the preset range.

[0038] Second, perform format conversion on the cleaned current running data to obtain the current running data with a unified format.

[0039] Third, perform verification on the current running data with a unified format, identify and correct the second abnormal data points to obtain the processed current running data; the second abnormal data points refer to the data points that do not conform to the historical behavior pattern or trend.

[0040] Among them, the first abnormal data point specifically refers to a data point in the server operation data whose characteristic values (such as CPU usage rate, memory occupancy, disk I / O, network traffic, etc.) exceed the preset normal range. By identifying and removing the first abnormal data point and filling it using an appropriate method (such as interpolation method), more accurate and reliable server operation data can be obtained.

[0041] The second abnormal data point refers to a data point in the current operation data after cleaning whose numerical value or behavior pattern seriously does not conform to the normal behavior or trend in the historical data. The second abnormal data point is usually identified through machine learning algorithms (such as time series analysis, clustering analysis, etc.). The machine learning algorithm can learn and understand the behavior pattern of the historical operation data. When the current operation data deviates from the corresponding behavior pattern, it is considered as the second abnormal data point. Correcting the second abnormal data point helps to further improve the data quality, reduce the noise interference in the training and prediction process of the fault prediction model, and thus improve the performance and reliability of the fault prediction model.

[0042] Optionally, the server initially cleans the current operation data of the server based on a statistical outlier detection algorithm, identifies and removes the first abnormal data points that significantly deviate from the normal range to obtain the current operation data after initial cleaning; applies the interpolation method to fill a small number of suspected first abnormal data points that still exist after cleaning to obtain the current operation data after final cleaning; converts the current operation data after final cleaning into the current operation data in a unified format according to the predefined data format standard; uses a machine learning algorithm to perform a secondary verification on the current operation data in the unified format, identifies and corrects the missing second abnormal data points, and obtains high-quality server operation data.

[0043] Among them, using a machine learning algorithm to perform a secondary verification on the current operation data in the unified format, identifying and correcting the missing second abnormal data points includes:

[0044] Perform feature selection and standardization processing on the current running data in a unified format, extract key performance indicators including CPU usage, memory occupancy, disk I / O, and network traffic, and construct a feature vector for anomaly detection; build an anomaly detection model based on the training dataset, and the machine learning algorithms used include one or more combinations of Isolation Forest, Local Outlier Factor, Autoencoder, or Support Vector Machine; input the current running data in a unified format into the anomaly detection model, and identify potential second anomaly data points through the anomaly score or classification result output by the anomaly detection model; for the identified second anomaly data points, use interpolation method for correction through the prediction model combined with time series context information, where the prediction model includes but is not limited to time series prediction models such as ARIMA or LSTM; finally, verify the corrected current running data to determine whether the corrected current running data meets the expectations. If there are still anomalies, return to adjust the parameters of the anomaly detection model or the correction strategy to achieve closed-loop processing and continuous optimization of the second anomaly data, and obtain high-quality server running data.

[0045] Through this embodiment, the first anomaly data points with feature values exceeding the preset range are removed, and interpolation method is used to fill in the missing values to ensure data integrity; subsequently, the cleaned data is converted into a unified format to enhance cross-system compatibility; finally, through in-depth verification by machine learning algorithms, the second anomaly data points that do not conform to historical behaviors are corrected, significantly improving the data quality. By performing detailed cleaning and standardization processing on the current running data of the server, the influence of noise and error information can be effectively reduced, the data quality can be improved, thereby enhancing the robustness and prediction accuracy of the fault prediction model, and effectively solving the problems of low data collection efficiency, difficult real-time monitoring, and inaccurate analysis of the impact of data anomalies in related server management.

[0046] In an exemplary embodiment, a fault feature set is extracted from the current running data, including:

[0047] I. Calculate the covariance matrix corresponding to the current running data; the covariance matrix characterizes the linear correlation between different indicators in the current running data.

[0048] Among them, the covariance matrix is a mathematical tool used to characterize the linear correlation between different indicators in the current running data of the server. The elements in the covariance matrix are covariance values, and the covariance value describes the strength and direction of the linear relationship between different variables in the current running data in pairs. Specifically, if the covariance matrix is constructed for the feature values of m dimensions ({x1, x2,..., x m}) in the current running data, then the element in the i-th row and j-th column of the covariance matrix is the variable x i and x jThe covariance between them. The larger the covariance value, the stronger the linear dependence between the two indicators.

[0049] Second, perform eigenvalue decomposition on the covariance matrix to obtain multiple principal component eigenvalues and multiple second eigenvectors corresponding to the multiple principal component eigenvalues.

[0050] Among them, eigenvalue decomposition is a linear algebra method that decomposes the covariance matrix into principal component eigenvalues and corresponding second eigenvectors. The principal component eigenvalues are a series of values obtained after eigenvalue decomposition of the covariance matrix of the current running data of the server. Each principal component eigenvalue corresponds to a second eigenvector of the covariance matrix, and the magnitude of the principal component eigenvalue reflects the amount of data variance in the direction of the corresponding second eigenvector. The second eigenvector refers to the eigenvector corresponding to the principal component eigenvalue obtained through eigenvalue decomposition. Each principal component eigenvalue corresponds to a second eigenvector, and the second eigenvectors form a new orthogonal coordinate system, where the direction of each second eigenvector represents the main change trend of the data in that direction. The second eigenvector, as the direction of the principal component, can indicate the weights of each dimension (such as CPU usage, memory usage, network traffic, disk I / O) in the current running data of the server in the new coordinate system.

[0051] Third, arrange the multiple second eigenvectors in descending order of the principal component eigenvalues, start selecting the second eigenvectors from the largest principal component eigenvalue, and when the selected second eigenvectors meet the preset conditions, determine the selected second eigenvectors as multiple first eigenvectors to obtain a fault feature set; the preset conditions refer to the condition that the cumulative contribution rate of the selected second eigenvectors is greater than the preset contribution rate threshold.

[0052] Among them, the cumulative contribution rate refers to the sum of the proportions of the total variance explained by the selected second eigenvectors in the total variance of the current running data of the server in principal component analysis (PCA) or factor analysis. The fault feature set in this embodiment refers to a data set obtained by selecting a group of second eigenvectors from the eigenvalue decomposition result of the covariance matrix according to the condition that the cumulative contribution rate is greater than the preset contribution rate threshold. The cumulative contribution rate is used to measure the importance of the second eigenvectors. The larger the principal component eigenvalue, the stronger the explanatory ability of the corresponding second eigenvector for the data.

[0053] Optionally, based on the current running data of the high-quality server, calculate the covariance matrix corresponding to the current running data of the server, perform eigenvalue decomposition on the covariance matrix to obtain multiple principal component eigenvalues and corresponding second eigenvectors; arrange the multiple second eigenvectors in descending order of the principal component eigenvalues, start selecting k second eigenvectors from the largest principal component eigenvalue, where k starts from 1, and calculate the cumulative contribution rate of the selected k second eigenvectors each time until the cumulative contribution rate of the selected k second eigenvectors is greater than the preset contribution rate threshold. Take the selected k second eigenvectors as k first eigenvectors, and these k first eigenvectors are also called k principal components. Determine the selected second eigenvectors as the fault feature set. Among them, the cumulative contribution rate of the k second eigenvectors can be calculated according to the following formula (1):

[0054]

[0055] Among them, is the cumulative contribution rate of the k second eigenvectors, is the principal component eigenvalue corresponding to the j-th second eigenvector selected, is the sum of all principal component eigenvalues obtained after the covariance matrix decomposition, represents the total number of principal component eigenvalues obtained after the covariance matrix decomposition.

[0056] For example, assume that eigenvalue decomposition is performed on the covariance matrix to obtain four principal component eigenvalues [2.4, 0.5, 0.08, 0.02] and a 4x4 second eigenvector matrix. Each column in the second eigenvector matrix represents a second eigenvector, which corresponds one-to-one with the principal component eigenvalues. Assume that the second eigenvector corresponding to the principal component eigenvalue 2.4 is [0.6, 0.7, 0.1, 0.3], which means that the weights of the four features of CPU usage rate, memory usage, network traffic, and disk I / O in the main change direction of the data are 0.6, 0.7, 0.1, and 0.3 respectively. After arranging the multiple second eigenvectors in descending order of the principal component eigenvalues, the cumulative contribution rate of the selected 2 second eigenvectors is 95%, and the preset contribution rate threshold is 90%. Then, take the selected 2 second eigenvectors as the fault feature set.

[0057] Through this embodiment, the linear correlation between different metrics in the current operating data of the server is analyzed in real time using the covariance matrix, so as to quickly capture the potential associations between metrics, improving the sensitivity of fault prediction; eigenvalue decomposition is performed on the covariance matrix to extract key information from the covariance matrix, and eigenvalues and corresponding eigenvectors representing the direction and intensity of data change can be obtained. Then, the eigenvectors are sorted according to the eigenvalue magnitudes, and a fault feature set is selected based on the principle that the cumulative contribution rate is greater than the preset contribution rate threshold, avoiding the interference of redundant information and ensuring that the fault prediction model can focus on the most important data dimensions, thereby improving the accuracy of the fault prediction model.

[0058] In an exemplary embodiment, calculating the importance score of the current eigenvalue of a dimension in the current operating data includes: obtaining the historical operating data of the server; the historical operating data includes historical eigenvalues of multiple dimensions; determining the historical mean and standard deviation of the eigenvalue of a dimension in the current operating data according to the historical operating data; and determining the importance score of the current eigenvalue of a dimension in the current operating data according to the current eigenvalue, historical mean, standard deviation, and first preset weight of a dimension in the current operating data.

[0059] Among them, the historical mean of a dimension in the current operating data reflects the average operating state of the server over a long period of time, and it is an estimate of the central tendency of the server performance metrics under normal operating conditions in the past. When calculating the importance score of the current eigenvalue, the historical mean serves as a reference benchmark for measuring whether the current eigenvalue deviates from the normal working level over a long period of time.

[0060] The standard deviation is a statistical indicator that measures the fluctuation range of the eigenvalues of this dimension in the historical operating data, showing the average degree of difference between the data points in the historical operating data and the historical mean. The larger the standard deviation, the higher the variability of the eigenvalue, which may be more sensitive to the health state of the server and also means a higher potential risk. When determining the importance score, the standard deviation can be used to determine the degree of abnormality of the current eigenvalue, that is, how severely it deviates from the historical mean.

[0061] The first preset weight is preset according to the occurrence frequency and impact degree of the current eigenvalue in historical fault cases, and it assigns an importance level to each eigenvalue in fault warning.

[0062] For example, the historical eigenvalues of 4 dimensions (CPU usage rate, memory usage, network traffic, disk I / O) at different time points (such as 10 time points) of the server are shown in Table 1:

[0063] Table 1

[0064]

[0065] Based on the eigenvalue of each column in Table 1, calculate the historical mean and standard deviation of the current eigenvalue of each dimension in the current running data. For example, for the dimension of CPU usage, calculate the historical mean and standard deviation of the eigenvalue of the CPU usage dimension based on the historical eigenvalues at 10 time points of the CPU usage.

[0066] In one embodiment, the importance score of the current eigenvalue of a dimension in the current running data can be calculated according to the following formula (2):

[0067]

[0068] Where, is the importance score of the current eigenvalue of a dimension in the current running data; is the current eigenvalue of a dimension in the current running data; is the historical mean of the current eigenvalue of a dimension in the current running data; is the standard deviation of the current eigenvalue of a dimension in the current running data; is the first preset weight.

[0069] Through this embodiment, when calculating the importance score of the eigenvalue in the server running data, the current eigenvalue, historical mean, standard deviation, and the first preset weight are combined. Among them, the historical mean and standard deviation represent the running mode and data fluctuation range within a long time. By comparing the current eigenvalue with the historical mean and standard deviation, the deviation degree of the current state from the normal range can be evaluated, so as to determine the normal fluctuation and potential fault signals in the current running data of the server, and the importance score is dynamically adjusted in combination with the first preset weight, so as to allocate a reasonable importance score to the current eigenvalue of a dimension in the current running data, improving the accuracy of the importance score of the current eigenvalue of a dimension in the current running data.

[0070] In an exemplary embodiment, adjust the single first eigenvector in the fault feature set according to the importance score of the current eigenvalue of a dimension in the current running data to obtain an updated fault feature set, including:

[0071] First, perform a weighted sum of the current eigenvalue of a dimension in the current running data according to the importance score of the current eigenvalue of a dimension in the current running data to obtain the first contribution degree corresponding to the current running data.

[0072] Among them, as can be seen from the above embodiments, the multiple first feature vectors are the feature vectors extracted after performing principal component analysis on the current running data; each single first feature vector among the multiple first feature vectors corresponds to a principal component eigenvalue. The calculation processes of the first feature vectors and the principal component eigenvalues are not elaborated here again.

[0073] The first contribution degree is the degree of influence of the current eigenvalue of a dimension in the current running data on the overall data set calculated based on the current eigenvalue of a dimension in the current running data and its importance score, and is used to quantify the direct impact of the eigenvalue in the current running data on the server health status. In this embodiment, during the process of calculating the first contribution degree corresponding to the current running data, the importance score of the current eigenvalue of a dimension in the current running data is used as an adaptive weight to quantify the sensitivity and influence of the current eigenvalue of this dimension on the server state change. Specifically, the importance score measures the direct contribution degree of the current eigenvalue of a specific dimension (such as CPU usage rate, memory occupancy, etc.) in representing the server running state to predicting potential failures at a given time point.

[0074] Optionally, the server calculates the product of the current eigenvalue of each dimension in the current running data and its importance score, and then determines the sum of the products of the current eigenvalue of each dimension in the current running data and its importance score as the first contribution degree corresponding to the current running data.

[0075] Second, calculate the product of each single first feature vector in the fault feature set and the corresponding principal component eigenvalue, and determine the product result as the second contribution degree of the single first feature vector in the fault feature set.

[0076] Among them, as can be seen from the above embodiments, the fault feature set includes k first feature vectors extracted through principal component analysis, where the k first feature vectors are k second feature vectors selected from the covariance matrix, and each principal component eigenvalue corresponds to a second feature vector. In this embodiment, the second contribution degree refers to the product of each single first feature vector in the fault feature set and its corresponding principal component eigenvalue, reflecting the contribution degree of the single first feature vector to fault prediction in the fault feature set.

[0077] Third, determine the sum of the second contribution degree of the single first feature vector in the fault feature set and the first contribution degree corresponding to the current running data as the updated single first feature vector in the fault feature set, and obtain the updated fault feature set.

[0078] Among them, the updated single first feature vector in the fault feature set can be expressed as shown in the following formula (3):

[0079]

[0080] Among them, denotes the updated eigenvector of the j-th first eigenvector, denotes the principal component eigenvalue corresponding to the j-th first eigenvector, denotes the th first eigenvector, denotes the second contribution degree of the th first eigenvector, denotes the total number of current eigenvalues in the current operating parameters of the server, and f denotes the f-th feature in the current operating parameters of the server; is the importance score of the current eigenvalue of a dimension in the current operating data; is the current eigenvalue of a dimension in the current operating data; denotes the first contribution degree corresponding to the current operating data.

[0081] For example, assume that the total number of current eigenvalues in the current operating parameters of the server is 4, that is, m is 4, and assume the value of is 8. Perform eigenvalue decomposition on the covariance matrix to obtain four principal component eigenvalues [2.4, 0.5, 0.08, 0.02] and a 4x4 second eigenvector matrix. Each column in the second eigenvector matrix represents a second eigenvector, which corresponds one-to-one with the principal component eigenvalues. The fault feature set includes the second eigenvectors corresponding to the principal component eigenvalues [2.4, 0.5]. Among them, the second eigenvector corresponding to the principal component eigenvalue 2.4 is [0.6, 0.7, 0.1, 0.3]. Then, the updated eigenvector of the second eigenvector [0.6, 0.7, 0.1, 0.3] is represented as shown in formula (4) below:

[0082]

[0083] Through this embodiment, the sum of the second contribution degree of the single first eigenvector in the fault feature set and the first contribution degree corresponding to the current operating data is determined as the updated single first eigenvector in the fault feature set, and the updated fault feature set is obtained, so that the updated fault feature set integrates real-time data and historical data analysis results, ensuring that the fault prediction model not only has robustness based on historical data, but also can timely reflect changes in recent operation trends, thereby making more accurate and timely fault warnings.

[0084] In an exemplary embodiment, the related fault warning system has problems such as warning delay or frequent false alarms, which affect the stable operation of the server. To solve this problem, the above fault feature extraction method further includes the following steps:

[0085] 1. Input the updated fault feature set into a pre-trained fault prediction model to obtain a prediction value, which represents the likelihood of the server experiencing a fault.

[0086] Among them, the fault prediction model is constructed based on a machine learning algorithm. It is used to output a prediction value of the server experiencing a fault according to the input fault features. This prediction value is an estimated probability of the server potentially experiencing a fault output by the fault prediction model. For example, this prediction value can be a probability value of the server experiencing a fault, or a probability score of the server experiencing a fault, etc.

[0087] In some embodiments, the updated fault feature set can also be used to train the fault prediction model. For example, the fault prediction model selects the Random Forest (RF) algorithm. As a powerful ensemble learning method, Random Forest has good generalization ability and anti-overfitting ability. Input the updated fault feature set as the training set into the fault prediction model for training. The training process expression is as shown in the following formula (5):

[0088]

[0089] Among them, represents the trained fault prediction model, is the training set.

[0090] After obtaining the trained fault prediction model, a new dataset to be predicted can be prepared, and a new set of fault features can be extracted according to the above fault feature extraction method . Input the fault features into the trained fault prediction model for prediction. The prediction process expression is as shown in the following formula (6):

[0091]

[0092] Among them, represents the prediction value output by the fault prediction model .

[0093] 2. Generate a warning message when the prediction value is greater than the warning threshold; the warning threshold is determined according to the current operation data; the warning message includes response operation process information, and the response operation process information includes multiple operation steps.

[0094] Among them, the warning threshold is a standard value used to determine whether the operating state of the server is approaching or reaching the critical point of possible failure. The warning threshold is dynamically generated based on the current operating data. By setting a reasonable warning threshold, the failure risk can be accurately judged, the warning threshold can be dynamically adjusted, and combined with detailed warning information, potential problems can be discovered in time and corresponding measures can be taken, thereby minimizing the impact of failures on the system.

[0095] The warning information is a notification automatically generated by the server when the predicted value exceeds the warning threshold. For example, the warning information may include key information such as server identification, failure probability, recommended response operation steps and their execution order. The warning information serves as a bridge connecting the failure prediction and the automatic response mechanism, guiding the server or operation and maintenance personnel to quickly identify and respond to potential server failures.

[0096] The response operation process information is an integral part of the warning information, which defines a series of predefined operation steps that the server should execute when a failure risk is detected to restore the server from the warning state to the normal operating state. For example, a series of predefined operation steps include data backup, service migration, resource allocation adjustment, etc.

[0097] In an exemplary embodiment, when the predicted value is greater than the warning threshold, generating the warning information includes:

[0098] When the predicted value is greater than the warning threshold, determine the target warning level from multiple warning levels, where each warning level corresponds to a threshold range, and the threshold range is obtained based on real-time data analysis of the current operating state of the server; each warning level corresponds to a set of response operation process templates, and each response operation process template contains a series of operation steps, such as resource adjustment, service degradation, fault isolation, emergency shutdown, etc.; according to the fault type and the current operating data of the server, determine the target response operation process template from a set of response operation process templates corresponding to the target warning level, and adjust the target response operation process template according to the current operating data of the server, and generate a response operation process list according to the adjusted target response operation process template. The multiple operation steps in the response operation process list are sorted according to the execution priority, and each step will clearly indicate the required commands and expected effects; generate warning information according to the response operation process list.

[0099] Among them, the multiple warning levels are set in advance according to historical fault data. For example, collect historical fault data, including information such as fault type, fault impact, processing time, business loss, etc.; use clustering algorithms or decision tree models to divide the historical data of faults into three levels according to the historical data of faults: minor faults, medium faults, and severe faults.

[0100] When selecting a target response operation process template, extract the scenario feature vectors from each response operation process template in a group of response operation process templates corresponding to the target warning level. Vectorize the fault type and the current operation data of the server to obtain an operation state vector. Match the operation state vector with the scenario feature vectors of each response operation process template, and select the response operation process template with the highest matching degree as the target response operation process template.

[0101] For example, if the CPU usage rate of the server exceeds the dynamically adjusted warning threshold, the system automatically identifies this fault as a medium fault (warning level II). Based on the response operation process template for medium faults, the generated warning information includes: fault description (such as abnormal increase in CPU usage rate), estimated impact (such as possible service response delay), and recommended response operation process (such as Step 1: Increase CPU and memory resource allocation; Step 2: Monitor the fault recovery situation, and if it is ineffective, execute Step 3: Start the standby server; Step 4: Back up key data; Step 5: Analyze the cause of high CPU load and perform targeted optimization).

[0102] Through this embodiment, compare the predicted value with the dynamically adjusted warning threshold, and combine the real-time operation state of the server to determine the target warning level. For each warning level, preset a group of response operation process templates containing multiple operation steps. By extracting the context feature vectors related to the fault type and current operation data and matching them with the scenario feature vectors of each response operation process template, it is possible to select the target response operation process template that best suits the current context, improve the pertinence and effectiveness of the response operation, and reduce the possibility of ineffective or excessive operations.

[0103] Third, issue a warning according to the warning information, and execute multiple operation steps according to the time sequence in the response operation process information.

[0104] Optionally, the server determines the warning threshold based on the current operation data, and according to the predicted value output by the fault prediction model judge whether there is a potential fault risk; when the prediction result exceeds the warning threshold, then generate the corresponding warning information. The expression for generating the warning information is shown in the following formula (7):

[0105]

[0106] Among them, indicates whether to issue a warning message, is the warning threshold; when = When it does, detailed warning information is generated; the warning information includes the server identifier, the probability of failure, the recommended operation steps, and the timestamp. The server executes multiple operation steps according to the execution sequence of each operation step in the warning information.

[0107] Through this embodiment, the warning threshold is dynamically generated according to the current running data, ensuring the timeliness and accuracy of the warning, and avoiding the problems of warning delay or frequent false alarms caused by a fixed threshold; after the warning information is triggered, the predefined response operation process is automatically executed, including a series of operation steps for fault recovery, without manual intervention, reducing the waiting time for fault handling, minimizing the impact of the fault on the system, and reducing the burden on the operation and maintenance personnel.

[0108] In an exemplary embodiment, after executing multiple operation steps according to the sequence in the response operation process information, the above-mentioned fault feature extraction method further includes:

[0109] Perform a performance evaluation on the server after executing multiple operation steps to obtain a fault response processing result; in the case where the fault response processing result is greater than a preset response value, stop the fault response.

[0110] In an exemplary embodiment, performing a performance evaluation on the server after executing multiple operation steps to obtain a fault response processing result includes:

[0111] Evaluate the preset performance indicators of the server after executing a single operation step among multiple operation steps, and perform fusion processing on the preset performance indicators of the server after executing a single operation step among multiple operation steps to obtain a fault response processing result.

[0112] Among them, the preset performance indicators refer to a series of key performance parameters defined to evaluate the health status of the server, such as CPU utilization rate, memory occupancy rate, disk I / O rate, network bandwidth usage rate, etc. The preset performance indicators provide a performance benchmark for the normal operation of the server and are used to measure the change in the server performance before and after executing the fault response operation.

[0113] A single operation step among multiple operation steps is a specific action in the response operation process information, such as restarting the service, adjusting resource allocation, performing data recovery, etc. Multiple operation steps are designed to directly solve or alleviate the server fault, and each operation step will have an immediate impact on the server performance.

[0114] The fault response processing result refers to the overall evaluation result of the server's preset performance metrics after the server has completed the operation steps in a series of response operation process information. The fault response processing result is obtained by comparing the performance metrics before and after the operation and integrating the change trends of multiple metrics, and is used to evaluate the overall effect of the fault response operation. If the fault response processing result is less than or equal to the preset response value, it means that the response operation fails to effectively improve the specified performance of the server to a non-fault state, and there may even be a performance decline. In this case, the operation steps need to be adjusted again. If the fault response processing result is greater than the preset response value, it means that the response operation effectively improves the specified performance of the server to a non-fault state, and in this case, there is no need to adjust the operation steps.

[0115] The preset response value is used to evaluate whether the fault response processing result has achieved the expected performance recovery or improvement goal. The preset response value is determined by analyzing historical fault cases and performance metrics in the normal operation state, combined with business requirements and the tolerance of the server performance, aiming to ensure that the performance of the server after fault recovery is not lower than or is better than the level before the fault.

[0116] Optionally, the server executes a single operation step according to the response operation process information in the warning information, such as restarting the server, adjusting resource allocation, etc., and monitors the server in real time after executing a single operation step, collects preset performance metrics, such as CPU utilization, memory occupancy, disk I / O, etc.; integrates and processes the performance metrics after executing multiple operation steps, and uses methods such as weighted average and maximum value selection to obtain a comprehensive fault response processing result; compares the obtained fault response processing result with the preset response value to judge whether the server performance has recovered to the expected level. If the fault response processing result is less than or equal to the preset response value, it indicates that the specified performance metric of the server is lower than the non-fault state, and further fault diagnosis or more advanced response measures may be required. If the fault response processing result is greater than the preset response value, it indicates that the specified performance metric of the server is higher than the non-fault state, and there is no need to replace the response measure at this time.

[0117] Through this embodiment, in the fault response, the performance of the server after executing each operation step is immediately evaluated, and the preset performance metrics of the server after executing multiple operation steps are integrated and processed to obtain a comprehensive fault response processing result, comprehensively considering the recovery of each performance parameter, avoiding the limitations of single-metric evaluation, and ensuring the comprehensiveness and high pertinence of the fault response strategy; when the fault response processing result reaches or exceeds the preset response value, it means that the server performance has been restored or improved to a non-fault state, and the subsequent fault response operation steps are automatically stopped, effectively reducing the excessive interference with the server operation.

[0118] In an exemplary embodiment, the preset performance metrics of the server after performing a single operation step among multiple operation steps are subjected to fusion processing to obtain a fault response processing result, including:

[0119] Calculate the difference between the first parameter value of the preset performance metrics of the server after performing a single operation step among multiple operation steps and the second parameter value of the preset performance metrics of the server after performing the previous operation step to obtain multiple differences; determine the ratio between the sum of the multiple differences and the total number of steps of the multiple operation steps as the fault response processing result.

[0120] Among them, the first parameter value specifically refers to the value of the preset performance metrics of the server measured and calculated immediately after performing a single operation step, reflecting the direct impact of the operation execution on the server performance. The second parameter value refers to the value of the preset performance metrics of the server recorded before performing the current operation step, that is, after the previous operation step is completed, and is used to provide a comparison baseline so that the server can evaluate whether the performance change brought by the current operation step is positive or negative. The difference between the first parameter value and the second parameter value reflects the change direction of the server performance metrics before and after the execution of the current operation step. A positive value indicates performance improvement, and a negative value indicates performance degradation.

[0121] In this embodiment, the fault response processing result is obtained by adding up all the differences obtained during the execution of multiple operation steps, and then using the ratio between the obtained total difference and the total number of operation steps as the final evaluation index. The fault response processing result can be expressed by the following formula (8):

[0122]

[0123] Among them, represents the fault response processing result, represents the first parameter value of the preset performance metrics after the -th step of operation, represents the second parameter value of the preset performance metrics after the -th step of operation, is the total number of operation steps.

[0124] Through this embodiment, after each operation step is executed, the change difference of the preset performance metrics of the server is calculated immediately. The first parameter value of the preset performance metrics after the execution of the current step is compared with the second parameter value of the preset performance metrics after the execution of the previous step to obtain a difference list. This dynamic evaluation mechanism allows real-time monitoring of the actual impact of each operation on the server performance during the fault recovery process, enabling timely identification of which operations contribute positively to the server performance improvement, and which operations are ineffective or have negative impacts, so as to quickly adjust the response strategy, avoid the continuation of ineffective operations, and improve the resource utilization efficiency and the fault recovery speed. Further, all the differences obtained after the execution of the operation steps are added up, and then the ratio is calculated with the total number of operation steps to finally obtain the fault response processing result. This fault response processing result reflects the average degree of the server performance improvement, facilitating subsequent judgment based on the fault response processing result whether the response operation is effective and whether the expected performance improvement degree is achieved, so as to decide whether to continue the subsequent operations or stop when the preset response value is reached, avoid over-intervention, ensure the server is restored to the best state while minimizing the interference to normal business.

[0125] In an exemplary embodiment, the above-mentioned fault feature extraction method further includes:

[0126] In the case where the fault response processing result is less than or equal to the preset response value, adjust the resource allocation ratios of multiple types of resources in the server to obtain a target resource configuration plan, and adjust the resource configuration of the server according to the target resource configuration plan.

[0127] Among them, in the case where the fault response processing result is less than or equal to the preset response value, the fault response processing result indicates that the specified performance of the server is less than the performance at the time of the server fault. In this case, the resource configuration of the server needs to be adjusted; in the case where the fault response processing result is greater than the preset response value, the fault response processing result indicates that the specified performance of the server is greater than the performance at the time of the server fault. In this case, the fault response can be ended. For example, assuming the preset response value is 0, if the fault response processing result > 0, it means that the server performance has been improved after executing multiple operation steps in the response operation process information; if the fault response processing result ≤ 0, it means that after executing multiple operation steps in the response operation process information, further investigation or other measures need to be taken, such as adjusting the resource allocation ratios of multiple types of resources in the server.

[0128] The resource allocation ratio refers to the proportion of various resources (such as CPU, memory, storage, network bandwidth, etc.) in the total resources of the server. When the result of the fault response processing fails to reach the preset response value, that is, it indicates that the current resource allocation status is insufficient to effectively handle faults or improve performance, the server will recalculate the resource allocation ratio according to the resource optimization algorithm, in order to expect to further improve the operating condition of the server through the reconfiguration of resources.

[0129] The target resource configuration plan is a detailed plan for re-planning the server resources after the adjustment of the resource allocation ratio. The target resource configuration plan includes the specific allocation ratio and adjustment strategy for each type of resource, ensuring that the server can operate in an optimal state.

[0130] Optionally, the server continuously monitors the result of the fault response processing. If it is found that the result of the fault response processing does not reach the preset response value, the resource optimization process is immediately started, that is, based on the latest server operation data, analyze where the resource bottleneck is, and identify the deficiencies in multiple types of resources such as CPU, memory, storage, and network; use resource dynamic adjustment strategies, such as prediction models based on historical data, load balancing algorithms, etc., to calculate the optimal allocation ratio of each resource, and form a resource allocation ratio adjustment plan; integrate the resource allocation ratio adjustment plan to formulate a target resource configuration plan, and the adjustment target and priority of each type of resource are clearly defined in the target resource configuration plan; according to the target resource configuration plan, automatically adjust the resource configuration of the server, including but not limited to adjusting the CPU and memory quotas, optimizing the storage allocation strategy, and regulating the network bandwidth allocation, etc. After the adjustment, the server re-evaluates the server performance to ensure that the resource optimization has indeed improved the result of the fault response processing, and loops until it reaches or exceeds the preset response value.

[0131] Through this embodiment, a mechanism for dynamically adjusting resource configuration according to the result of the fault response processing is introduced. If the result of the fault response processing is lower than or equal to the preset response value, that is, it indicates that the current resource allocation plan fails to effectively handle the fault, the server will automatically adjust the allocation ratio of multiple types of resources to generate a new target resource configuration plan. By immediately monitoring and analyzing the result of the fault response processing, the server can quickly identify the deficiencies in resource allocation and make timely adjustments, avoiding poor recovery effects or resource waste caused by unreasonable resource allocation, and solving the problems of poor flexibility in resource scheduling methods and lagging evaluation of fault response effects in the related technologies.

[0132] In an exemplary embodiment, adjusting the resource allocation ratio of multiple types of resources in the server to obtain a target resource configuration plan includes:

[0133] Determine the new resource allocation ratio of multiple types of resources in the server according to the ratio between the current utilization rate of a single type of resource in multiple types of resources and the maximum resource utilization rate in multiple types of resources, and the second preset weight of the single type of resource in multiple types of resources, and determine the set of the new resource allocation ratios of multiple types of resources in the server as the target resource configuration plan.

[0134] Among them, the current utilization rate of a single type of resource specifically refers to the ratio of the amount of a certain resource (such as CPU, memory, disk I / O, network bandwidth, etc.) being used on the server to the total amount of the resource at a certain moment. The maximum resource utilization rate refers to the highest ratio of the amount of all resources being used on the server to the total amount of the resource among all resources of the server.

[0135] The second preset weight is a numerical value preset according to the type and importance of the resource, reflecting the priority and influence of the resource in the overall resource configuration. When calculating the new resource allocation ratio, the second preset weight is used to assign appropriate adjustment strength and priority to different resources to ensure that the resource configuration plan takes into account both the immediate state and the long-term value and business requirements of the resources.

[0136] The new resource allocation ratio refers to the optimized value for guiding the future resource allocation of the server calculated by the server according to the resource dynamic adjustment strategy after evaluating the fault response processing result. In this embodiment, the new resource allocation ratio can be calculated according to the following formula (9):

[0137]

[0138] Among them, is the resource allocation ratio of the th type of resource, is the current utilization rate of the th type of resource, is the maximum resource utilization rate of all resource utilizations; is the second preset weight, which is determined according to historical data and current requirements.

[0139] In this embodiment, the target resource configuration plan is a comprehensive resource optimization plan formed by integrating all single-type resource optimization strategies based on the new resource allocation ratio. Figure 3 The following is a flowchart of a server fault handling provided by an embodiment of the present application. As Figure 3 shown, it includes:

[0140] Step S301: Collect the runtime data of each server node by using a distributed data collection framework to obtain the current runtime data of the server;

[0141] Step S302: Process the current runtime data by using data cleaning and formatting methods to obtain high-quality current runtime data;

[0142] Step S303: Based on the high-quality current operation data, extract key features using statistical methods and machine learning algorithms to obtain a fault feature set.

[0143] Step S304: Construct a fault prediction model based on a machine learning algorithm, input the fault feature set into the fault prediction model, and output a predicted value.

[0144] Step S305: Set a threshold based on the server operation data, compare the predicted value with the warning threshold, and generate a warning message when the predicted value is greater than the warning threshold.

[0145] Step S306: Based on the warning message, adopt an automated response mechanism to execute a predefined response operation process to obtain a fault response processing result.

[0146] Step S307: Based on the fault response processing result, adopt a resource dynamic adjustment strategy to optimize the server resource configuration to obtain a target resource configuration plan.

[0147] Figure 4 As shown in the structure diagram of a server management system provided by an embodiment of the present application, Figure 4 it includes a data acquisition module, a data cleaning module, a feature extraction module, a model construction module, a warning generation module, a response execution module, and a resource optimization module. Among them, the data acquisition module is used to implement the above-mentioned step S301, the data cleaning module is used to implement the above-mentioned step S302, the feature extraction module is used to implement the above-mentioned step S303, the model construction module is used to implement the above-mentioned step S304, the warning generation module is used to implement the above-mentioned step S305, the response execution module is used to implement the above-mentioned step S306, and the resource optimization module is used to implement the above-mentioned step S307.

[0148] Through this embodiment, by adjusting the resource allocation ratio according to the ratio between the current utilization rate of a single type of resource and the maximum resource utilization rate, it is possible to accurately identify resource bottlenecks, timely adjust resource allocation, avoid affecting the overall server performance due to overuse of a certain type of resource, and solve the problem in the related art that resource management ignores the instant comparison of resource utilization efficiency, resulting in uneven resource allocation or waste.

[0149] Through the above description of the embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. However, in many cases, the former is a better implementation method.

[0150] An embodiment of the present application also provides a fault feature extraction device, as Figure 5 shown, including:

[0151] The acquisition module 502 is configured to obtain the current operation data of the server; the current operation data includes multi-dimensional current feature values;

[0152] The feature extraction module 504 is configured to extract a fault feature set from the current operation data and calculate the importance score of the current feature value of one dimension in the current operation data; the fault feature set includes a plurality of first feature vectors;

[0153] The feature update module 506 is configured to adjust the single first feature vectors in the fault feature set according to the importance score of the current feature value of one dimension in the current operation data to obtain an updated fault feature set; the updated fault feature set is used for fault detection.

[0154] For the description of the features in the embodiments corresponding to the fault feature extraction device, reference may be made to the relevant descriptions of the embodiments corresponding to the fault feature extraction method, which will not be elaborated here one by one.

[0155] In an exemplary embodiment, the acquisition module 502 is further configured to acquire the current operation data through an acquisition component pre-deployed in the server; the configuration parameters of the acquisition component are configured according to the number of nodes and the network environment of the server cluster where the server is located.

[0156] In an exemplary embodiment, the feature extraction module 504 is further configured to identify, remove, and fill the first abnormal data points in the current operation data to obtain the cleaned current operation data; the first abnormal data points refer to the data points whose feature values are not within the preset range; perform format conversion on the cleaned current operation data to obtain the current operation data in a unified format; perform verification on the current operation data in the unified format to identify and correct the second abnormal data points to obtain the processed current operation data; the second abnormal data points refer to the data points that do not conform to the historical behavior pattern or trend.

[0157] In an exemplary embodiment, the feature extraction module 504 is further configured to calculate the covariance matrix corresponding to the current operation data; the covariance matrix characterizes the linear correlation between different metrics in the current operation data; perform eigenvalue decomposition on the covariance matrix to obtain a plurality of principal component feature values and a plurality of second feature vectors corresponding to the plurality of principal component feature values; arrange the plurality of second feature vectors in descending order of the principal component feature values, select the second feature vectors starting from the largest principal component feature value, and when the selected second feature vectors meet the preset conditions, determine the selected second feature vectors as the plurality of first feature vectors to obtain the fault feature set; the preset condition refers to the condition that the cumulative contribution rate of the selected second feature vectors is greater than the preset contribution rate threshold.

[0158] In an exemplary embodiment, the feature update module 506 is further configured to obtain the historical operation data of the server; the historical operation data includes multi-dimensional historical feature values; determine the historical mean and standard deviation of the current feature value of one dimension in the current operation data according to the historical operation data; and determine the importance score of the current feature value of one dimension in the current operation data according to the current feature value, historical mean, standard deviation, and first preset weight of one dimension in the current operation data.

[0159] In an exemplary embodiment, the multiple first feature vectors are the feature vectors extracted after performing principal component analysis on the current operation data; each single first feature vector in the multiple first feature vectors corresponds to a principal component eigenvalue; the feature update module 506 is further configured to perform weighted summation on the current feature value of one dimension in the current operation data according to the importance score of the current feature value of one dimension in the current operation data to obtain the first contribution degree corresponding to the current operation data; calculate the product between each single first feature vector in the fault feature set and the corresponding principal component eigenvalue, and determine the product result as the second contribution degree of each single first feature vector in the fault feature set; and determine the sum of the second contribution degree of each single first feature vector in the fault feature set and the first contribution degree corresponding to the current operation data as the updated single first feature vector in the fault feature set, so as to obtain the updated fault feature set.

[0160] In an exemplary embodiment, the apparatus further includes a fault prediction module, configured to input the updated fault feature set into a pre-trained fault prediction model to obtain a prediction value, where the prediction value represents the possibility of the server having a fault; generate a warning message when the prediction value is greater than a warning threshold; the warning threshold is determined according to the current operation data; the warning message includes response operation process information, and the response operation process information includes multiple operation steps; perform warning according to the warning message, and execute the multiple operation steps according to the time sequence in the response operation process information.

[0161] In an exemplary embodiment, the apparatus further includes a response module, configured to perform performance evaluation on the server after executing the multiple operation steps according to the time sequence in the response operation process information to obtain a fault response processing result; stop the fault response when the fault response processing result is greater than a preset response value.

[0162] In an exemplary embodiment, the response module is further configured to evaluate the preset performance index of the server after executing a single operation step in the multiple operation steps, and perform fusion processing on the preset performance index of the server after executing a single operation step in the multiple operation steps to obtain a fault response processing result.

[0163] In an exemplary embodiment, the response module is further configured to calculate the difference between the first parameter value of the preset performance metric of the server after executing a single operation step among multiple operation steps and the second parameter value of the preset performance metric of the server after executing the previous operation step, to obtain a plurality of differences; and determine the ratio between the sum of the plurality of differences and the total number of steps of the multiple operation steps as the fault response processing result.

[0164] In an exemplary embodiment, the response module is further configured to, when the fault response processing result is less than or equal to a preset response value, adjust the resource allocation ratio of multiple types of resources in the server to obtain a target resource configuration scheme, and adjust the resource configuration of the server according to the target resource configuration scheme.

[0165] In an exemplary embodiment, the response module is further configured to determine the new resource allocation ratio of multiple types of resources in the server according to the ratio between the current utilization rate of a single type of resource among the multiple types of resources and the maximum resource utilization rate among the multiple types of resources, and the second preset weight of the single type of resource among the multiple types of resources, and determine the set of the new resource allocation ratios of the multiple types of resources in the server as the target resource configuration scheme.

[0166] An embodiment of the present application further provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any of the above embodiments of the fault feature extraction method.

[0167] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps in any of the above embodiments of the fault feature extraction method when running.

[0168] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: a USB flash drive, a read-only memory (ROM for short), a random access memory (RAM for short), a mobile hard disk, a magnetic disk, or an optical disc, and other various media that can store a computer program.

[0169] An embodiment of the present application further provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, the steps in any of the above embodiments of the fault feature extraction method are implemented.

[0170] An embodiment of the present application further provides another computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in any of the above embodiments of the fault feature extraction method are implemented.

[0171] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of this application.

[0172] The above has introduced in detail a method for extracting fault features, an electronic device, and a storage medium provided by this application. Specific examples are used herein to elaborate on the principle and implementation manner of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the principle of this application, several improvements and modifications can be made to this application, and these improvements and modifications also fall within the protection scope of the claims of this application.

Claims

1. A method for extracting fault features, characterized in that, Including: Obtain the current operation data of the server; The current operation data includes multi-dimensional current eigenvalue; Extract a fault feature set from the current operation data, and calculate the importance score of the current eigenvalue of one dimension in the current operation data; The fault feature set includes multiple first feature vectors; Adjust the single first feature vector in the fault feature set according to the importance score of the current eigenvalue of one dimension in the current operation data, and obtain the updated fault feature set; The updated fault feature set is used for fault detection.

2. The method according to claim 1, wherein The obtaining the current operation data of the server includes: Collect the current operation data through a collection component pre-deployed in the server; the configuration parameters of the collection component are configured according to the number of nodes and network environment of the server cluster where the server is located.

3. The method according to claim 1, characterized in that, Before extracting the fault feature set from the current operation data, the method further includes: Identify, remove and fill the first abnormal data points in the current operation data to obtain the cleaned current operation data; the first abnormal data points refer to the data points whose eigenvalues are not within the preset range; Perform format conversion on the cleaned current operation data to obtain the current operation data in a unified format; Verify the current operation data in the unified format, identify and correct the second abnormal data points to obtain the processed current operation data; the second abnormal data points refer to the data points that do not conform to the historical behavior pattern or trend.

4. The method according to claim 1, characterized in that The extracting the fault feature set from the current operation data includes: Calculate the covariance matrix corresponding to the current operation data; the covariance matrix characterizes the linear correlation between different indicators in the current operation data; Perform eigenvalue decomposition on the covariance matrix to obtain multiple principal component eigenvalues and multiple second feature vectors corresponding to the multiple principal component eigenvalues; Arrange the multiple second feature vectors in descending order of the principal component eigenvalues, select the second feature vectors starting from the largest principal component eigenvalue, and when the selected second feature vectors meet the preset conditions, determine the selected second feature vectors as the multiple first feature vectors to obtain the fault feature set; the preset condition refers to the condition that the cumulative contribution rate of the selected second feature vectors is greater than the preset contribution rate threshold.

5. The method according to claim 1, wherein The calculating the importance score of the current eigenvalue of one dimension in the current operation data includes: Obtain the historical operation data of the server; the historical operation data includes the multi-dimensional historical eigenvalues; Determine the historical mean and standard deviation of the current eigenvalue of one dimension in the current operation data according to the historical operation data; Determine the importance score of the current eigenvalue of one dimension in the current operation data according to the current eigenvalue of one dimension in the current operation data, historical mean, standard deviation and the first preset weight.

6. The method according to claim 1, wherein The multiple first feature vectors are the feature vectors extracted after performing principal component analysis on the current operation data; a single first feature vector in the multiple first feature vectors corresponds to a principal component eigenvalue; the adjusting the single first feature vector in the fault feature set according to the importance score of the current eigenvalue of a dimension in the current operation data to obtain the updated fault feature set includes: According to the importance score of the current eigenvalue of a dimension in the current operation data, performing weighted summation on the current eigenvalues of a dimension in the current operation data to obtain a first contribution degree corresponding to the current operation data; Calculating the product between the single first feature vector in the fault feature set and the corresponding principal component eigenvalue, and determining the product result as the second contribution degree of the single first feature vector in the fault feature set; Determining the sum of the second contribution degree of the single first feature vector in the fault feature set and the first contribution degree corresponding to the current operation data as the updated single first feature vector in the fault feature set, so as to obtain the updated fault feature set.

7. The method according to claim 1, wherein The method further includes: Inputting the updated fault feature set into a pre-trained fault prediction model to obtain a prediction value, where the prediction value represents the possibility of a fault occurring in the server; Generating a warning message when the prediction value is greater than a warning threshold; the warning threshold is determined according to the current operation data; the warning message includes response operation process information, and the response operation process information includes a plurality of operation steps; Performing a warning according to the warning message, and executing the plurality of operation steps in the order of time sequence in the response operation process information.

8. The method according to claim 7, characterized in that, After executing the plurality of operation steps in the order of time sequence in the response operation process information, the method further includes: Performing a performance evaluation on the server after executing the plurality of operation steps to obtain a fault response processing result; stopping the fault response when the fault response processing result is greater than a preset response value.

9. The method according to claim 8, characterized in that The performing a performance evaluation on the server after executing the plurality of operation steps to obtain a fault response processing result includes: Evaluating a preset performance index of the server after executing a single operation step in the plurality of operation steps, and performing fusion processing on the preset performance index of the server after executing a single operation step in the plurality of operation steps to obtain the fault response processing result.

10. The method according to claim 9, characterized in that, The performing fusion processing on the preset performance index of the server after executing a single operation step in the plurality of operation steps to obtain a fault response processing result includes: Calculating the difference between a first parameter value of the preset performance index of the server after executing a single operation step in the plurality of operation steps and a second parameter value of the preset performance index of the server after executing the previous operation step to obtain a plurality of differences; Determining the ratio between the sum of the plurality of differences and the total number of steps of the plurality of operation steps as the fault response processing result.

11. The method according to claim 8, wherein The method further includes: In the case that the fault response processing result is less than or equal to the preset response value, adjust the resource allocation ratios of multiple types of resources in the server to obtain a target resource configuration plan, and adjust the resource configuration of the server according to the target resource configuration plan.

12. The method according to claim 11, wherein The adjusting the resource allocation ratios of multiple types of resources in the server to obtain a target resource configuration plan includes: Determine the new resource allocation ratios of multiple types of resources in the server according to the ratio between the current utilization rate of a single type of resource in the multiple types of resources and the maximum resource utilization rate in the multiple types of resources, and the second preset weight of the single type of resource in the multiple types of resources, and determine the set of the new resource allocation ratios of multiple types of resources in the server as the target resource configuration plan.

13. An electronic device, characterized in that, including: A memory for storing a computer program; A processor for implementing the steps of the fault feature extraction method according to any one of claims 1 to 12 when executing the computer program.

14. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, wherein the computer program implements the steps of the fault feature extraction method according to any one of claims 1 to 12 when being executed by a processor.

15. A computer program product, comprising a computer program, characterized in that, The computer program implements the steps of the fault feature extraction method according to any one of claims 1 to 12 when being executed by a processor.

Citation Information

Patent Citations

  • Fault monitoring method and device, computer equipment and storage medium

    CN117130886A

  • Distributed storage node fault detection system

    CN119718741A

  • Machine room equipment remote operation and maintenance management method and system based on Internet of Things technology

    CN119863234A

  • Equipment fault identification method and system and storage medium

    CN119939277A

  • Latency-aware resource allocation for stream processing applications

    US20240394110A1

Cited By

  • Hardware fault information management method and device, electronic equipment and storage medium

    CN120540925A

  • Bridge health detection method and device, electronic equipment and readable storage medium

    CN121580212A