Intelligent fault self-healing method based on multi-dimensional indexes

By collecting multi-dimensional indicator data and performing feature engineering, combined with XGBoost classifier and service dependency graph optimization, the problems of misjudgment and inflexible repair strategies in existing technologies for fault detection and repair are solved, and an efficient and accurate fault self-healing process is achieved.

CN120561758BActive Publication Date: 2025-10-17北京科杰科技有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511063157.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-31
Publication Date
2025-10-17
Estimated Expiration
2045-07-31

AI Technical Summary

Technical Problem

Existing technologies for fault detection and repair suffer from reliance on a single indicator, leading to misjudgments or omissions. They lack flexibility and context awareness, cannot effectively handle complex distributed system faults, and their repair strategies lack consideration for inter-service dependencies.

Method used

By collecting multi-dimensional indicator data and performing feature engineering, combined with the XGBoost classifier for fault classification, a service dependency graph optimization and repair strategy is constructed to achieve automated repair operations.

Benefits of technology

It improves the accuracy of fault identification and the precision of repair strategies, reduces manual intervention, shortens fault recovery time, and enhances system availability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120561758B_ABST
    Figure CN120561758B_ABST
Patent Text Reader

Abstract

The application provides an intelligent fault self-healing method based on multidimensional indexes, relates to the technical field of fault detection and repair, comprises over-collecting multi-source data and performing feature engineering processing, using an XGBoost classifier to perform fault classification, selecting and combining a repair strategy according to a fault type, a controller type and a service level feature, and performing a repair operation through a service dependency graph atlas optimization. The application can improve fault diagnosis accuracy, shorten fault repair time, reduce the demand for manual intervention, and improve the reliability of cloud platform services.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to fault detection and repair technology, and in particular to an intelligent fault self-healing method based on multi-dimensional indicators. BACKGROUND

[0002] With the widespread application of cloud native architecture, Kubernetes-based container orchestration platforms have become the core components of enterprise IT infrastructure. In complex distributed systems, application failures are inevitable, which may be caused by resource competition, configuration errors, abnormal dependent services, and other factors. Traditional fault handling methods usually rely on manual intervention, and operation and maintenance personnel need to collect logs, analyze indicators, and manually execute repair operations. This method is inefficient and prone to errors when facing large-scale clusters.

[0003] In order to improve system availability and reduce operation and maintenance costs, the industry has begun to explore automated fault detection and repair solutions. Existing technologies mainly use single indicator-based or simple rule-based fault detection mechanisms, such as setting threshold values based on CPU utilization or memory usage to trigger alarms, and combining pre-defined simple repair strategies to perform automated operations. However, this method has obvious limitations in practical application:

[0004] Existing fault detection mechanisms rely too much on single-dimensional indicator data and lack comprehensive analysis capabilities for multi-source heterogeneous data. Single indicators are difficult to fully reflect the health status of the system, which can easily lead to misjudgment or omission, especially for complex distributed system failures, single indicators often cannot accurately identify the root cause.

[0005] Existing fault repair strategies lack flexibility and context awareness. Most automated repair solutions use static pre-defined strategies, which cannot dynamically adjust repair strategies based on workload types, service importance, and other factors, resulting in repair operations that may not be suitable for specific scenarios, and even may trigger a chain reaction, further worsening the system state.

[0006] Existing technologies usually ignore the impact of service dependencies on fault repair. In microservice architecture, there are complex dependencies between services, and the abnormality of one service may affect multiple related services. Existing repair mechanisms lack awareness and consideration of such dependencies, making it difficult to achieve globally optimal repair decisions, and easily leading to new problems caused by repair operations or unsatisfactory repair results. SUMMARY

[0007] The embodiments of the present application provide an intelligent fault self-healing method based on multi-dimensional indicators, which can solve the problems in the prior art.

[0008] In a first aspect of the embodiments of the present application, an intelligent fault self-healing method based on multi-dimensional indicators is provided, comprising:

[0009] The pod restart number data is collected through the Kubernetes API, the CPU utilization and memory usage data is collected through Prometheus, the file damage keyword data is collected and parsed through the ELK Stack, and the error rate data is collected through application of a custom monitoring, to form multi-source data;

[0010] The multi-source data is subjected to feature engineering processing, including feature encoding of a controller type to which a pod belongs to obtain a controller type feature, and feature encoding of a service level in a pod label to obtain a service level feature, to generate a multi-dimensional feature vector;

[0011] The multi-dimensional feature vector is input into an XGBoost classifier to obtain a fault type; and a basic strategy template is selected according to the fault type;

[0012] The basic strategy template is combined based on the controller type feature and the service level feature to form a combined repair strategy; and a service dependency graph is constructed to optimize the combined repair strategy;

[0013] The optimized combined repair strategy is sent to a Kubernetes control plane, and a corresponding repair operation is executed by calling a Kubernetes API, the repair operation including one or more combinations of a pod restart operation, a node expansion operation, and a node eviction operation.

[0014] In an optional implementation,

[0015] The feature engineering processing of the multi-source data includes:

[0016] The controller type to which the pod belongs is one-hot encoded to generate the controller type feature, the CPU utilization and memory usage data is normalized to generate a resource usage feature, the pod restart number data is raw-counted to generate a restart number feature, the file damage keyword data is Boolean feature converted to generate a log abnormality feature, the service level in the pod label is ordinal encoded to generate a service level feature, the service level is divided into multiple levels according to a business importance degree from high to low, a database storage type in a pod storage volume declaration configuration is Boolean feature converted to generate a storage volume feature, and the error rate data is time windowed to generate a service call feature;

[0017] The contribution of each feature to fault classification is calculated based on information gain, the correlation between features is calculated based on a Pearson correlation coefficient, and an optimal feature set is selected based on the contribution and the correlation;

[0018] Based on the preferred feature set, the controller type feature is cross-combined with the service level feature to generate a service attribute feature, the resource usage feature is weighted combined with the restart number feature to generate a stability feature, the log anomaly feature is associated combined with the storage volume feature to obtain an initial dependency feature, and the service call feature is associated combined with the initial dependency feature to generate an enhanced dependency feature; and the service attribute feature, the stability feature, and the enhanced dependency feature are spliced to generate a multi-dimensional feature vector.

[0019] In an optional implementation,

[0020] Inputting the multi-dimensional feature vector into an XGBoost classifier for fault classification includes:

[0021] Performing wavelet transform on the resource usage feature to obtain a time-frequency feature, performing BERT encoding on the log anomaly feature to obtain a semantic feature, splicing the time-frequency feature and the semantic feature to the multi-dimensional feature vector to form an enhanced feature vector, and performing mean-variance standardization processing on the enhanced feature vector to obtain a standardized feature vector;

[0022] Performing sample balancing processing on historical fault labeled data to obtain training data, and filtering the training data based on labeling consistency and repair effect to form a core training set;

[0023] Layering the core training set according to fault types and the service level feature, training an XGBoost classifier using the standardized feature vector to obtain a layered fault classification model, determining a level to which a failure case belongs when a repair strategy corresponding to the fault type fails, adding the failure case to a negative sample pool of the corresponding level, and performing incremental training on the layered fault classification model of the corresponding level based on a time decay weight of samples in the negative sample pool;

[0024] Calculating a prediction probability distribution of the XGBoost classifier to obtain an uncertainty index of sample classification, and adding samples with the uncertainty index exceeding a preset index threshold to the core training set after manual confirmation;

[0025] Building an evaluation index system including accuracy, recall rate, and F1 score to evaluate the XGBoost classifier, and optimizing model parameters according to consistency between a prediction result and an actual repair effect when evaluation indexes are lower than a preset evaluation threshold.

[0026] In an optional implementation,

[0027] Layering the core training set for training and performing incremental training based on failure samples includes:

[0028] The core training set is allocated according to fault type and service level to obtain an initial stratified sample set. The sample sharing coefficient is calculated by calculating the ratio of the number of samples in each initial stratified sample set to the maximum number of samples in all initial stratified sample sets. When the sample sharing coefficient is less than a preset sample threshold, the cosine similarity between the adjacent level samples and the current level class center is calculated. The adjacent level samples whose cosine similarity is greater than the preset similarity threshold are added to the current level sample set to form an expanded stratified sample set. The samples in the expanded stratified sample set are trained to obtain a stratified fault classification model.

[0029] Receive failed samples from a negative sample pool, calculate a time decay factor based on the time interval of the failed samples, calculate a degree of repair failure based on the ratio of the number of retries of the failed samples during the repair process to a preset maximum number of retries, calculate a sample importance coefficient based on the degree of repair failure, and use the product of the time decay factor and the sample importance coefficient as a time decay weight;

[0030] The ratio of the number of samples in the negative sample pool to the preset maximum number of samples is calculated to obtain the sample capacity ratio, and the product of the difference between the preset benchmark learning rate and the sample capacity ratio is used as the adaptive learning rate; the ratio of the classification accuracy change to the time interval is calculated to obtain the accuracy change ratio, and the product of the preset benchmark update threshold and the accuracy change ratio is used as the dynamic update threshold; when the sum of the time decay weights is greater than the dynamic update threshold, the hierarchical fault classification model is incrementally trained using the adaptive learning rate to obtain an optimized hierarchical fault classification model.

[0031] In an optional embodiment,

[0032] The basic policy template includes:

[0033] A first basic policy template for data-type failures, which includes basic operations such as traffic suspension through the Ingress API, data backup through the CSI snapshot interface, and safe restart through the Kubernetes API;

[0034] The second basic policy template for resource-based failures includes basic operations such as current limiting and degradation through the HPA interface, resource restriction through the Cgroup interface, and node eviction through the Kubernetes drain interface.

[0035] The third basic policy template for dependency-type failures includes basic operations of performing dependency chain traversal through the dependency graph interface and performing circuit breaking and degradation through the Hystrix API.

[0036] In an optional implementation,

[0037] Combining the base strategy templates based on the controller type features and the service level features to form a combined repair strategy includes:

[0038] When the controller type feature is a StatefulSet and the fault type is a data type fault, sorting the operations in the first base strategy template based on historical repair success rates, selecting a traffic suspension operation, a data backup operation with the highest repair success rate, and adding a storage volume retention operation based on storage volume state evaluation, an image switching operation based on available image version evaluation, to generate a first combined strategy;

[0039] When the controller type feature is a Deployment and the fault type is a resource type fault, determining a throttling degradation threshold based on resource usage features, selecting a throttling degradation operation, and adding a Pod restart operation based on Pod restart history evaluation, a number of replica expansion operations based on cluster resource margin evaluation, and connection pool fuse parameter configuration based on service call volume evaluation, to generate a second combined strategy;

[0040] When the service level feature is higher than a preset level threshold and the fault type is a resource type fault, selecting a limit parameter of a resource limit operation based on multi-dimensional resource profiling, and selecting an optimal migration target node based on target node health score, to generate a third combined strategy;

[0041] According to different service level features, different execution priorities are set for the first combined strategy, the second combined strategy, or the third combined strategy.

[0042] In an optional implementation,

[0043] Optimizing the combined repair strategy based on the service dependency graph includes:

[0044] The service dependency graph includes service node information, service call relationship edge information, and storage dependency edge information; the service node information includes Pod name, controller type feature, and service level feature, the service call relationship edge information records the direction and error rate data of service request calls, and the storage dependency edge information records the association relationship between services and storage volumes declared by the services;

[0045] Optimize the combined repair strategy based on the service dependency graph, analyze the error rate data propagation path in the service call relationship edge information, analyze the storage volume access state in the storage dependency edge information, and determine the fault root cause and propagation link in combination with the resource usage characteristics; determine the fault influence range based on the controller type characteristics and service level characteristics in the service node information, and optimize the execution order of the first combined strategy, the second combined strategy or the third combined strategy according to the fault root cause and the propagation link on the premise of maintaining the service level priority constraint, and generate the optimized repair strategy.

[0046] The second aspect of the embodiment of the application provides an electronic device, comprising:

[0047] A processor;

[0048] A memory for storing processor-executable instructions;

[0049] The processor is configured to call the instructions stored in the memory to execute the method described above.

[0050] The third aspect of the embodiment of the application provides a computer-readable storage medium having computer program instructions stored thereon, wherein the computer program instructions are executed by a processor to implement the method described above.

[0051] The application realizes accurate classification of fault types in the Kubernetes environment by constructing a multi-dimensional data collection and feature engineering system, improves the accuracy of fault identification, and reduces the misjudgment rate, so that the self-healing process is more accurate and efficient.

[0052] The application combines the controller type and service level characteristics to optimize the basic strategy template, can formulate differentiated repair strategies for different types of applications and services, and further optimizes the repair strategy through the service dependency graph to ensure that the repair operation conforms to the business logic and dependency relationship, and improves the accuracy and effectiveness of the repair strategy.

[0053] The application automatically executes the repair operation through the Kubernetes API, realizes the full-process automation from fault detection, classification to repair, significantly reduces manual intervention, shortens the fault recovery time, improves the system availability, and has important practical value for large-scale cluster management. BRIEF DESCRIPTION OF DRAWINGS

[0054] Figure 1 It is a flowchart of the intelligent fault self-healing method based on multi-dimensional indexes of the embodiment of the application;

[0055] Figure 2 It is a flowchart of constructing a service dependency graph to optimize the combined repair strategy. DETAILED DESCRIPTION

[0056] In order to make the objects, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative work fall within the protection scope of the present application.

[0057] The technical solutions of the present application will be described in detail below with specific embodiments. The following specific embodiments can be combined with each other, and some embodiments may not be described again for the same or similar concepts or processes.

[0058] Figure 1 The flowchart of the intelligent fault self-healing method based on multi-dimensional indexes of the embodiments of the present application is shown in FIG. 1, which comprises the following steps. Figure 1

[0059] The number of Pod restarts is collected through the Kubernetes API, the CPU utilization and memory usage are collected through the Prometheus, the file damage keyword data is collected and parsed through the ELK Stack, and the error rate data is collected through the application of a custom monitoring, thereby forming multi-source data.

[0060] The multi-source data is subjected to feature engineering processing, including feature encoding of the controller type to which the Pod belongs to obtain the controller type feature, feature encoding of the service level in the Pod label to obtain the service level feature, and generation of a multi-dimensional feature vector.

[0061] The multi-dimensional feature vector is input into an XGBoost classifier to obtain a fault type; and a basic strategy template is selected according to the fault type.

[0062] The basic strategy template is combined based on the controller type feature and the service level feature to form a combined repair strategy; and the combined repair strategy is optimized by constructing a service dependency graph.

[0063] The optimized combined repair strategy is sent to the Kubernetes control plane, and a corresponding repair operation is performed by calling the Kubernetes API, wherein the repair operation includes one or more combinations of a Pod restart operation, a node expansion operation and a node eviction operation.

[0064] In an optional embodiment, the feature engineering processing of the multi-source data comprises:

[0065] ​The controller type of the Pod is one-hot encoded to generate a controller type feature, the CPU utilization and memory usage data are normalized to generate a resource usage feature, the Pod restart number data is raw counted to generate a restart number feature, the file damage keyword data is Boolean feature converted to generate a log abnormality feature, the service level in the Pod label is ordinal encoded to generate a service level feature, the service level is divided into multiple levels from high to low according to the business importance, the database storage type in the Pod storage volume declaration configuration is Boolean feature converted to generate a storage volume feature, and the error rate data is time windowed to generate a service call feature.

[0066] The contribution of each feature to the fault classification is calculated based on information gain, the correlation between features is calculated based on Pearson correlation coefficient, and the preferred feature set is obtained based on the contribution and correlation.

[0067] Based on the preferred feature set, the controller type feature and the service level feature are cross combined to generate a service attribute feature, the resource usage feature and the restart number feature are weighted combined to generate a stability feature, the log abnormality feature and the storage volume feature are associated combined to obtain an initial dependency feature, and the service call feature and the initial dependency feature are associated combined to generate an enhanced dependency feature; and the service attribute feature, the stability feature and the enhanced dependency feature are spliced to generate a multi-dimensional feature vector.

[0068] For example, the embodiment obtains Pod related data in a Kubernetes cluster, including controller type, CPU utilization, memory usage, restart number, log keyword, Pod label and storage volume configuration and the like, and then converts them into feature vectors available for machine learning models through feature engineering.

[0069] When the controller type of the Pod is one-hot encoded to generate a controller type feature, the controller types such as Deployment, StatefulSet and DaemonSet are respectively mapped into binary vectors. For example, assuming that there are three controller types, a Pod belongs to the Deployment type, then a vector of [1, 0, 0] is generated; if it belongs to the StatefulSet type, a vector of [0, 1, 0] is generated. This encoding method effectively preserves the category information of the controller type and avoids introducing inappropriate numerical order relationships.

[0070] When generating the resource usage features by normalizing the CPU utilization and memory usage data, the min-max normalization method is used to convert the CPU utilization and memory usage to the range of 0-1. Specifically, for an original value x, the normalized value is (x-min) / (max-min), where min and max are the minimum and maximum values of the indicator, respectively. For example, if the CPU utilization is 75%, the minimum value is 10%, and the maximum value is 90%, the normalized value is (75-10) / (90-10)=0.8125. This processing allows resource indicators of different scales to be compared on the same dimension.

[0071] When generating the restart count feature by performing raw count on the Pod restart count data, the integer count value is directly used. For example, if a Pod restarts 3 times within the observation period, the restart count feature value is 3. Frequent restarts of a Pod are usually an important signal of stability problems, and preserving the raw count can directly reflect this characteristic.

[0072] When generating the log exception feature by performing Boolean feature conversion on the file corruption keyword data, it is detected whether the log contains keywords such as "file corruption", "checksum error", etc. If it contains, the feature value is 1, otherwise it is 0. For example, if the log of a Pod contains the words "data file checksum mismatch", its log exception feature value is 1, indicating that there is a file corruption problem.

[0073] When generating the service level feature by ordinal encoding the service level in the Pod label, the service level is mapped to an integer according to the business importance from high to low. For example, core service, critical service, ordinary service, and non-critical service are encoded as 3, 2, 1, and 0, respectively. If the label of a Pod contains "service-level: critical", its service level feature value is 2.

[0074] When generating the storage volume feature by performing Boolean feature conversion on the database storage type in the Pod storage volume claim configuration, it is checked whether the storage volume is of a database-specific type. If it is, the feature value is 1, otherwise it is 0. For example, if the storage volume claim of a Pod specifies the "database-storage" type, its storage volume feature value is 1, indicating that the Pod runs a database service.

[0075] When generating the service call feature by performing time window statistics on the error rate data, the service call error rate is calculated within a 5-minute sliding window. For example, a Pod receives 100 requests in the last 5 minutes, of which 5 return errors, so the error rate is 5%. In order to capture the trend of the error rate, the error rates of multiple consecutive windows can be calculated, such as [3%, 4%, 5%, 4%, 6%], forming an error rate sequence as the service call feature.

[0076] When calculating the feature contribution degree based on information gain, the association degree of the distribution of each feature with the fault type is counted, the continuous features are discretized and divided into multiple intervals; then the distinguishing ability of each value of the feature to fault classification is calculated, and the higher the value, the greater the contribution of the feature to classification.

[0077] When evaluating feature redundancy based on Pearson correlation coefficient, the linear correlation strength between each pair of features is calculated, with a value range of [-1, 1]. The original data needs to be standardized before calculation, and then the ratio of the covariance between the feature pairs to the standard deviation is calculated. A value close to ±1 indicates a high correlation, and a value close to 0 indicates a low correlation. Set 0.8 as the redundancy threshold to filter out highly overlapping features.

[0078] The preferred feature set screening adopts a two-step method: first, arrange all features in descending order of information gain, and retain the top 80% high-contribution features; then detect the correlation coefficient of the feature pairs, and retain the feature with higher information gain when the correlation coefficient exceeds 0.8. For example, the correlation coefficient between the Pod status code and the restart count is 0.87, and the restart count feature with higher information gain (0.33) is retained.

[0079] Based on the preferred feature set, when generating service attribute features by cross-combining controller type features and service level features, vector outer product or feature value connection method is used. For example, if the controller type is [1, 0, 0] and the service level is 2, a service attribute feature of [2, 0, 0] can be generated, indicating that the Pod is running a critical business Deployment.

[0080] When generating stability features by weighted combination of resource usage features and restart count features, the normalized CPU utilization, memory usage, and restart count are weighted and summed. For example, if the CPU utilization is 0.8, the memory usage is 0.7, and the restart count is normalized to 0.2, with weights of 0.4, 0.4, and 0.2 respectively, then the stability feature value is 0.8×0.4+0.7×0.4+0.2×0.2=0.64.

[0081] When generating initial dependency features by associating log anomaly features and storage volume features, the combination pattern of log anomaly and storage volume type is detected. For example, when the log anomaly feature is 1 and the storage volume feature is 1, the initial dependency feature is 1, indicating that there is a database storage related fault; otherwise, it is 0.

[0082] When the service call feature is combined with the initial dependency feature to generate an enhanced dependency feature, the error rate trend is combined with the initial dependency feature. For example, if the recent error rate is in an increasing trend [3%, 4%, 5%, 7%, 10%] and the initial dependency feature is 1, the enhanced dependency feature can be represented as [0.03, 0.04, 0.05, 0.07, 0.10, 1], capturing the combined information of service call anomalies and dependency relationships.

[0083] Finally, the service attribute feature, the stability feature, and the enhanced dependency feature are spliced to generate a multi-dimensional feature vector. For example, if the service attribute feature is [2, 0, 0], the stability feature is 0.64, and the enhanced dependency feature is [0.03, 0.04, 0.05, 0.07, 0.10, 1], the final generated feature vector is [2, 0, 0, 0.64, 0.03, 0.04, 0.05, 0.07, 0.10, 1]. This multi-dimensional feature vector comprehensively considers the service attributes, running stability, and dependency relationships of the Pod, providing a rich information foundation for subsequent fault detection and classification.

[0084] The present application realizes comprehensive extraction and optimized combination of fault-related data through multi-dimensional feature engineering processing and feature combination, improves the expression ability and discrimination of the features. The feature screening mechanism based on information gain and correlation ensures the effectiveness of the features, and the cross combination of the features enhances the associated expression between the features, thereby improving the accuracy of subsequent fault classification.

[0085] In an optional implementation, inputting the multi-dimensional feature vector into an XGBoost classifier for fault classification includes:

[0086] Performing wavelet transform on the resource usage feature to obtain a time-frequency feature, performing BERT encoding on the log anomaly feature to obtain a semantic feature, splicing the time-frequency feature and the semantic feature to the multi-dimensional feature vector to form an enhanced feature vector, and performing mean-variance standardization processing on the enhanced feature vector to obtain a standardized feature vector;

[0087] Performing sample balancing processing on historical fault labeled data to obtain training data, and screening the training data based on labeled consistency and repair effect to form a core training set;

[0088] Layering the core training set according to fault types and the service level feature, training an XGBoost classifier using the standardized feature vector to obtain a layered fault classification model, determining the level to which a failed case belongs when a repair strategy corresponding to the fault type fails, adding the failed case to a negative sample pool of the corresponding level, and performing incremental training on the layered fault classification model of the corresponding level based on the time decay weight of the samples in the negative sample pool;

[0089] The prediction probability distribution of the XGBoost classifier is calculated to obtain an uncertainty index of sample classification, and samples with an uncertainty index exceeding a preset index threshold are added to the core training set after manual confirmation;

[0090] An evaluation index system including accuracy, recall rate, and F1 score is constructed to evaluate the XGBoost classifier, and when the evaluation index is lower than a preset evaluation threshold, the model parameters are optimized according to the consistency of the prediction result and the actual repair effect.

[0091] For example, the step of inputting a multi-dimensional feature vector into the XGBoost classifier for fault classification includes wavelet transformation of resource usage features, BERT encoding of log anomaly features, feature splicing processing, and model training and optimization.

[0092] For resource usage features, DB4 wavelet bases are used to decompose 3 layers of time series data of resource indicators such as CPU utilization, memory occupancy, and disk IO, extract low-frequency coefficients as stable trend features, and high-frequency coefficients as fluctuation features. For example, wavelet decomposition is performed on the CPU utilization curve of a server 30 minutes before the fault to obtain 4 groups of coefficients: low-frequency approximation coefficients [0.82, 0.85, 0.79, 0.91, 0.95] and three groups of high-frequency detail coefficients [0.03, -0.02, 0.04, -0.05, 0.01], [0.01, -0.03, 0.02, -0.01, 0.02], [0.005, 0.002, -0.004, 0.003, -0.001], which together constitute a time-frequency feature vector.

[0093] The log anomaly features are processed, and first, the key logs are preprocessed, including timestamp extraction, variable replacement, and error code standardization. The preprocessed logs are used to extract semantic representations using a pre-trained BERT-Base model (containing 12 layers, 768-dimensional hidden layers, and 12 attention heads). For example, the error log "Connection timeout while accessing database server" is encoded to obtain a 768-dimensional semantic vector, and the first 5 dimensions are [0.153, -0.278, 0.421, -0.087, 0.309].

[0094] The time-frequency feature and the semantic feature are spliced to form an enhanced feature vector. Assuming that the time-frequency feature dimension is 120 and the semantic feature dimension is 768, an enhanced feature vector of 888 dimensions is formed after splicing. The feature vector is subjected to mean-variance standardization processing, and the mean and standard deviation of each feature dimension on the training set are calculated. The feature value is subtracted from the mean and then divided by the standard deviation. For example, the original value of the first dimension feature is 0.82, the mean of the dimension on the training set is 0.75, and the standard deviation is 0.08. The standardized value is (0.82-0.75) / 0.08=0.875.

[0095] When the historical failure labeled data is subjected to sample balancing processing to obtain training data, the number of samples of each failure type is counted. For rare failure types (such as network partition failure) with a number less than 100, SMOTE algorithm is used to generate synthetic samples. For example, a rare failure type has only 20 samples. Through the SMOTE algorithm, 80 synthetic samples are generated in the feature space, so that the number of samples of this category reaches 100. Usually, the sampling ratio is set to 80% of the overall average sample number. When the training data is screened based on labeling consistency, the success rate of the repair strategy execution of each failure sample is calculated, and samples with a success rate greater than 85% are selected as the core training set.

[0096] The core training set is layered according to the failure type and service level features. For example, the failures are divided into four categories: "network failure", "disk failure", "memory failure", and "CPU failure", and the service levels are divided into three levels: "critical business", "important business", and "ordinary business", forming 12 hierarchical sub-training sets. The XGBoost classifier is trained using the standardized feature vector to obtain a hierarchical failure classification model. Different model parameters are configured for each layer: a deep tree depth (maximum depth of 8) and a small learning rate (0.01) are set for the critical business layer to ensure classification accuracy; a shallow tree depth (maximum depth of 5) and a large learning rate (0.05) are set for the ordinary business layer to improve training speed.

[0097] When the repair strategy corresponding to the failure type fails, the failure case is determined to belong to a certain level, and the failure case is added to the negative sample pool of the corresponding level. The samples in the negative sample pool are assigned a time-decay-based weight. The hierarchical failure classification model of the corresponding level is incrementally trained based on the time-decay-based weight of the samples in the negative sample pool. Incremental updating is performed once a week, and the negative samples in the last 30 days are used for each update.

[0098] The prediction probability distribution of the XGBoost classifier is calculated to obtain an uncertainty index of sample classification, and the prediction probability entropy value is used as the uncertainty measurement. For example, the prediction probability distribution of a certain fault sample is [0.6, 0.3, 0.05, 0.05], and the calculated entropy value is 1.05. The preset index threshold is set to 0.9, and the samples with an entropy value greater than 0.9 are added to the core training set after manual confirmation. The manual confirmation link determines the fault type by expert voting, and updates the label after reaching a consensus.

[0099] An evaluation index system including accuracy, recall rate, and F1 score is constructed to evaluate the XGBoost classifier, and these indexes are calculated periodically on the test set. When the evaluation index is lower than the preset evaluation threshold (accuracy < 0.85, recall rate < 0.80, or F1 score < 0.82), the model parameters are optimized according to the consistency of the prediction result and the actual repair effect. Parameter optimization includes adjusting the number of trees (in the range of 50 to 500), the maximum depth (in the range of 3 to 10), the learning rate (in the range of 0.01 to 0.2), and the sampling ratio (in the range of 0.6 to 0.9). The optimal parameter combination is determined by grid search and applied to the next round of model training, so that the model maintains high classification performance in new fault scenarios.

[0100] Traditional Kubernetes fault detection mainly relies on threshold rules and simple statistical methods, which are difficult to cope with complex and variable cloud-native environments, especially when the fault forms are diverse and interrelated. The present invention introduces a hierarchical XGBoost classification architecture, which layers the model according to service levels and fault types, solving the practical problem of requiring different precision for different important level services. At the same time, a negative sample incremental learning mechanism based on time decay weight is designed, enabling the model to continuously learn from failed cases and optimize itself. Through sample balancing and incremental training mechanism, the problem of unbalanced fault sample distribution is solved, and the reliability and adaptability of the model are improved based on the uncertainty index manual confirmation mechanism.

[0101] In an optional implementation, the hierarchical training of the core training set and the incremental training based on failed samples include:

[0102] The core training set is allocated according to fault types and service levels to obtain an initial hierarchical sample set, and the ratio of the sample quantity of each initial hierarchical sample set to the maximum sample quantity in all initial hierarchical sample sets is calculated to obtain a sample sharing coefficient; when the sample sharing coefficient is less than a preset sample threshold, the cosine similarity between adjacent level samples and the current level class center is calculated, and the adjacent level samples with a cosine similarity greater than a preset similarity threshold are added to the current level sample set to form an expanded hierarchical sample set; the samples in the expanded hierarchical sample set are trained to obtain a hierarchical fault classification model;

[0103] receive failure samples in the negative sample pool, calculate a time decay factor based on a time interval of the failure samples, calculate a repair failure degree based on a ratio of a number of retries of the failure samples in a repair process to a preset maximum number of retries, calculate a sample importance coefficient based on the repair failure degree, and take a product of the time decay factor and the sample importance coefficient as a time decay weight;

[0104] calculate a sample capacity ratio by calculating a ratio of a number of samples in the negative sample pool to a preset maximum number of samples, take a product of a preset baseline learning rate and a difference between the sample capacity ratio as an adaptive learning rate, calculate an accuracy change ratio by calculating a ratio of a classification accuracy change to a time interval, take a product of a preset baseline update threshold and the accuracy change ratio as a dynamic update threshold, and use the adaptive learning rate to perform incremental training on the hierarchical fault classification model to obtain an optimized hierarchical fault classification model when a sum of the time decay weights is greater than the dynamic update threshold.

[0105] For example, the fault types are divided into multiple categories such as network fault, hardware fault, software fault, and the service levels are divided into three levels of high, medium and low. For example, for a core training set containing 10,000 fault samples, according to the combination of fault types and service levels, multiple initial hierarchical sample sets such as network fault-high level, network fault-medium level, network fault-low level, hardware fault-high level, and the like can be obtained.

[0106] Taking the sample distribution of the initial hierarchical sample set as an example, there are 500 samples of network fault-high level, 2000 samples of hardware fault-medium level, and 300 samples of software fault-low level. Assuming that the maximum number of samples is 2000, the sample sharing coefficient of network fault-high level is 500 / 2000=0.25, and the sample sharing coefficient of software fault-low level is 300 / 2000=0.15. When the sample sharing coefficient is less than a preset sample threshold, sample expansion is needed. Assuming that the preset sample threshold is 0.2, the sample sharing coefficient of software fault-low level 0.15 is less than the preset sample threshold, and sample expansion is needed. Calculate the cosine similarity between the adjacent level samples and the current level class center. For software fault-low level, the adjacent levels are software fault-medium level and other fault types-low level. Through feature vector calculation, assuming that there are 30 samples of software fault-medium level whose cosine similarity with the class center of software fault-low level is greater than a preset similarity threshold 0.8, the 30 samples are added to the software fault-low level sample set to form an expanded hierarchical sample set, and the number of samples after expansion reaches 330.

[0107] A classification model is trained using the random forest algorithm on the samples in the augmented stratified sample set, for example, for the augmented stratified sample set of software failure-low, a stratified failure classification model is trained using 330 samples, and the classification accuracy of the model is 92%.

[0108] Failure samples in the negative sample pool are received, which are samples that the classification model cannot correctly classify. Assume that after a week of operation, 50 failure samples have accumulated in the negative sample pool. Calculate the time decay factor based on the time interval of the failure samples. For example, set the time decay factor to 0.7 for failure samples that are 3 days old, and set the time decay factor to 0.9 for failure samples that are 1 day old.

[0109] Assume that the preset maximum number of retries is 10, and a failure sample has been retried 7 times and still has not been repaired, then the repair failure degree is 7 / 10=0.7. Calculate the sample importance coefficient based on the repair failure degree, which can be directly used as the sample importance coefficient, or can be converted through a mapping function, for example, sample importance coefficient=1-e^(-repair failure degree), and the corresponding sample importance coefficient is 0.5. The product of the time decay factor and the sample importance coefficient is used as the time decay weight. For this failure sample, the time decay weight=0.7×0.5=0.35.

[0110] Assume that the preset maximum number of samples is 200, and the current negative sample pool has 50 samples, then the sample capacity ratio=50 / 200=0.25. The product of the difference between the preset baseline learning rate and the sample capacity ratio is used as the adaptive learning rate. Set the preset baseline learning rate to 0.1, then the adaptive learning rate=0.1×(1-0.25)=0.075. Assume that the classification accuracy was 92% a week ago and is now 90%, the time interval is 7 days, then the accuracy change ratio=(92%-90%) / 7=0.286% / day. The product of the preset baseline update threshold and the accuracy change ratio is used as the dynamic update threshold. Set the preset baseline update threshold to 10, then the dynamic update threshold=10×0.286%=2.86.

[0111] Calculate the sum of the time decay weights of all failure samples, assume it is 15.5. When the sum of the time decay weights 15.5 is greater than the dynamic update threshold 2.86, use the adaptive learning rate 0.075 to incrementally train the stratified failure classification model, and obtain an optimized stratified failure classification model.

[0112] The existing system adopts a static batch training mode and cannot timely adapt to new fault modes. When facing a continuously changing cloud environment, the model accuracy will significantly decrease over time and frequent retraining is required. The cross-level sample supplement strategy based on the sample sharing coefficient of the application solves the problem of sample scarcity in a specific level; a cosine similarity-based adjacent level sample selection method is designed to ensure that the supplemented samples have semantic relevance with the target level. The core innovation lies in the introduction of a multi-dimensional adaptive incremental learning mechanism, which captures the timeliness of faults through a time decay factor, quantifies sample importance through a repair failure degree, and dynamically adjusts learning parameters based on sample capacity and accuracy changes to realize an intelligent updating strategy for the model. This method is particularly suitable for scenarios such as network devices and cloud computing platforms that require real-time fault diagnosis, and can significantly improve reliability and service quality.

[0113] In an optional implementation, the basic strategy template includes:

[0114] The first basic strategy template for data-type faults includes basic operations of executing traffic suspension through an Ingress API, executing data backup through a CSI snapshot interface, and executing a secure restart through a Kubernetes API;

[0115] The second basic strategy template for resource-type faults includes basic operations of executing flow limiting and degradation through an HPA interface, executing resource limitation through a Cgroup interface, and executing node eviction through a Kubernetes drain interface;

[0116] The third basic strategy template for dependency-type faults includes basic operations of executing dependency chain traversal through a dependency graph interface and executing a fuse degradation through a Hystrix API;

[0117] The basic strategy template is selected according to the fault type.

[0118] For example, the embodiment describes in detail an adaptive strategy selection method based on fault types, which realizes fast response and recovery of faults by pre-setting basic strategy templates for different fault types.

[0119] Applications deployed in a Kubernetes environment face various types of faults, including data-type faults, resource-type faults, and dependency-type faults. After observing the occurrence of a fault, an administrator needs to determine the fault type based on the fault characteristics and select an appropriate basic strategy template based on the fault type to execute fault recovery operations. The basic strategy template in the embodiment is a template set composed of multiple basic operations for a specific fault type.

[0120] The first base strategy template for data-type failures includes three core operations. When a database persistent volume is detected to be corrupted or data is inconsistent, the first operation is to perform traffic suspension through the Ingress API. In the specific implementation, the Kubernetes Ingress API is called to modify the routing rules, temporarily routing user requests to a fault prompt page to prevent new requests from continuing to write to the damaged database. For example, the “kubectl patch ingress database-ingress -p '{"spec":{"rules":[{"http":{"paths":[{"path":" / api","pathType":"Prefix","backend":{"service":{"name":"maintenance-service","port":{"number":80}}}}]}}]}}'” command can be used to redirect traffic originally directed to the database service to the maintenance service.

[0121] After traffic suspension, data backup is performed through the CSI snapshot interface. The Kubernetes Container Storage Interface (CSI) snapshot function is called to create a persistent volume data snapshot for subsequent data recovery. The implementation is to create a VolumeSnapshot resource, such as “kubectl create -f snapshot.yaml”, where snapshot.yaml defines the data volume snapshot configuration, including snapshot name, data source, and other information. The data backup operation needs to be performed in a timely manner before the damage spreads, and the backup timeout is usually set to 30 seconds to ensure that critical data protection can be completed even under high load.

[0122] After completing the data backup, a safe restart operation is performed through the Kubernetes API, which recreates the Pod instance to make the application load the backup data or repair minor data inconsistencies. The implementation is to call the “kubectl delete pod database-pod-name” and “kubectl apply -f database-deployment.yaml” command sequences to ensure that the data service is restarted in a healthy state. In the case of multiple database instances, the rolling restart parameter maxUnavailable=25% is set to ensure service availability.

[0123] The second base strategy template for resource-type failures also contains three core operations. When a resource-type failure is detected, such as CPU usage exceeding 90%, memory usage exceeding 85%, or disk I / O saturation, first perform throttling degradation through the HPA (HorizontalPod Autoscaler) interface. Create or modify HPA resources to adjust the number of application instances and distribute load pressure. For example, for CPU-intensive applications, call “kubectl apply -f hpa-config.yaml” to create an HPA configuration, set the CPU target utilization to 60%, and the maximum number of replicas to 10, allowing automatic scaling based on load.

[0124] When throttling degradation cannot immediately alleviate resource pressure, perform resource limitation through the Cgroup interface. Modify the resource configuration of the Pod to limit the upper limit of CPU and memory usage of the Pod, preventing a single service from excessively consuming cluster resources. The specific implementation is to call “kubectl patch deployment resource-intensive-app-p '{"spec" :{ "template" :{ "spec" : { "containers" :[{ "name" : "main-container", "resources":{ "limits" :{ "cpu" : "200m", "memory" : "512Mi"}}}]}}}}'”, and adjust the resource upper limit of the problematic container to a lower safe value.

[0125] If the resource pressure is still severe, perform node eviction operations through the Kubernetes drain interface. Identify the Node node with resource anomalies and safely migrate the Pod from the node to a healthy node. The implementation is to call “kubectl drain node-name --ignore-daemonsets --delete-local-data” to ensure safe migration of workloads on the node. In actual application, the node eviction operation will send a SIGTERM signal to the application 15 seconds in advance, giving the application enough time to complete state saving and connection closing.

[0126] The third base strategy template for dependency failure includes two core operations. When a service call timeout, third-party API response exception, or other dependency failure is detected, the dependency chain traversal is first performed through the dependency graph interface. The service dependency graph is queried to identify the source of the failure and the range of affected services. The implementation is to call the dependency graph API "GET / api / v1 / dependencies / service-name" to obtain the complete dependency relationship list and analyze the dependent services that have failed and their impact links. In actual application, the traversal depth limit is set to 5 layers to avoid excessive analysis causing recovery delay.

[0127] After determining the scope of the failure, the Hystrix API is executed to perform the circuit breaker degradation operation. The circuit breaker strategy is configured for the failed dependent service to block unstable calls and return a degraded result or call the backup service. The specific implementation is to call the Hystrix configuration API to set the circuit breaker parameters, such as setting the error rate threshold to 50%, the waiting time after circuit breaking to 5000 milliseconds, and the sampling window to 10 requests. This ensures that when the dependent service fails, abnormal calls can be quickly cut off to keep the core functions available.

[0128] After identifying the fault type, the corresponding base strategy template is selected according to the fault type, and the fault classification result is associated with the pre-defined strategy template through the strategy mapping mechanism. A strategy mapping table is maintained, where the key is the fault type identifier and the value is the corresponding base strategy template reference. For example, when the XGBoost classifier outputs the fault type as "DataCorruption", the corresponding first base strategy template is obtained by querying the strategy mapping table; when the fault type is "ResourceExhaustion", the second base strategy template is obtained; when the fault type is "DependencyFailure", the third base strategy template is obtained. The strategy selection logic is executed through the StrategySelector component, which receives the fault type parameter and returns the corresponding strategy template instance. For example, when "checksumfailed" appears in the database Pod log and the I / O error rate is higher than the threshold, the fault is classified as a data type fault, and the first base strategy template is automatically selected; when the node CPU usage continuously exceeds 90%, the fault is classified as a resource type fault, and the second base strategy template is selected. The urgency of the fault is also considered in the strategy selection process. For faults with a severity of "Critical" level, the accelerated execution mode of the strategy template is enabled to shorten the operation interval time and improve the response speed of the repair.

[0129] The application establishes specialized basic strategy templates for different fault types, and realizes precise execution of fault repair through combined calling of API interfaces. The template design improves the reusability of repair strategies, while ensuring the standardization and controllability of repair operations. It can take corresponding recovery measures for different types of faults, significantly improving service availability and fault recovery efficiency.

[0130] In an optional embodiment, combining the basic strategy templates based on the controller type feature and the service level feature to form a combined repair strategy includes:

[0131] When the controller type feature is StatefulSet and the fault type is data fault, the operations in the first basic strategy template are sorted based on historical repair success rates, the traffic suspension operation and the data backup operation with the highest repair success rates are selected, the storage volume reservation operation is added based on storage volume state evaluation, the image switching operation is added based on available image version evaluation, and a first combined strategy is generated;

[0132] The CPU utilization and memory usage data are normalized to generate resource usage features. When the controller type feature is Deployment and the fault type is resource fault, the resource usage features are used to determine the flow limiting degradation threshold, select the flow limiting degradation operation, and add the Pod restart operation based on Pod restart history evaluation, increase the number of replica expansion operations based on cluster resource margin evaluation, and increase the connection pool fuse parameter configuration based on service call volume evaluation to generate a second combined strategy;

[0133] When the service level feature is higher than the preset level threshold and the fault type is resource fault, the limit parameters of the resource limit operation are selected based on multi-dimensional resource profiling, and the optimal migration target node is selected based on the target node health score to generate a third combined strategy;

[0134] Different execution priorities are set for the first combined strategy, the second combined strategy, or the third combined strategy according to different service level features.

[0135] For example, when the controller type is StatefulSet and the fault type is data fault, the repair operations and their success rates for this type of fault are first extracted from the historical repair record database. For example, for a StatefulSet controller of a certain database service, the historical records show that the success rate of the traffic suspension operation is 92.3%, the success rate of the data backup operation is 87.6%, the success rate of the storage volume reconstruction operation is 65.2%, and the success rate of the container restart operation is 58.7%. Based on these data, the operations in the first basic strategy template are sorted, and the traffic suspension and data backup operations with the highest success rates are preferentially selected.

[0136] Get the current storage volume state information through API calls, including storage volume health status, data integrity check results, and mounting status. According to the preset storage volume evaluation rules, calculate the storage volume retention necessity score. When the score exceeds the threshold of 75 points (out of 100 points), add the storage volume retention operation in the combined strategy to ensure that the faulty data is not cleared, facilitating subsequent analysis. At the same time, retrieve the image repository to obtain the list of available application image versions, and select the latest image version with a score not less than 80 points based on the version stability score (calculated through historical running data). If there is a qualified image, add the image switching operation in the combined strategy and specify the target version. Through the above steps, the first combined strategy is generated, containing traffic suspension, data backup, storage volume retention, and image switching operations, each of which contains detailed parameter configurations.

[0137] For the scenario of resource type failure of Deployment controller, first collect the CPU utilization and memory usage data in the recent period (such as 30 minutes). Taking a web service as an example, the original data shows that the CPU utilization fluctuates between 60%-95%, and the memory usage fluctuates between 75%-88%. Normalize these data using the maximum and minimum value normalization method to map the value range to the 0-1 interval to obtain the normalized resource usage feature values. Based on these feature values, determine the trigger threshold of traffic limiting degradation. For example, when the normalized CPU usage exceeds 0.85 and the normalized memory usage exceeds 0.8, trigger traffic limiting, and set the degradation ratio to 30%, i.e. only receive 70% of the normal request amount.

[0138] At the same time, analyze the Pod restart history record, and when it is found that the Pod restarts less than 3 times in the past 24 hours and the service recovers normally after each restart, add the Pod restart operation in the second combined strategy. Also, get the resource surplus data of each node in the cluster through the cluster resource scheduler API, and calculate the upper limit of the resources available for expansion. When the node CPU allocatable resource exceeds 25% of the total and the memory allocatable resource exceeds 20%, add the replica expansion operation in the strategy, and calculate the optimal expansion number as 50% of the current number of replicas (rounded up). Based on the service recent average requests per second, average response time, and other service call volume data, configure the connection pool fuse parameters, including setting the maximum connection number to 85% of the current value, setting the timeout threshold to 200% of the average response time, and setting the circuit breaker closing time to 30 seconds. Through these operations, the second combined strategy is formed.

[0139] For scenarios where the service level feature is higher than the preset level threshold (such as SLA availability requirement ≥ 99.9%) and the fault type is resource fault, the resource restriction operation parameter is selected based on the multi-dimensional resource portrait. The resource portrait includes CPU usage history distribution, memory occupation trend, disk IO frequency, network throughput and other dimensions. Taking a key microservice as an example, the resource usage mode is analyzed, and the CPU limit is determined to be 150% of the request value, and the memory limit is determined to be 125% of the request value, to avoid performance fluctuations caused by excessive resource competition.

[0140] Meanwhile, the health score of each node in the cluster is calculated, and the score indicators include node resource usage, recent fault frequency, network delay, disk health status, etc. For example, node A scores 85 points, node B scores 92 points, and node C scores 78 points. The node B with the highest score is selected as the migration target node, and precise migration parameters are set in the third combination strategy, including migration time window, resource reservation ratio and tolerance interruption time, etc.

[0141] According to the service level feature, different combination strategies are set to have execution priorities. The third combination strategy of high-level service (SLA ≥ 99.99%) has the highest execution priority "P0", the first combination strategy of medium-level service (99.9% ≤ SLA < 99.99%) has "P1" priority, and the second combination strategy of low-level service (SLA < 99.9%) has "P2" priority. The priority determines the execution order of the strategy in the case of resource competition, ensuring that the key business is repaired first.

[0142] The strategy combination mechanism based on the controller type and the service level of the application realizes the scenario-based customization of the repair strategy. Through multi-dimensional evaluation and parameter adjustment, the accuracy and adaptability of the repair strategy are improved, and the priority mechanism based on the service level ensures the priority protection of the key business.

[0143] In an optional implementation, constructing a service dependency graph for the combination repair strategy optimization includes:

[0144] A service dependency graph is constructed, which includes service node information, service call relationship edge information and storage dependency edge information. The service node information includes Pod name, controller type feature, service level feature, the service call relationship edge information records the direction and error rate data of service request call, and the storage dependency edge information records the association relationship between the service and the storage volume declared by it;

[0145] Based on the service dependency graph, the combined repair strategy is optimized, the error rate data propagation path in the service call relationship edge information is analyzed, the storage volume access state in the storage dependency edge information is analyzed, and the fault root cause and propagation link are determined in combination with the resource use characteristics; based on the controller type characteristics and service level characteristics in the service node information, the fault influence range is determined, and on the premise of maintaining the service level priority constraint, the execution order of the first combined strategy, the second combined strategy or the third combined strategy is optimized according to the fault root cause and propagation link, and an optimized repair strategy is generated.

[0146] Figure 2 To construct the service dependency graph for the combined repair strategy optimization process, exemplary, in the process of constructing the service dependency graph, first, the service node information is collected, including the Pod name, the controller type characteristics and the service level characteristics. The Pod name is directly obtained from the cluster management, such as "payment-service-5f7d9c5b5b-2xvzp"; the controller type characteristics include the deployment mode (Deployment, StatefulSet, DaemonSet, etc.), for example, "payment-service" belongs to "Deployment" type controller; the service level characteristics are divided into three levels of key service, core service and general service according to the service importance, for example, the payment service is marked as "key service" and the log service is marked as "general service".

[0147] The construction of service call relationship edge information is realized by analyzing the telemetry data of service mesh. The source service, target service, request success rate, error type and other information are extracted from the request log collected by the service mesh agent. For example, the edge information of the order service calling the payment service is recorded as: source node "order-service", target node "payment-service", request total amount 1000 times / minute, error rate 4.5%, and error type mainly "timeout error". The direction of the edge indicates the request flow direction, from the caller to the callee.

[0148] The storage dependency edge information is obtained by analyzing the volume mounting configuration declared by the Pod. The volume declaration information in the Pod configuration is extracted to establish the association between the Pod and the storage volume. For example, the storage dependency edge is established between the database service Pod "mysql-0" and the persistent volume "mysql-pv-data", and the access mode is recorded as "ReadWriteOnce" and the state is "Bound".

[0149] After constructing the complete service dependency graph, the combination repair strategy is optimized based on the graph. The optimization process first analyzes the error rate data propagation path of the service call relationship edge, and identifies the abnormal error rate propagation direction. For example, when it is found that the payment service has a 60% error rate, and the order service calling the payment service has an error rate of 45%, and the inventory service has an error rate of 50%, according to the error rate transmission link, it is determined that the payment service is the root cause of the fault.

[0150] Storage dependency edge information analysis involves checking the storage volume access state, such as "Available", "Bound" or "Pending". When an abnormal state of the storage volume is detected, such as the state of "mysql-pv-data" changing to "Failed", the associated database service is marked as a potential fault source. At the same time, combined with resource usage characteristics, including CPU usage increasing from normal 30% to 95%, memory usage increasing from 50% to 88%, network traffic increasing by 200% and other abnormal patterns, the fault root cause and its propagation link are further confirmed.

[0151] When determining the scope of the fault, the controller type characteristics of the service node are used to evaluate the fault diffusion potential. For example, a database service fault managed by a StatefulSet controller causes data consistency problems, while a stateless service fault managed by a Deployment controller has a relatively controllable impact. Service level characteristics are used to determine the degree of business impact, and the priority of key service faults is higher than that of general services.

[0152] When generating the optimized repair strategy, the service level priority constraint is maintained, i.e. key services are prioritized over core services, and core services are prioritized over general services. For the fault root cause and propagation link that has been determined, the execution order of the first combination strategy (for Pod-level faults), the second combination strategy (for Node-level faults) or the third combination strategy (for cluster-level faults) is optimized.

[0153] For example, when it is identified that the memory leak in the payment service (key service) Pod causes multiple service call failures, the first combination strategy is optimized to execute the Pod restart operation on the payment service first, rather than handling the affected order service at the same time. The specific execution command is "kubectl rollout restart deployment payment-service", and after restarting, the memory usage is reduced to below normal 40%, and the error rate is restored to below 0.5%.

[0154] When a storage failure is found to cause multiple service exceptions, the strategy combination order is adjusted, such as first performing the storage repair operation "kubectl patch pv mysql-pv-data -p '{"spec":{"claimRef":null}}'" to release the storage volume exception binding, and then performing the rescheduling of the related Pod. During the entire repair process, the service state and error rate information in the service dependency graph are continuously updated, and the priority and combination order of the repair strategy are adjusted in real time.

[0155] The optimized repair strategy is converted into a sequence of Kubernetes API calls. For the Pod restart operation, a "DELETE / api / v1 / namespaces / {namespace} / pods / {name}" request is constructed, and the "gracePeriodSeconds" parameter is set to 30 seconds to ensure that the application has enough time to complete state saving. For the node scaling operation, by modifying the deployment configuration, a "PATCH / apis / apps / v1 / namespaces / {namespace} / deployments / {name}" request is sent to increase the "spec.replicas" value, and the "maxSurge" parameter is set to ensure smooth scaling. For the node eviction operation, a "POST / api / v1 / nodes / {name} / drain" request is constructed, and the "--delete-local-data" and "--ignore-daemonsets" flags are configured to ensure safe migration of Pods on the node to healthy nodes. A transactional execution mode is adopted, with timeout and retry mechanisms set for each API operation, and when a single operation fails, the executed repair steps are automatically rolled back to ensure cluster state consistency. After execution, repair operation logs and execution results are recorded for subsequent effect evaluation and model optimization.

[0156] Traditional Kubernetes fault repair methods usually adopt isolation strategies, lack understanding of complex dependencies between services, and may cause chain reactions when repairing a component, which damages other services that depend on it. The invention introduces a multi-dimensional service dependency graph technology, breaking through the limitations of traditional single-dimensional topology graphs, and integrating three types of key information: service node attributes (including controller types and business levels), service call relationships (including directions and error rates), and storage dependency relationships. The core innovation is the graph-based strategy optimization mechanism, which identifies fault causes by analyzing error rate propagation paths, evaluates data impact based on storage access states, and determines priorities based on service attributes. Unlike existing technologies, the invention realizes repair strategy optimization considering global dependency constraints, ensuring that repair operations are executed in the optimal order in a complex microservice environment.

[0157] In a second aspect, the present application provides an electronic device, comprising:

[0158] a processor;

[0159] a memory for storing processor-executable instructions;

[0160] wherein the processor is configured to invoke the instructions stored in the memory to perform the method described above.

[0161] In a third aspect, the present application provides a computer-readable storage medium having stored thereon computer program instructions, which when executed by a processor, implement the method described above.

[0162] The present application can be a method, apparatus, system, and / or computer program product. Computer program products can include computer-readable storage media having computer-readable program instructions loaded thereon for performing various aspects of the present application.

[0163] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for part or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application.

Claims

1. An intelligent fault self-healing method based on multi-dimensional indicators, characterized by: include: Collect Pod restart data through the Kubernetes API, CPU utilization and memory usage data through Prometheus, collect and parse file corruption keyword data through the ELK Stack, and collect error rate data through custom application monitoring to form multi-source data. Perform feature engineering on the multi-source data, including feature encoding of the controller type to which the Pod belongs to obtain controller type features, feature encoding of the service level in the Pod label to obtain service level features, and generating a multi-dimensional feature vector, specifically including: performing unique hot encoding on the controller type to which the Pod belongs to generate controller type features, normalizing the CPU utilization and memory utilization data to generate resource usage features, performing raw counting on the Pod restart count data to generate restart count features, performing Boolean feature conversion on the file damage keyword data to generate log anomaly features, ordinal encoding of the service level in the Pod label to generate service level features, the service level is divided into multiple levels according to the business importance from high to low, and the database storage type in the Pod storage volume declaration configuration is Boolean feature conversion generates storage volume features, and time window statistics are performed on the error rate data to generate service call features; the contribution of each feature to fault classification is calculated based on information gain, the correlation between features is calculated based on the Pearson correlation coefficient, and a preferred feature set is obtained based on the contribution and correlation; based on the preferred feature set, the controller type feature and the service level feature are cross-combined to generate service attribute features, the resource usage feature and the restart number feature are weightedly combined to generate stability features, the log anomaly feature and the storage volume feature are correlated and combined to obtain initial dependency features, and the service call feature and the initial dependency feature are correlated and combined to generate enhanced dependency features; the service attribute features, the stability features, and the enhanced dependency features are spliced ​​to generate a multi-dimensional feature vector; Inputting the multi-dimensional feature vector into the XGBoost classifier to classify the fault to obtain a fault type; selecting a basic strategy template according to the fault type; Combining the basic policy templates based on the controller type characteristics and the service level characteristics to form a combined repair strategy; constructing a service dependency graph to optimize the combined repair strategy; The optimized combined repair strategy is sent to the Kubernetes control plane, and the corresponding repair operation is performed by calling the Kubernetes API. The repair operation includes one or more combinations of Pod restart operation, node expansion operation, and node eviction operation.

2. The method according to claim 1, characterized in that Inputting the multi-dimensional feature vector into the XGBoost classifier for fault classification includes: Performing wavelet transform on the resource usage features to obtain time-frequency features, performing BERT encoding on the log anomaly features to obtain semantic features, concatenating the time-frequency features and the semantic features into the multi-dimensional feature vector to form an enhanced feature vector; performing mean-variance normalization on the enhanced feature vector to obtain a standardized feature vector; Perform sample balancing on historical fault annotated data to obtain training data, and select the training data based on annotation consistency and repair effect to form a core training set; The core training set is stratified according to the fault type and the service level feature, and the standardized feature vector is used to train the XGBoost classifier to obtain a hierarchical fault classification model; when the repair strategy corresponding to the fault type fails to execute, the layer to which the failure case belongs is determined, and the failure case is added to the negative sample pool of the corresponding layer; the hierarchical fault classification model of the corresponding layer is incrementally trained based on the time decay weights of the samples in the negative sample pool; Calculate the predicted probability distribution of the XGBoost classifier to obtain the uncertainty index of the sample classification, and manually confirm the samples whose uncertainty index exceeds the preset index threshold and add them to the core training set; An evaluation index system including accuracy, recall rate, and F1 score is constructed to evaluate the XGBoost classifier. When the evaluation index is lower than the preset evaluation threshold, the model parameters are optimized based on the consistency between the predicted results and the actual repair effect.

3. The method according to claim 2, characterized in that The core training set is trained in layers and incremental training is performed based on failed samples. The core training set is allocated according to fault type and service level to obtain an initial stratified sample set. The sample sharing coefficient is calculated by calculating the ratio of the number of samples in each initial stratified sample set to the maximum number of samples in all initial stratified sample sets. When the sample sharing coefficient is less than a preset sample threshold, the cosine similarity between the adjacent level samples and the current level class center is calculated. The adjacent level samples whose cosine similarity is greater than the preset similarity threshold are added to the current level sample set to form an expanded stratified sample set. The samples in the expanded stratified sample set are trained to obtain a stratified fault classification model. Receive failed samples from a negative sample pool, calculate a time decay factor based on the time interval of the failed samples, calculate a degree of repair failure based on the ratio of the number of retries of the failed samples during the repair process to a preset maximum number of retries, calculate a sample importance coefficient based on the degree of repair failure, and use the product of the time decay factor and the sample importance coefficient as a time decay weight; The ratio of the number of samples in the negative sample pool to the preset maximum number of samples is calculated to obtain the sample capacity ratio, and the product of the difference between the preset benchmark learning rate and the sample capacity ratio is used as the adaptive learning rate; the ratio of the classification accuracy change to the time interval is calculated to obtain the accuracy change ratio, and the product of the preset benchmark update threshold and the accuracy change ratio is used as the dynamic update threshold; when the sum of the time decay weights is greater than the dynamic update threshold, the hierarchical fault classification model is incrementally trained using the adaptive learning rate to obtain an optimized hierarchical fault classification model.

4. The method according to claim 1, wherein The basic policy template includes: A first basic policy template for data-type failures, which includes basic operations such as traffic suspension through the Ingress API, data backup through the CSI snapshot interface, and safe restart through the Kubernetes API; The second basic policy template for resource-based failures includes basic operations such as current limiting and degradation through the HPA interface, resource restriction through the Cgroup interface, and node eviction through the Kubernetes drain interface. The third basic policy template for dependency-type failures includes basic operations of performing dependency chain traversal through the dependency graph interface and performing circuit breaking and degradation through the Hystrix API.

5. The method according to claim 4, characterized in that Combining the basic policy templates based on the controller type characteristics and the service level characteristics to form a combined repair policy includes: When the controller type feature is StatefulSet and the fault type is a data fault, the operations in the first basic policy template are sorted based on the historical repair success rate, and the traffic suspension operation and data backup operation with the highest repair success rate are selected. The storage volume retention operation is added based on the storage volume status evaluation, and the image switching operation is added based on the available image version evaluation to generate a first combined policy; When the controller type feature is Deployment and the fault type is a resource-based fault, the current limiting degradation threshold is determined based on the resource usage feature, the current limiting degradation operation is selected, and the Pod restart operation is added based on the Pod restart history evaluation, the number of replica expansion operations is increased based on the cluster resource margin evaluation, and the connection pool fuse parameter configuration is increased based on the service call volume evaluation to generate a second combination strategy; When the service level characteristic is higher than the preset level threshold and the fault type is a resource-based fault, the restriction parameters of the resource restriction operation are selected based on the multi-dimensional resource profile, and the optimal migration target node is selected based on the target node health score to generate a third combination strategy; According to different service level characteristics, different execution priorities are set for the first combination strategy, the second combination strategy or the third combination strategy.

6. The method according to claim 5, characterized in that Building a service dependency graph to optimize the combined repair strategy includes: Construct a service dependency graph, which includes service node information, service call relationship edge information, and storage dependency edge information. The service node information includes Pod name, controller type characteristics, and service level characteristics. The service call relationship edge information records the direction and error rate data of request calls between services. The storage dependency edge information records the association between services and the storage volumes they declare to use. Based on the service dependency graph, the combined repair strategy is optimized, the error rate data propagation path in the service call relationship side information is analyzed, the storage volume access status in the storage dependency side information is analyzed, and the root cause of the fault and the propagation link are determined in combination with the resource usage characteristics; based on the controller type characteristics and service level characteristics in the service node information, the fault impact range is determined, and under the premise of maintaining the service level priority constraint, the execution order of the first combined strategy, the second combined strategy or the third combined strategy is optimized according to the root cause of the fault and the propagation link to generate an optimized repair strategy.

7. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the method according to any one of claims 1 to 6.

8. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Adaptive operation and maintenance root cause positioning method and system based on deep learning

    CN119691576A

  • Network anomaly detection and automatic processing method and system, storage medium and equipment

    CN119892476A