Methods, apparatus, electronic devices and storage media for self-healing fault detection of safe atomic capabilities

By using real-time analysis and risk prediction models, combined with anomaly indicators and fault levels, targeted fault self-healing strategies are developed, which solves the problem of insufficient efficiency and accuracy in the detection of safety atomic capabilities in existing technologies, and achieves stable and safe operation of the system.

CN121435222BActive Publication Date: 2026-06-30BEIJING YIAN TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING YIAN TECHNOLOGY CO LTD
Filing Date
2025-12-18
Publication Date
2026-06-30

AI Technical Summary

Technical Problem

Existing technologies suffer from low efficiency and insufficient accuracy in detecting faults in secure atomic capabilities, making it difficult to detect potential faults in complex systems in a timely manner, which affects system stability and security.

Method used

By acquiring the target instance's metric data, we can analyze and train risk prediction models, use anomaly and risk indicators to determine the fault type and level, and formulate targeted fault self-healing strategies, including traffic forwarding and instance reconstruction strategies.

Benefits of technology

It enables real-time monitoring and accurate prediction of safety atomic capability faults, improves the timeliness of fault detection and self-healing capability, and ensures stable system operation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121435222B_ABST
    Figure CN121435222B_ABST
Patent Text Reader

Abstract

This application relates to a method, apparatus, electronic device, and storage medium for self-healing fault detection in secure atomic capabilities, applied in the field of system security technology. The method includes: acquiring indicator data of a target instance; analyzing the indicator data to determine current fault information; inputting the indicator data into a risk prediction model for fault prediction to obtain predicted fault information, the predicted fault information including fault probability and risk indicators; if the current fault information includes anomaly indicators and / or the fault probability is greater than a preset fault probability, then determining a target fault type based on the anomaly indicators and / or the risk indicators; determining a fault level based on the target fault type; and determining a fault self-healing strategy based on the fault level. This application improves the timeliness and accuracy of self-healing fault detection in secure atomic capabilities.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of system security, and in particular to a method, apparatus, electronic device and storage medium for self-healing of fault detection in secure atomic capabilities. Background Technology

[0002] In an era of rapid information technology development, system security and stability are becoming increasingly important. Security atomic capabilities (such as vulnerability scanning, traffic scrubbing, and threat intelligence detection) serve as fundamental units for system operation, and their proper functioning plays a crucial role in the stable operation of the entire system. With the continuous emergence of various complex systems, higher demands are placed on the reliable operation of security atomic capabilities. Failure of a security atomic capability may lead to the malfunction of some system functions, affecting normal business operations, and may even trigger security risks, causing serious consequences such as data breaches and system paralysis, resulting in huge losses for enterprises and users. Therefore, fault detection and self-healing of security atomic capabilities have become an important research direction for ensuring stable system operation, as they relate to whether the system can operate continuously and efficiently in complex and ever-changing environments.

[0003] Currently, a common approach to handling safety atomic capability faults is to rely on regular manual system inspections. Maintenance personnel manually check system log files and monitoring metrics for anomalies, and then conduct further investigation and repair when problems are found. Another approach is to set simple threshold alarms; when certain system metrics exceed preset thresholds, an alert is issued to remind maintenance personnel. These methods can detect and handle some faults to a certain extent, but each has its limitations. Manual inspections are inefficient, unable to monitor the system's operational status in real time, and may lead to delayed fault detection. Threshold alarms can only detect some simple faults and are difficult to accurately identify complex or potential faults, failing to meet the needs of modern complex systems for safety atomic capability fault detection and self-healing. Summary of the Invention

[0004] To improve the timeliness and accuracy of self-healing for safe atomic capability fault detection, this application provides a method, apparatus, electronic device, and storage medium for self-healing for safe atomic capability fault detection.

[0005] In a first aspect, this application provides a self-healing method for detecting faults in safe atomic capabilities, employing the following technical solution:

[0006] A self-healing method for detecting faults in safe atomic capabilities includes:

[0007] Obtain the target instance's metric data;

[0008] The aforementioned indicator data is analyzed to determine the current fault information;

[0009] The indicator data is input into the risk prediction model to predict the fault, and the predicted fault information includes the fault probability and risk indicators.

[0010] If the current fault information includes abnormal indicators and / or the fault probability is greater than the preset fault probability, then the target fault type is determined based on the abnormal indicators and / or the risk indicators.

[0011] Determine the fault level based on the target fault type;

[0012] A fault self-healing strategy is determined based on the fault level.

[0013] By adopting the above technical solutions and analyzing indicator data in real time, current fault information can be detected in a timely manner. Fault prediction can be performed through risk prediction models, allowing for the prediction of impending faults. Real-time fault analysis and prediction improve the real-time performance and diversity of fault monitoring. Based on abnormal and risk indicators, the target fault type can be further determined, thereby determining the fault level. More targeted fault self-healing strategies can be adopted based on the fault level, improving the timeliness and accuracy of fault detection and self-healing in safety atomic capabilities.

[0014] Optionally, before inputting the indicator data into the risk prediction model for fault prediction, the method further includes:

[0015] Acquire historical operational data, which includes historical fault data and historical fault precursor data;

[0016] The risk prediction model is constructed by training the preset model based on the historical operating data.

[0017] By adopting the above technical solution, a risk prediction model is constructed by training the preset model with historical fault data and historical fault precursor data. In this way, the risk prediction model can more accurately predict future faults based on current indicator data.

[0018] Optionally, determining the target fault type based on the anomaly indicators and / or the risk indicators includes:

[0019] Acquire historical fault data, which includes fault indicators and fault types;

[0020] By exhaustively combining the fault indicators in the historical fault data, multiple indicator combinations are obtained;

[0021] Statistical analysis of the correlation probability between each of the aforementioned indicator combinations and various fault types;

[0022] If the correlation probability between the indicator combination and the fault type exceeds a preset correlation probability, then the indicator combination and the fault type are correlated.

[0023] The correlation between the indicators and the fault types is determined based on the relationship between the indicator combinations and the fault types.

[0024] The current fault type is determined based on the abnormal indicators and the fault correlations of the indicators;

[0025] The predicted fault type is determined based on the risk indicators and the fault correlations of the indicators.

[0026] The target fault type is determined based on the current fault type and the predicted fault type.

[0027] By adopting the above technical solution, and by statistically analyzing the correlation probability between various indicator combinations and various fault types, the correlation between indicator combinations and fault types can be determined. This facilitates the rapid identification of target fault types based on current abnormal and risk indicators, thereby improving the efficiency of fault detection.

[0028] Optionally, determining the predicted fault type based on the risk indicators and the fault correlation of the indicators includes:

[0029] The risk indicator is used to match a first association probability from the fault associations of the indicator;

[0030] Calculate the type probability of each fault type based on the fault probability and the first association probability;

[0031] The fault type whose type probability exceeds the preset type probability is determined as the predicted fault type.

[0032] By adopting the above technical solution, the probability of each fault type is determined by comprehensively considering the predicted fault probability and the first correlation probability between the risk index and the fault type, thereby determining the predicted fault type and improving the reliability of fault prediction.

[0033] Optionally, determining the fault self-healing strategy based on the fault level includes:

[0034] If the fault level does not exceed the preset fault level, the fault self-healing strategy is retrieved from the fault self-healing strategy library based on the target fault type.

[0035] If the fault level exceeds the preset fault level, then the traffic forwarding strategy and the instance reconstruction strategy of the target instance are determined; the fault self-healing strategy is determined based on the traffic forwarding strategy and the instance reconstruction strategy.

[0036] Optionally, determining the traffic forwarding strategy for the target instance includes:

[0037] Determine whether the target instance has a backup instance;

[0038] If the target instance has a backup instance, then the traffic forwarding strategy is determined to be to forward the traffic of the target instance to the corresponding backup instance;

[0039] If the target instance does not have the backup instance, then an optional instance is obtained based on the instance type and business group of the target instance;

[0040] The candidate instances are filtered based on their instance status and instance load to determine the candidate instances;

[0041] The candidate instances are sorted based on historical fault conditions, historical fault recovery times, and historical load stability to determine the sorting result;

[0042] The traffic forwarding strategy is determined based on the sorting results.

[0043] By adopting the above technical solution, if a backup instance exists for the target instance, the traffic of the target instance is directly forwarded to the corresponding backup instance, simplifying the traffic forwarding process; if no backup instance exists for the target instance, candidate instances are determined by instance type, service group, instance status, and instance load, improving the availability of candidate instances; candidate instances are sorted by historical fault conditions, historical fault recovery time, and historical load stability, and the traffic forwarding strategy is determined based on the sorting results, improving the reliability of traffic forwarding.

[0044] Optionally, before determining whether a backup instance exists for the target instance, the method further includes:

[0045] The business importance score is determined based on the business type and service level agreement requirements corresponding to the target instance;

[0046] The fault impact score is determined based on the scale of users covered by the target instance and the number of associated instances.

[0047] The operating load score is determined based on the load fluctuation amplitude and historical failure frequency of the target instance.

[0048] The target instance score is determined based on the business importance score, the fault impact score, and the operational load score.

[0049] If the target instance score is greater than the preset score, then the target instance has a backup instance;

[0050] If the target instance score is less than or equal to the preset score, then the target instance does not have a backup instance.

[0051] By adopting the above technical solution, the target instance score is determined by the business type, service level agreement requirements, user coverage scale, number of associated instances, load fluctuation range, and historical failure frequency of the target instance. Based on the target instance score, it is determined whether to set up a backup instance for the target instance. Instead of setting up backup instances for all instances, instances with higher scores (i.e., higher importance) have higher reliability, and resource waste is reduced.

[0052] Secondly, this application provides a self-healing device for detecting and resolving security atomic capability faults, employing the following technical solution:

[0053] A self-healing device for detecting and resolving faults in secure atomic capabilities, comprising:

[0054] The indicator data acquisition module is used to acquire indicator data for the target instance;

[0055] The fault information determination module is used to analyze the indicator data and determine the current fault information;

[0056] The fault information prediction module is used to input the indicator data into the risk prediction model to predict faults and obtain predicted fault information, which includes fault probability and risk indicators.

[0057] The fault type determination module is used to determine the target fault type based on the abnormal indicators and / or the risk indicators if the current fault information includes abnormal indicators and / or the fault probability is greater than a preset fault probability.

[0058] A fault level determination module is used to determine the fault level based on the target fault type;

[0059] The self-healing strategy determination module is used to determine the fault self-healing strategy based on the fault level.

[0060] By adopting the above technical solutions and analyzing indicator data in real time, current fault information can be detected in a timely manner. Fault prediction can be performed through risk prediction models, allowing for the prediction of impending faults. Real-time fault analysis and prediction improve the real-time performance and diversity of fault monitoring. Based on abnormal and risk indicators, the target fault type can be further determined, thereby determining the fault level. More targeted fault self-healing strategies can be adopted based on the fault level, improving the timeliness and accuracy of fault detection and self-healing in safety atomic capabilities.

[0061] Thirdly, this application provides an electronic device that adopts the following technical solution:

[0062] An electronic device includes a processor coupled to a memory;

[0063] The memory stores a computer program that can be loaded by a processor and executed by the self-healing method for detecting faults in the secure atomic capability as described in any of the first aspects.

[0064] Fourthly, this application provides a computer-readable storage medium, which adopts the following technical solution:

[0065] A computer-readable storage medium storing a computer program capable of being loaded by a processor and executing the self-healing method for detecting faults in the security atomic capability as described in any of the first aspects. Attached Figure Description

[0066] Figure 1 This is a flowchart illustrating a self-healing method for detecting faults in the security atomic capability provided in an embodiment of this application.

[0067] Figure 2 This is a structural block diagram of a self-healing device for detecting faults in the security atomic capability provided in an embodiment of this application.

[0068] Figure 3 This is a structural block diagram of the electronic device provided in the embodiments of this application. Detailed Implementation

[0069] The present application will be further described in detail below with reference to the accompanying drawings.

[0070] This application provides a self-healing method for detecting and resolving security atomic capability faults. This method can be executed by an electronic device, which can be a server or a terminal device. The server can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. The terminal device can be a smartphone, tablet computer, desktop computer, etc., but is not limited to these.

[0071] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0072] Furthermore, the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this article, unless otherwise specified, generally indicates that the preceding and following related objects have an "or" relationship.

[0073] like Figure 1 As shown, a self-healing method for detecting faults in secure atomic capabilities is described in the following main process flow (steps S101 to S106):

[0074] Step S101: Obtain the target instance's metric data.

[0075] The target instance is the instance that needs to be detected for faults. All instances in the system can be used as target instances for fault detection.

[0076] Metrics data for the target instance, such as CPU utilization, memory utilization, network throughput, log data, and heartbeat data, are obtained in real time through methods such as built-in agent collection, business log streaming collection, and API calls + heartbeat detection.

[0077] Step S102: Analyze the indicator data to determine the current fault information.

[0078] The indicator data is analyzed according to preset anomaly judgment rules to determine whether there are any anomalies in the indicator data. If there are anomalies in the indicator data, the target instance and all abnormal indicators (i.e., indicator data with anomalies) are jointly identified as the current fault information. The preset anomaly judgment rules are as follows: if no heartbeat packets are received from the target instance three times in a row, the heartbeat data is abnormal; if the log data includes "network unreachable", the log data is abnormal; if the CPU utilization exceeds 95% and lasts for more than 10 seconds, the CPU utilization is abnormal.

[0079] If all metrics data are normal, the current fault information is that the target instance is not abnormal.

[0080] Step S103: Input the indicator data into the risk prediction model to predict the fault and obtain the predicted fault information, which includes the fault probability and risk indicators.

[0081] In this embodiment, not only can the real-time fault status of the target instance be detected, but the future fault status of the target instance can also be predicted. The indicator data is input into the risk prediction model, and the risk prediction model can predict the future fault probability of the target instance and the risk indicators that may cause the fault to occur based on the current indicator data, that is, predict the fault information.

[0082] Specifically, before inputting the indicator data into the risk prediction model for fault prediction, the method also includes: acquiring historical operating data, which includes historical fault data and historical fault precursor data; and training the preset model based on the historical operating data to construct the risk prediction model.

[0083] In this embodiment, historical fault data and corresponding historical fault precursor data of the target instance are obtained from the database; a preset model (e.g., Transformer model) is trained using the historical fault data and corresponding historical fault precursor data to construct a risk prediction model.

[0084] Step S104: If the current fault information includes abnormal indicators and / or the fault probability is greater than the preset fault probability, then determine the target fault type based on the abnormal indicators and / or risk indicators.

[0085] The target fault type includes the current fault type and / or the predicted fault type.

[0086] If the current fault information includes abnormal indicators, it indicates that the target instance is currently faulty, and the fault type needs to be further determined based on the abnormal indicators.

[0087] If the predicted failure probability is greater than the preset failure probability (preset, not specifically limited here, for example: 70%), then the predicted failure type needs to be further determined based on the risk indicators.

[0088] Specifically, determining the target fault type based on abnormal indicators and / or risk indicators includes: acquiring historical fault data, which includes fault indicators and fault types; combining fault indicators from the historical fault data using an exhaustive method to obtain multiple indicator combinations; calculating the association probability between each indicator combination and various fault types; if the association probability between an indicator combination and a fault type exceeds a preset association probability, then an association exists between the indicator combination and the fault type; determining the indicator-fault association based on the association between the indicator combination and the fault type; determining the current fault type based on abnormal indicators and indicator-fault associations; determining the predicted fault type based on risk indicators and indicator-fault associations; and determining the target fault type based on the current fault type and the predicted fault type.

[0089] In this embodiment, historical fault data is retrieved from the database. An exhaustive method is used to combine the fault indicators in each historical fault data set to obtain multiple indicator combinations. For example, if historical fault data A includes fault indicator 1, fault indicator 2, and fault indicator 3, and historical fault data B includes fault indicator 4 and fault indicator 5, then the indicator combinations include (fault indicator 1), (fault indicator 2), (fault indicator 3), (fault indicator 1, fault indicator 2), (fault indicator 1, fault indicator 3), (fault indicator 2, fault indicator 3), (fault indicator 1, fault indicator 2, fault indicator 3), and (fault indicator 4). (Fault Indicator 5), (Fault Indicator 4, Fault Indicator 5); The fault type corresponding to a combination of indicators is the fault type in the historical fault data corresponding to that fault indicator. The probability of association between an indicator combination and a fault type = the number of times the indicator combination corresponds to that fault type / the number of times the indicator combination appears in all historical fault data; If the probability of association between an indicator combination and a fault type exceeds the preset probability of association (pre-set, for example: 60%), then the indicator combination is associated with that fault type; The association between all indicator combinations and fault types (including the probability of association) is jointly determined as the indicator-fault association.

[0090] All abnormal indicators are combined into a current indicator combination. The current fault type is then matched with the fault association from the indicator fault association (there may be multiple types, and all matched fault types are retained). By analyzing the risk indicators and the fault association, the predicted fault type is determined. The current fault type and the predicted fault type are then combined to determine the target fault type.

[0091] More specifically, determining the predicted fault type based on risk indicators and indicator fault associations includes: matching a first association probability from indicator fault associations using risk indicators; calculating the type probability of various fault types based on fault probabilities and the first association probability; and identifying fault types whose type probabilities exceed preset type probabilities as predicted fault types.

[0092] In this embodiment, all risk indicators are combined into a risk indicator combination. The association probability between the risk indicator combination and the fault type is matched from the indicator fault association, i.e., the first association probability. Since the risk indicator combination may have an association probability with multiple fault types that exceeds the preset association probability, there may be multiple first association probabilities. The type probability of a fault type = fault probability × the first association probability corresponding to the fault type. Fault types whose type probability exceeds the preset type probability (pre-set, for example, 60%) are determined as predicted fault types.

[0093] Step S105: Determine the fault level based on the target fault type.

[0094] Different fault types correspond to different fault levels. The relationship between fault types and fault levels is pre-stored in the database. The fault level is matched from the database according to the target fault type. If the target instance corresponds to multiple target fault types, the highest matched fault level is determined as the fault level of the target instance.

[0095] Step S106: Determine the fault self-healing strategy based on the fault level.

[0096] Specifically, determining the fault self-healing strategy based on the fault level includes: if the fault level does not exceed the preset fault level, then retrieving the fault self-healing strategy from the fault self-healing strategy library based on the target fault type; if the fault level exceeds the preset fault level, then determining the traffic forwarding strategy and the instance reconstruction strategy for the target instance; and determining the fault self-healing strategy based on the traffic forwarding strategy and the instance reconstruction strategy.

[0097] In this embodiment, if the fault level of the target instance does not exceed the preset fault level (pre-set, not specifically limited here, for example: moderate fault), that is, the fault level is low, then there is no need to rebuild the fault instance. It is only necessary to retrieve the corresponding fault self-healing strategy from the fault self-healing strategy library according to the target fault type. For example, if the target fault type is that the vulnerability scanning process is unresponsive, the fault self-healing strategy is to call the systemd command to restart the abnormal process. The fault self-healing strategy can quickly restore the abnormality.

[0098] If the fault level of the target instance exceeds the preset fault level, the target instance needs to be rebuilt. The preset instance rebuild strategy is retrieved from the database. At the same time, during the instance rebuild process, the current traffic of the target instance needs to be distributed to other instances. Therefore, the traffic forwarding strategy of the target instance also needs to be determined. The traffic forwarding strategy and the instance rebuild strategy are jointly determined as the fault self-healing strategy.

[0099] More specifically, determining the traffic forwarding strategy for the target instance includes: determining whether a backup instance exists for the target instance; if a backup instance exists for the target instance, determining the traffic forwarding strategy to forward the traffic of the target instance to the corresponding backup instance; if no backup instance exists for the target instance, obtaining optional instances based on the instance type and service group of the target instance; filtering the optional instances based on the instance status and instance load to determine candidate instances; sorting the candidate instances based on historical fault conditions, historical fault recovery time, and historical load stability to determine the sorting result; and determining the traffic forwarding strategy based on the sorting result.

[0100] In this embodiment, it is first determined whether there is a backup instance for the target instance; if there is a backup instance for the target instance, the traffic forwarding strategy is to forward all traffic of the target instance to the corresponding backup instance.

[0101] If no backup instance exists for the target instance, the instance with the same instance type (e.g., WAF, IPS) and business group (e.g., financial business WAF group) as the target instance will be identified as an optional instance; the optional instance that meets all three conditions of "the instance status is normal operation, the CPU available load rate exceeds the preset load (e.g., 20%), and the memory utilization rate is lower than the preset utilization rate (e.g., 70%)" will be identified as a candidate instance.

[0102] The database stores scoring rules for candidate instances based on historical fault occurrences, historical fault recovery times, and historical load stability. Based on these rules, each candidate instance is scored for its historical fault occurrences (e.g., number of faults in the last 30 days), historical fault recovery times (e.g., average recovery time in the last 30 days), and historical load stability (e.g., CPU utilization fluctuation coefficient in the last 30 days). All scores are summed to obtain the candidate instance score. Candidate instances are then sorted from highest to lowest score to determine the ranking. The traffic capacity of each candidate instance is obtained from the monitoring system. The traffic forwarding strategy proceeds sequentially from the target instance to the candidate instances according to the ranking, until a candidate instance's traffic capacity reaches zero. Then, traffic is forwarded to the next candidate instance. The traffic forwarding process ends when the traffic to the target instance has been completely forwarded.

[0103] Furthermore, before determining whether a backup instance exists for the target instance, the method also includes: determining a business importance score based on the business type and service level agreement requirements corresponding to the target instance; determining a fault impact score based on the user coverage scale and the number of associated instances corresponding to the target instance; determining an operational load score based on the load fluctuation range and historical fault frequency corresponding to the target instance; determining a target instance score based on the business importance score, fault impact score, and operational load score; if the target instance score is greater than a preset score, then a backup instance exists for the target instance; if the target instance score is less than or equal to the preset score, then a backup instance does not exist for the target instance.

[0104] In this embodiment, the database stores rules for determining the importance score of a business based on the business type and service level agreement requirements, rules for determining the fault impact score based on the scale of covered users and the number of associated instances, and rules for determining the operating load score based on the load fluctuation range and historical fault frequency. The database is used to obtain the business type, service level agreement requirements, scale of covered users, number of associated instances, load fluctuation range, and historical fault frequency corresponding to the target instance.

[0105] The target instance is matched with a business importance score from the database based on its business type and service level agreement requirements; a fault impact score is matched with the database based on the user coverage scale and number of associated instances of the target instance; and an operational load score is matched with the database based on the load fluctuation range and historical fault frequency of the target instance. The target instance score is calculated as: Business Importance Score + Fault Impact Score + Operational Load Score. If the target instance score is greater than the preset score (pre-set), a backup instance needs to be set for the target instance, meaning that a backup instance exists for the target instance. If the target instance score is less than or equal to the preset score, then no backup instance exists for the target instance.

[0106] Figure 2 This is a structural block diagram of a self-healing device 200 for detecting faults in a safe atomic capability, provided in an embodiment of this application.

[0107] like Figure 2 As shown, the self-healing device 200 for detecting and resolving security atomic capability faults mainly includes:

[0108] The indicator data acquisition module 201 is used to acquire indicator data of the target instance;

[0109] The fault information determination module 202 is used to analyze the indicator data and determine the current fault information;

[0110] The fault information prediction module 203 is used to input indicator data into the risk prediction model to predict faults and obtain predicted fault information, which includes fault probability and risk indicators.

[0111] The fault type determination module 204 is used to determine the target fault type based on the abnormal indicators and / or risk indicators if the current fault information includes abnormal indicators and / or the fault probability is greater than the preset fault probability.

[0112] Fault level determination module 205 is used to determine the fault level based on the target fault type;

[0113] The self-healing strategy determination module 206 is used to determine the fault self-healing strategy based on the fault level.

[0114] As an optional implementation of this embodiment, the fault information prediction module 203 is specifically used to: acquire historical operating data, including historical fault data and historical fault precursor data, before inputting the indicator data into the risk prediction model for fault prediction; and train the preset model based on the historical operating data to construct the risk prediction model.

[0115] As an optional implementation of this embodiment, the fault type determination module 204 is specifically used to determine the target fault type based on abnormal indicators and / or risk indicators, including: acquiring historical fault data, which includes fault indicators and fault types; combining the fault indicators in the historical fault data using an exhaustive method to obtain multiple indicator combinations; calculating the correlation probability between each indicator combination and various fault types; if the correlation probability between the indicator combination and the fault type exceeds a preset correlation probability, then the indicator combination and the fault type are correlated; determining the indicator fault association based on the correlation between the indicator combination and the fault type; determining the current fault type based on the abnormal indicators and the indicator fault association; determining the predicted fault type based on the risk indicators and the indicator fault association; and determining the target fault type based on the current fault type and the predicted fault type.

[0116] As an optional implementation of this embodiment, the fault type determination module 204 is specifically used to determine the predicted fault type based on risk indicators and indicator fault associations, including: matching a first association probability from indicator fault associations through risk indicators; calculating the type probability of various fault types based on fault probabilities and the first association probability; and determining the fault type with a type probability exceeding a preset type probability as the predicted fault type.

[0117] As an optional implementation of this embodiment, the self-healing strategy determination module 206 is specifically used to determine the fault self-healing strategy based on the fault level, including: if the fault level does not exceed the preset fault level, then retrieve the fault self-healing strategy from the fault self-healing strategy library based on the target fault type; if the fault level exceeds the preset fault level, then determine the traffic forwarding strategy and the instance reconstruction strategy of the target instance; and determine the fault self-healing strategy based on the traffic forwarding strategy and the instance reconstruction strategy.

[0118] As an optional implementation of this embodiment, the self-healing strategy determination module 206 is specifically used to determine the traffic forwarding strategy of the target instance, including: determining whether the target instance has a backup instance; if the target instance has a backup instance, determining the traffic forwarding strategy as forwarding the traffic of the target instance to the corresponding backup instance; if the target instance does not have a backup instance, obtaining optional instances based on the instance type and service group of the target instance; filtering the optional instances based on the instance status and instance load to determine candidate instances; sorting the candidate instances based on historical fault conditions, historical fault recovery time, and historical load stability to determine the sorting result; and determining the traffic forwarding strategy based on the sorting result.

[0119] As an optional implementation of this embodiment, the security atomic capability fault detection self-healing device 200 is further specifically used to determine, before determining whether a backup instance exists for the target instance, include: determining a business importance score based on the business type and service level agreement requirements corresponding to the target instance; determining a fault impact score based on the coverage user scale and the number of associated instances corresponding to the target instance; determining an operational load score based on the load fluctuation amplitude and historical fault frequency corresponding to the target instance; determining a target instance score based on the business importance score, fault impact score, and operational load score; if the target instance score is greater than a preset score, then a backup instance exists for the target instance; if the target instance score is less than or equal to the preset score, then a backup instance does not exist for the target instance.

[0120] In one example, the module in any of the above devices may be one or more integrated circuits configured to implement the above methods, such as one or more application-specific integrated circuits (ASICs), or one or more digital signal processors (DSPs), or one or more field-programmable gate arrays (FPGAs), or a combination of at least two of these integrated circuit forms.

[0121] For example, when modules in a device can be implemented via a processing element scheduler, the processing element can be a general-purpose processor, such as a central processing unit (CPU) or other processor capable of calling programs. Alternatively, these modules can be integrated together as a system-on-a-chip (SOC).

[0122] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the above-described device and module can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0123] Figure 3 This is a structural block diagram of an electronic device 300 provided in an embodiment of this application.

[0124] like Figure 3 As shown, the electronic device 300 includes a processor 301 and a memory 302, and may further include one or more of an information input / output (I / O) interface 303, a communication component 304, and a communication bus 305.

[0125] The processor 301 controls the overall operation of the electronic device 300 to complete all or part of the steps of the aforementioned self-healing method for detecting and resolving security atomic capability faults. The memory 302 stores various types of data to support the operation of the electronic device 300. This data may include, for example, instructions for any application or method operating on the electronic device 300, as well as application-related data. The memory 302 may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as one or more of Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read-Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

[0126] I / O interface 303 provides an interface between processor 301 and other interface modules, such as keyboards, mice, and buttons. These buttons can be virtual or physical. Communication component 304 is used for wired or wireless communication between electronic device 300 and other devices. Wireless communication includes Wi-Fi, Bluetooth, Near Field Communication (NFC), 2G, 3G, or 4G, or a combination thereof. Therefore, the corresponding communication component 304 may include a Wi-Fi component, a Bluetooth component, and an NFC component.

[0127] The electronic device 300 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to execute the self-healing method for detecting security atomic capability faults given in the above embodiments.

[0128] The communication bus 305 may include a path for transmitting information between the aforementioned components. The communication bus 305 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. The communication bus 305 may be divided into an address bus, a data bus, a control bus, etc.

[0129] Electronic device 300 may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital radio receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), and in-vehicle terminals (such as in-vehicle navigation terminals), as well as fixed terminals such as digital TVs and desktop computers, and may also be servers.

[0130] This application also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the above-described self-healing method for detecting and resolving security atomic capability faults.

[0131] The computer-readable storage medium may include various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0132] The terms “comprising,” “including,” or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.

[0133] The above description is merely a preferred embodiment of this application and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this application is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the foregoing application concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features with similar functions claimed in this application.

Claims

1. A self-healing method for detecting faults in safe atomic capabilities, characterized in that, include: Obtain the target instance's metric data; The aforementioned indicator data is analyzed to determine the current fault information; The indicator data is input into the risk prediction model to predict the fault, and the predicted fault information includes the fault probability and risk indicators. If the current fault information includes abnormal indicators and / or the fault probability is greater than the preset fault probability, then the target fault type is determined based on the abnormal indicators and / or the risk indicators. Determine the fault level based on the target fault type; Determine the fault self-healing strategy based on the fault level; The determination of the target fault type based on the abnormal indicators and / or the risk indicators includes: Acquire historical fault data, which includes fault indicators and fault types; By exhaustively combining the fault indicators in the historical fault data, multiple indicator combinations are obtained; The association probability between each of the aforementioned indicator combinations and each of the aforementioned fault types is calculated. The association probability between an indicator combination and a fault type is equal to the number of times the indicator combination corresponds to the fault type / the number of times the indicator combination appears in all the historical fault data. If the correlation probability between the indicator combination and the fault type exceeds a preset correlation probability, then the indicator combination and the fault type are correlated. The correlation between the indicators and the fault types is determined based on the relationship between the indicator combinations and the fault types. The current fault type is determined based on the abnormal indicators and the fault correlations of the indicators; The predicted fault type is determined based on the risk indicators and the fault correlations of the indicators. The target fault type is determined based on the current fault type and the predicted fault type; The step of determining the predicted fault type based on the risk indicators and the fault correlation of the indicators includes: The risk indicator is used to match a first association probability from the fault associations of the indicator; Calculate the type probability of each of the fault types based on the fault probability and the first association probability; The fault type whose type probability exceeds the preset type probability is determined as the predicted fault type; The method for determining a fault self-healing strategy based on the fault level includes: If the fault level does not exceed the preset fault level, the fault self-healing strategy is retrieved from the fault self-healing strategy library based on the target fault type. If the fault level exceeds the preset fault level, then the traffic forwarding strategy and the instance reconstruction strategy of the target instance are determined; the fault self-healing strategy is determined based on the traffic forwarding strategy and the instance reconstruction strategy. The determination of the traffic forwarding strategy for the target instance includes: Determine whether the target instance has a backup instance; If the target instance has a backup instance, then the traffic forwarding strategy is determined to be to forward the traffic of the target instance to the corresponding backup instance; If the target instance does not have the backup instance, then an optional instance is obtained based on the instance type and business group of the target instance; The candidate instances are filtered based on their instance status and instance load to determine the candidate instances; The candidate instances are sorted based on historical fault conditions, historical fault recovery times, and historical load stability to determine the sorting result; The traffic forwarding strategy is determined based on the sorting result. The traffic forwarding strategy is to forward the traffic of the target instance to the candidate instance in sequence according to the sorting result until the traffic value that a candidate instance can handle is 0. Then, the traffic is forwarded to the next candidate instance. When the traffic of the target instance is forwarded, the traffic forwarding process ends. Before determining whether a backup instance exists for the target instance, the method further includes: The business importance score is determined based on the business type and service level agreement requirements corresponding to the target instance; The fault impact score is determined based on the scale of users covered by the target instance and the number of associated instances. The operating load score is determined based on the load fluctuation amplitude and historical failure frequency of the target instance. The target instance score is determined based on the business importance score, the fault impact score, and the operational load score. If the target instance score is greater than the preset score, then the target instance has a backup instance; If the target instance score is less than or equal to the preset score, then the target instance does not have a backup instance.

2. The method according to claim 1, characterized in that, Before inputting the indicator data into the risk prediction model for fault prediction, the method further includes: Acquire historical operational data, which includes historical fault data and historical fault precursor data; The risk prediction model is constructed by training the preset model based on the historical operating data.

3. A self-healing device for detecting faults in safe atomic capabilities, characterized in that, For implementing the method of claim 1, comprising: The indicator data acquisition module is used to acquire indicator data for the target instance; The fault information determination module is used to analyze the indicator data and determine the current fault information; The fault information prediction module is used to input the indicator data into the risk prediction model to predict faults and obtain predicted fault information, which includes fault probability and risk indicators. The fault type determination module is used to determine the target fault type based on the abnormal indicators and / or the risk indicators if the current fault information includes abnormal indicators and / or the fault probability is greater than a preset fault probability. A fault level determination module is used to determine the fault level based on the target fault type; The self-healing strategy determination module is used to determine the fault self-healing strategy based on the fault level.

4. An electronic device, characterized in that, Includes a processor, which is coupled to a memory; The processor is configured to execute a computer program stored in the memory, causing the electronic device to perform the method as described in any one of claims 1 to 2.

5. A computer-readable storage medium, characterized in that, It includes a computer program or instructions that, when run on a computer, cause the computer to perform the method as described in any one of claims 1 to 2.

Citation Information

Patent Citations

  • Fault self-recovery method and device

    CN117370054A

  • Equipment fault alarm processing method and device, electronic equipment and medium

    CN119814527A