A method, device, equipment, medium and program product for determining a repair strategy
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHINA UNITED NETWORK COMM GRP CO LTD
- Filing Date
- 2026-04-27
- Publication Date
- 2026-08-07
AI Technical Summary
[0004]本申请通过一种修复策略的确定方法、装置、设备、介质和程序产品,解决了现有技术存在的匹配到的修复策略实际上与当前异常事件之间的适配度不高的问题,实现了快速确定最适配与当前异常事件的修复策略,提高对异常事件的修复成功率,使得异常链路能够快速恢复正常
[0016]This application provides a method for determining a repair strategy, comprising: obtaining the abnormal propagation path of a current abnormal event, and the state change sequence and node characteristics of each service node in the abnormal propagation path; determining the composite similarity between the current abnormal event and each reference abnormal case in a case library based on the abnormal propagation path and the state change sequence of each service node, wherein the case library stores the reference propagation path of each reference abnormal case, the reference state change sequence of each service node in the reference propagation path, and the repair strategy corresponding to the reference abnormal case; determining the adaptability score of each repair strategy in the current abnormal event based on the node characteristics of each service node and the abnormal propagation path; and determining the repair strategy for the current abnormal event in the case library based on the adaptability score and the composite similarity. This method solves the problem in existing technologies where the matched repair strategies do not actually have a high degree of adaptability to the current abnormal event, enabling rapid determination of the most suitable repair strategy for the current abnormal event, improving the success rate of repairing abnormal events, and allowing the abnormal link to quickly return to normal.
Smart Images

Figure CN122533928A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the technical field of node anomaly repair, and specifically relates to a method, apparatus, device, medium and program product for determining a repair strategy. Background Technology
[0002] With the rapid development of information technology, the complexity of enterprise business systems is also increasing. The service call chain in the distributed architecture is extended, which makes the propagation path of abnormal events more hidden and difficult to trace. As a result, when an abnormal event is discovered, its impact has often spread to multiple business nodes, which leads to longer fault location time and significantly increased repair costs.
[0003] While existing technologies can quickly pinpoint the root cause of anomalies and determine remediation strategies based on reference anomaly cases in a database, they rely on strategies derived from similar anomaly cases. Furthermore, they only consider the similarity of anomaly features when identifying such cases. However, in reality, while two anomalies may be highly similar in their features, they can differ significantly in actual circumstances. This can lead to a poor fit between the chosen remediation strategy and the current anomaly, resulting in a low remediation success rate. Therefore, remediation strategies determined using existing technologies suffer from low success rates and difficulty in quickly restoring abnormal paths to normal operation. Summary of the Invention
[0004] This application solves the problem in the prior art where the matched repair strategy is not well adapted to the current abnormal event by providing a method, apparatus, device, medium, and program product for determining a repair strategy. It enables the rapid determination of the repair strategy that is most suitable for the current abnormal event, improves the success rate of repairing abnormal events, and enables the abnormal link to quickly return to normal.
[0005] To address the aforementioned technical problems, this application proposes five aspects.
[0006] In a first aspect, this application proposes a method for determining a repair strategy, comprising: obtaining the abnormal propagation path of a current abnormal event, and the state change sequence and node characteristics of each business node in the abnormal propagation path; determining the composite similarity between the current abnormal event and each reference abnormal case in a case library based on the abnormal propagation path and the state change sequence of each business node, wherein the case library stores the reference propagation path of each reference abnormal case, the reference state change sequence of each business node in the reference propagation path, and the repair strategy corresponding to the reference abnormal case; determining the adaptability score of each repair strategy in the current abnormal event based on the node characteristics of each business node and the abnormal propagation path; and determining the repair strategy for the current abnormal event in the case library based on the adaptability score and the composite similarity.
[0007] In some embodiments, determining the composite similarity between the current abnormal event and each reference abnormal case in the case library based on the abnormal propagation path and the state change sequence of each service node includes: determining the propagation similarity between the current abnormal event and each of the reference abnormal cases based on the abnormal propagation path and the reference propagation path; determining the abnormal similarity between the current abnormal event and each of the reference abnormal cases based on the state change sequence of each service node and the reference state change sequence of each service node in the reference propagation path; and determining the composite similarity between the current abnormal event and each of the reference abnormal cases based on the propagation similarity and the abnormal similarity.
[0008] In some embodiments, determining the propagation similarity between the current anomalous event and each of the reference anomalous cases based on the anomalous propagation path and the reference propagation path includes: determining the path similarity between the current anomalous event and the reference anomalous cases based on the anomalous propagation path and the reference propagation path; determining the root cause similarity between the current anomalous event and the reference anomalous cases based on the current root cause node in the anomalous propagation path and the reference root cause node in the reference propagation path; and determining the propagation similarity between the current anomalous event and each of the reference anomalous cases based on the path similarity and the root cause similarity.
[0009] In some embodiments, determining the adaptability score of each repair strategy in the current abnormal event based on the node characteristics of each service node and the anomaly propagation path includes: obtaining the repair success rate of each repair strategy in historical repairs; determining the scenario matching degree of each repair strategy when applied to the current abnormal event based on the node characteristics of each service node; determining the rollback difficulty of each repair strategy after completing the repair of the current abnormal event based on the anomaly propagation path; and determining the adaptability score of each repair strategy in the current abnormal event based on the repair success rate, the scenario matching degree, and the rollback difficulty.
[0010] In some embodiments, obtaining the abnormal propagation path of the current abnormal event includes: obtaining the abnormal occurrence node of the current abnormal event, a unique association identifier between the abnormal occurrence node and the business entity, and the full-link topology map of the business entity; determining all upstream nodes of the abnormal occurrence node in the full-link topology map based on the unique association identifier, wherein the upstream node is the upstream node of the abnormal occurrence node belonging to the business entity in the full-link topology map; calculating the propagation contribution of each upstream node to the current abnormal event based on the historical call frequency and dependency strength between each upstream node and the abnormal occurrence node; determining the root cause node among each upstream node based on the propagation contribution; and forming the abnormal propagation path based on the root cause node.
[0011] In some embodiments, calculating the propagation contribution of each upstream node to the current abnormal event based on the historical call frequency and dependency strength between each upstream node and the node where the anomaly occurred includes: obtaining the time of the anomaly occurrence of the current abnormal event; obtaining the node data and behavioral baseline of each upstream node at the time of the anomaly occurrence; calculating the state contribution of each upstream node to the current abnormal event based on the node data and the behavioral baseline; calculating the structural contribution of each upstream node to the current abnormal event based on the historical call frequency and dependency strength between the upstream node and the node where the anomaly occurred; and calculating the propagation contribution of each upstream node to the current abnormal event based on the structural contribution and the state contribution.
[0012] Secondly, this application provides a device for determining a repair strategy, comprising: a first acquisition module, configured to acquire the abnormal propagation path of a current abnormal event, and the state change sequence and node characteristics of each business node in the abnormal propagation path; a first determination module, configured to determine the composite similarity between the current abnormal event and each reference abnormal case in a case library based on the abnormal propagation path and the state change sequence of each business node, wherein the case library stores the reference propagation path of each reference abnormal case, the reference state change sequence of each business node in the reference propagation path, and the repair strategy corresponding to the reference abnormal case; a second determination module, configured to determine the adaptability score of each repair strategy in the current abnormal event based on the node characteristics of each business node and the abnormal propagation path; and a third determination module, configured to determine the repair strategy for the current abnormal event in the case library based on the adaptability score and the composite similarity.
[0013] Thirdly, this application proposes a computer electronic production apparatus, including a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of any of the methods described in the first aspect.
[0014] Fourthly, this application proposes a computer-readable storage medium having a computer program stored thereon, characterized in that the computer program, when executed by a processor, implements the steps of the method described in any one of the first aspects.
[0015] Fifthly, this application proposes a computer program product, including a computer program, characterized in that, when executed by a processor, the computer program implements the steps of the method described in any one of the first aspects.
[0016] This application provides a method for determining a repair strategy, comprising: obtaining the abnormal propagation path of a current abnormal event, and the state change sequence and node characteristics of each service node in the abnormal propagation path; determining the composite similarity between the current abnormal event and each reference abnormal case in a case library based on the abnormal propagation path and the state change sequence of each service node, wherein the case library stores the reference propagation path of each reference abnormal case, the reference state change sequence of each service node in the reference propagation path, and the repair strategy corresponding to the reference abnormal case; determining the adaptability score of each repair strategy in the current abnormal event based on the node characteristics of each service node and the abnormal propagation path; and determining the repair strategy for the current abnormal event in the case library based on the adaptability score and the composite similarity. This method solves the problem in existing technologies where the matched repair strategies do not actually have a high degree of adaptability to the current abnormal event, enabling rapid determination of the most suitable repair strategy for the current abnormal event, improving the success rate of repairing abnormal events, and allowing the abnormal link to quickly return to normal. Attached Figure Description
[0017] One or more embodiments are illustrated by way of example with reference to the accompanying drawings, and these illustrative descriptions do not constitute a limitation on the embodiments.
[0018] Figure 1 The main flowchart of a method for determining a repair strategy provided in an embodiment of this application; Figure 2 A main structural block diagram of a device for determining a repair strategy provided in an embodiment of this application; Figure 3 This is a structural block diagram of a computer electronic production equipment provided in an embodiment of this application. Detailed Implementation
[0019] With the rapid development of information technology, the complexity of enterprise business systems is also increasing. The service call chain in the distributed architecture is extended, which makes the propagation path of abnormal events more hidden and difficult to trace. As a result, when an abnormal event is discovered, its impact has often spread to multiple business nodes, which leads to longer fault location time and significantly increased repair costs.
[0020] While existing technologies can quickly pinpoint the root cause of anomalies and determine remediation strategies based on reference anomaly cases in a database, they rely on strategies derived from similar anomaly cases. Furthermore, they only consider the similarity of anomaly features when identifying such cases. However, in reality, while two anomalies may be highly similar in their features, they can differ significantly in actual circumstances. This can lead to a poor fit between the chosen remediation strategy and the current anomaly, resulting in a low remediation success rate. Therefore, remediation strategies determined using existing technologies suffer from low success rates and difficulty in quickly restoring abnormal paths to normal operation.
[0021] To address the aforementioned technical problems, this invention proposes a method for determining a repair strategy. The implementation details of this method are described below. The following details are provided for ease of understanding and are not essential for implementing this solution.
[0022] Example 1: like Figure 1 As shown, this application provides a method for determining a repair strategy. This method is applicable to electronic production equipment, which can be a server, mobile terminal, computer, cloud platform, etc. The functions implemented by the data processing of the production equipment provided in this application embodiment can be achieved by the processor of the electronic production equipment calling program code, wherein the program code can be stored in a computer storage medium. The method for determining the repair strategy includes: Step S1: Obtain the abnormal propagation path of the current abnormal event, as well as the state change sequence and node characteristics of each business node in the abnormal propagation path.
[0023] Existing technologies determine repair strategies by directly matching the data features of the current abnormal event's abnormal nodes with the data features of abnormal cases in a case library to identify the most similar abnormal case. Then, the repair strategy of that abnormal case is used to repair the current abnormal case. However, in reality, it's rare for two anomalies to be exactly the same; rather, different anomaly propagation paths, different business entities, different root causes, and different business nodes are more common. This means that repair strategies based solely on simple data feature matching are not well-suited to the current anomaly characteristics, resulting in a low repair success rate. Therefore, to improve the repair success rate, this application needs to match as many repair strategies as possible from the case library suitable for the current abnormal event. This requires obtaining the anomaly propagation path of the current abnormal event, the state change sequence of each business node along the propagation path, and the node characteristics of each business node.
[0024] The process of obtaining the state change sequence involves acquiring multiple time-series indicators (response time, call frequency, error rate, etc.) reflecting the operational status of business nodes within a period before and after an anomaly occurs, and constructing corresponding multi-dimensional time-series feature vectors accordingly. The multi-dimensional time-series feature vectors (including anomaly duration, number of affected users, number of associated nodes, CPU anomaly utilization, network latency, and business failure rate, etc.) undergo dimensionality reduction and feature extraction. Dimensionality reduction can be performed using classic methods such as Principal Component Analysis (PCA), which projects the original high-dimensional data into a low-dimensional space through linear transformation while retaining the main feature information. Feature extraction further filters out feature indicators closely related to the anomaly pattern from the dimensionality-reduced data, such as abrupt changes in response time and abnormal peaks in call frequency, thus forming the state change sequence. Specific indicators need to be set according to the actual business scenario.
[0025] The node characteristics refer to the scope of business impact of a business node, which can be quantified as the number of affected users, as well as the currently available computing resources and network bandwidth.
[0026] The anomaly propagation path refers to the path from the root cause node to a series of downstream nodes affected by the anomaly. The anomaly propagation path can characterize the scope of the anomaly's impact.
[0027] In some embodiments, step S1, "obtaining the anomaly propagation path of the current abnormal event", includes: Step S11: Obtain the node where the current abnormal event occurred, the unique association identifier between the node where the abnormal event occurred and the business entity, and the full-link topology diagram of the business entity.
[0028] To obtain the propagation path of an anomaly, it is necessary to identify the root cause node and upstream and downstream nodes of the anomaly event. Generally, the anomaly node is not the root cause node. The anomaly node is the node where the anomaly can be detected, while the root cause node is the node that caused the downstream anomaly node to appear. However, the root cause node itself does not meet the criteria of an anomaly node, meaning that the root cause node cannot be obtained directly. Therefore, this application needs to first identify the root cause node of the current anomaly event. Determining the root cause node inevitably involves the anomaly occurrence node, the business entity where the anomaly occurrence node is located, and the full-link topology diagram of the business entity. The full-link topology diagram is very important. In addition, this application does not directly obtain the business entity, but obtains a unique association identifier. Through the unique association identifier, other nodes belonging to the same business entity as the anomaly node can be directly found in the full-link topology diagram.
[0029] The process of obtaining the full-link topology map in this application is as follows: Identify the identifiers and states of the business entities based on the standardized event flow of each node, and reconstruct the flow path of the business entities based on the temporal relationship and semantic association between events to form a full-link business topology map.
[0030] Specifically, unique identifiers for business entities, such as order numbers or user IDs, are located from the standardized event stream to ensure real-time tracking of business entity status. State change events associated with these unique identifiers are also extracted. These events record the state changes of business entities at different business nodes, such as order creation, payment, and shipment. The extracted state change events are then sorted according to their timestamp sequence to ensure they are arranged in chronological order, providing a data foundation for building temporal dependencies. Semantic association rules are then used to identify call relationships. These rules define the correspondence between data and operations across different business systems, clarifying the flow logic of business entities between systems and identifying cross-system call relationships. This lays the foundation for building a complete business chain topology. Next, based on event timestamps and state transition rules (which are formulated according to the logical relationship of state transitions of business entities across different nodes), the temporal dependencies between events are further constructed. Then, multiple state transition sequences with the same unique identifier are aggregated. These sequences may originate from different business systems or different points in time, but they are all associated with the same business entity. By concatenating all state transition sequences of the same business entity in chronological order, a complete business entity flow trajectory is formed. This trajectory illustrates the state changes of the business entity at various business nodes. Then, the intersection nodes of multiple business entity flow trajectories are merged. These intersection nodes represent the points where different business entities meet in the business process. By extending the link boundaries through semantic association rules, the intersection points and their call relationships and dependency paths can be incorporated into the topology graph (dependency paths refer to the preceding nodes or resource paths that a business entity must traverse during the flow). Finally, a full-link topology graph of multiple business entities is constructed.
[0031] Step S12: Determine all upstream nodes of the anomaly occurrence node in the full-link topology diagram based on the unique association identifier, wherein the upstream nodes are the upstream nodes of the anomaly occurrence node belonging to the business entity in the full-link topology diagram.
[0032] Based on the full-link topology diagram, all upstream business nodes associated with the unique identifier are traced in reverse to find the source node that may cause abnormal events. By querying the full-link topology diagram, the flow path of business entities between various business nodes can be clarified, thereby identifying the upstream nodes associated with the abnormal nodes.
[0033] Step S13: Calculate the contribution of each upstream node to the propagation of the current abnormal event based on the historical call frequency and dependency strength between each upstream node and the node where the anomaly occurred.
[0034] In some embodiments, step S13, "calculating the propagation contribution of each upstream node to the current abnormal event based on the historical call frequency and dependency strength between each upstream node and the node where the anomaly occurred," includes: Step S131: Obtain the time when the current abnormal event occurred.
[0035] Step S132: Obtain the node data and behavior baseline of each upstream node at the time the anomaly occurred.
[0036] To determine the root cause node among the upstream nodes in this application, it is necessary to calculate the propagation contribution of each upstream node and determine the node with the highest propagation contribution as the root cause node. In this application, the propagation contribution mainly includes structural contribution and state contribution. The structural contribution is used to characterize the contribution of the node to the abnormal transmission in the structure of the whole link topology graph, while the state contribution characterizes the contribution of the node's data state to causing the downstream abnormality.
[0037] Therefore, when calculating the state contribution, it is necessary to first obtain the node data and behavioral baseline of the upstream node at the time of the anomaly. The behavioral baseline is used to determine the reference data for the node at that moment; however, in reality, the node data will fluctuate within a certain range of the behavioral baseline. If the fluctuation does not exceed a preset threshold, the node is not considered to have an anomaly. To improve the sensitivity and accuracy of identifying anomalous nodes, the behavioral baseline in this application is dynamic, thus requiring the acquisition of the behavioral baseline at the time of the anomaly.
[0038] Specifically, the behavior update process in this application is as follows: After a business node is determined to be normal, its corresponding node data (including request response latency, throughput, error rate, and resource utilization) will be included in the update sample set of the behavior baseline. Simultaneously, the node data will undergo corresponding timeliness and reasonableness checks. The purpose is to ensure that the node data in the update sample set accurately reflects the actual operating characteristics of the current business node. First, it is verified whether the timestamp of the node data is within a reasonable range. The reasonable range is a dynamic time window set based on the system clock synchronization mechanism; the specific span can be set according to actual needs. By comparing the reasonable range with the timestamp, it is possible to eliminate... Excluding delayed or premature node data due to network latency or system anomalies, this ensures that only valid and timely sample data is included in the updated sample set. The filtered updated sample set is recorded as the valid sample set. Then, based on a time-decay weighted method, the valid sample set is fused with the old samples stored in the historical behavior baseline to obtain the updated behavior baseline. Under stable operation, the mean and variance of the historical behavior baseline can be calculated based on a sliding window mechanism. During this process, the sample weights within the sliding window follow a Gaussian distribution, meaning that samples closer to the current time have larger weights, while the weights of samples farther from the current time decay exponentially. Specifically, this can be expressed as... Where λ is the decay coefficient and Δt is the time difference between the sample time and the current time, when business traffic changes drastically (e.g., within a preset short period, the standard deviation of the business node request rate exceeds a preset proportion of its long-term historical average, or the sliding coefficient of variation of the resource utilization index exceeds the variation threshold), in addition to considering time decay, a change compensation factor is also introduced to enhance the weight ratio of new samples, which can be specifically expressed as: In the formula, α is the compensation intensity coefficient, which is dynamically adjusted according to the intensity of business fluctuations. This indicates the volatility of node data within the current sample's time window, thereby accelerating the response speed of the behavioral baseline to the current operating status and ensuring that the behavioral baseline can reflect the latest operating characteristics of business nodes in a timely manner, avoiding misjudgments caused by outdated historical data.
[0039] Step S133: Calculate the state contribution of each upstream node to the current abnormal event based on the node data and the behavioral baseline.
[0040] The formula for calculating the contribution of each node in this application is as follows: , In the formula, Indicates the state contribution. This represents the mutation amount of the i-th abnormal indicator of node u, such as the difference between the current CPU utilization and the historical baseline mean. This represents the historical standard deviation of the i-th abnormal indicator at node u, used to reflect the normal fluctuation range. represents the weight of the i-th abnormal indicator, and n represents the number of abnormal indicators.
[0041] Step S134: Determine the structural contribution of the current abnormal event based on the historical call frequency and dependency strength between the upstream node and the node where the anomaly occurred.
[0042] The formula for calculating the structural contribution in this application is as follows: , In the formula, Indicates the structural contribution. and These represent the weighting coefficients for call frequency and dependency strength, respectively, which are determined based on historical call frequency and dependency strength. This indicates the frequency with which node u calls node v in the historical record. This represents the sum of the total call frequencies of all upstream nodes of node v. This indicates the strength of the dependency between node u and node v.
[0043] Step S135: Calculate the contribution of each upstream node to the propagation of the current abnormal event based on the structural contribution and the state contribution.
[0044] Considering the lag effect in the time series, a time window sliding mechanism is introduced to quantify the time delay relationship between the abnormal signal of the upstream node and the current node state change. The time series between upstream and downstream nodes are aligned by the dynamic time warping algorithm to calculate their similarity score. The fusion weight of structural contribution and state contribution is further corrected. Based on the corrected fusion weight, the structural contribution and state contribution are fused into the propagation contribution.
[0045] Step S14: Determine the root cause node among each of the upstream nodes based on the propagation contribution.
[0046] Step S15: Form the anomaly propagation path based on the root cause node.
[0047] By designating the upstream node with the highest contribution to the propagation as the root cause node, the source node most likely to trigger the abnormal event, i.e., the root cause node, can be quickly located. Then, the propagation path is extended downstream from the root cause node to form a proposed propagation path for the anomaly. The state correlation coefficient between each business node in the proposed propagation path and the root cause node is calculated. The state correlation coefficient can be calculated using cosine similarity or Pearson correlation coefficient to determine the propagation path of the abnormal event in the business chain and the degree of association between each business node and the root cause node. Then, business nodes with state correlation coefficients higher than a relevant threshold are marked as abnormal associated nodes, and all abnormal associated nodes are connected to form the anomaly propagation path. By using a pre-set relevant threshold, business nodes with a high degree of association with the root cause node, i.e., abnormal associated nodes, can be selected. Finally, all abnormal associated nodes are connected to form a complete anomaly propagation path.
[0048] Step S2: Determine the composite similarity between the current abnormal event and each reference abnormal case in the case library based on the abnormal propagation path and the state change sequence of each business node, wherein the case library stores the reference propagation path of each reference abnormal case, the reference state change sequence of each business node in the reference propagation path, and the repair strategy corresponding to the reference abnormal case.
[0049] In some embodiments, step S2, "determining the composite similarity between the current abnormal event and each reference abnormal case in the case library based on the abnormal propagation path and the state change sequence of each business node," includes: Step S21: Determine the propagation similarity between the current abnormal event and each of the reference abnormal cases based on the abnormal propagation path and the reference propagation path.
[0050] In some embodiments, step S21, "determining the propagation similarity between the current abnormal event and each of the reference abnormal cases based on the abnormal propagation path and the reference propagation path," includes: Step S211: Determine the path similarity between the current abnormal event and the reference abnormal case based on the abnormal propagation path and the reference propagation path.
[0051] Step S212: Determine the root cause similarity between the current abnormal event and the reference abnormal case based on the current root cause node in the abnormal propagation path and the reference root cause node in the reference propagation path.
[0052] Step S213: Determine the propagation similarity between the current abnormal event and each of the reference abnormal cases based on the path similarity and the root cause similarity.
[0053] Specifically, the current propagation characteristics of the current anomalous event are determined based on the anomalous propagation path. Reference propagation characteristics of each reference anomalous case are determined based on the reference propagation path. The propagation similarity between the current anomalous event and each of the reference anomalous cases is determined based on the current propagation characteristics and the reference propagation characteristics.
[0054] Then, the root cause similarity between the current anomalous event and the reference anomalous case is determined based on the current root cause feature in the current propagation features and the reference root cause feature in the reference propagation features. The path similarity between the current anomalous event and the reference anomalous case is then determined based on the current path feature in the current propagation features and the reference path feature in the reference propagation features.
[0055] The current propagation features in this application include root cause features and path features, where root cause features are the characteristics of the root cause node, and path features are the characteristics of the propagation path. Similarly, the case library in this application also stores the root cause features and path features of each case. Therefore, this application can determine the root cause similarity between the current abnormal event and each reference abnormal case in the case library based on the current root cause features and reference root cause features, and can determine the path similarity between the current abnormal event and reference abnormal cases based on the current path features and reference path features. The root cause similarity and path similarity are fused according to a preset weight to form the propagation similarity. Existing technologies do not consider propagation similarity, but the calculation of propagation similarity is beneficial for further identifying the commonalities between the final selected target reference abnormal case and the current abnormal event, making the repair strategy of the target reference abnormal case more applicable to the current abnormal event, and increasing the success rate of the determined repair strategy.
[0056] Step S22: Determine the anomaly similarity between the current abnormal event and each of the reference abnormal cases based on the state change sequence of each service node and the reference state change sequence of each service node in the reference propagation path.
[0057] Specifically, the current anomaly characteristics of the current anomaly event are determined based on the state change sequence of each business node in the anomaly propagation path. Reference anomaly characteristics of reference anomaly cases are determined based on the reference state change sequence of each business node in the reference propagation path. The anomaly similarity between the current anomaly event and each of the reference anomaly cases is determined based on the current anomaly characteristics and the reference anomaly characteristics.
[0058] Step S23: Determine the composite similarity between the current abnormal event and each of the reference abnormal cases based on the propagation similarity and the anomaly similarity.
[0059] Moreover, in this application, when determining reference anomalous cases, the reference anomalous cases are not determined solely based on conventional anomalous similarity and propagation similarity, but rather by merging the two according to preset weights to form a composite similarity.
[0060] Step S3: Determine the adaptability score of each repair strategy in the current abnormal event based on the node characteristics and the anomaly propagation path.
[0061] In some embodiments, step S3, "determining the adaptability score of each repair strategy in the current abnormal event based on the node characteristics of each service node and the abnormal propagation path," includes: Step S31: Obtain the repair success rate of each repair strategy in historical repairs.
[0062] The repair success rate includes the average repair time and success rate of the repair strategies corresponding to each reference anomaly case in historical cases.
[0063] Step S32: Determine the scenario matching degree of each repair strategy when applied to the current abnormal event based on the node characteristics of each business node.
[0064] Scene matching degree refers to the degree of matching between the node characteristics of the reference abnormal case repaired by the repair strategy and the node characteristics of the current abnormal event. Specifically, it is reflected in the actual business impact range, which can be quantified as the number of affected users, as well as the currently available computing resources and network bandwidth.
[0065] Step S33: Determine the rollback difficulty of each repair strategy after repairing the current abnormal event based on the abnormal propagation path.
[0066] Rollback difficulty refers to the time cost and operational complexity required to restore the system to its pre-repair state if the repair fails or new problems arise after a repair operation. The higher the rollback difficulty, the greater the execution risk. The larger the anomaly propagation chain, the more nodes are involved in the rollback, and the greater the rollback difficulty.
[0067] Step S34: Determine the adaptability score of each repair strategy in the current abnormal event based on the repair success rate, the scenario matching degree, and the rollback difficulty.
[0068] Weights are assigned to the success rate of the repair, the matching degree of the scenario, and the difficulty of the rollback. The overall utility score of each repair strategy is calculated in a weighted manner, which is the adaptability score.
[0069] Step S4: Determine a remediation strategy for the current anomalous event in the case library based on the adaptability score and the composite similarity.
[0070] When determining the remediation strategy based on adaptability scores and composite similarity, candidate reference anomaly cases that are similar to the current anomaly can be identified first based on composite similarity scores; these are candidate reference anomaly cases whose composite similarity to the current anomaly exceeds a preset threshold. Then, the candidate reference anomaly cases are ranked according to the adaptability scores of their corresponding remediation strategies, and the remediation strategy corresponding to the reference anomaly case with the highest adaptability score is selected as the remediation strategy for the current anomaly.
[0071] After the repair is completed, it will be fed back into the case library as a new case to update the historical handling efficiency data and realize the self-enhancement and iterative optimization of the case library.
[0072] In addition, the repair strategy may fail to execute, although this is a low-probability event. Therefore, when a repair strategy fails, a supplementary repair mechanism will be activated to attempt a supplementary repair. During this process, a secondary screening is first performed on the adaptive scores, sorted in descending order. Multiple candidate repair strategies are selected using a pre-set screening threshold. Specifically, the top N candidate repair strategies with scores higher than the screening threshold are selected, where N is the preset number of candidates, which can be adjusted according to system resources and processing capacity. The screening threshold is set based on the average adaptive score of historical cases. To ensure the reliability of the repair process, the average adaptive score is lowered by a certain percentage as the final screening threshold to avoid raising the overall standard due to a high score in a single case. The specific percentage reduction can be set according to actual needs. Then, the compatibility of the candidate repair strategies with previously executed but failed repair strategies is evaluated. The complementarity mechanism first eliminates candidate repair strategies that have already been fully implemented and are included in failed strategies, as re-execution of these strategies will not solve the problem and may even exacerbate the anomaly. Next, it considers the compatibility between remaining candidate repair strategies, i.e., whether executing multiple candidate repair strategies simultaneously will cause conflicts, such as resource contention or operational logic contradictions. Candidate repair strategies without operational conflicts are aggregated into a sub-repair combination. When multiple sub-repair combinations exist, weights are allocated based on the adaptability scores of each candidate repair strategy within each sub-repair combination, with higher scores resulting in greater weights. Finally, a weighted fusion score is obtained for each sub-repair combination, and serialization attempts are performed according to the fusion score ranking. Attempts cease upon successful attempts. Conversely, if all sub-repair combinations fail, a repair failure report is directly exported and sent to the management end for manual review.
[0073] In addition, when multiple abnormal events exist simultaneously, they need to be handled one by one. Therefore, appropriate priority ranking is required. First, it is necessary to collect the occurrence frequency of each abnormal event, that is, the number of times the abnormal event occurs per unit time. At the same time, it is necessary to record the business processing delay, that is, the time elapsed from the occurrence of the abnormal event to the start of processing, and count the number of related abnormal events, that is, the number of other abnormal events related to the current abnormal event. Based on this, a weighted fusion calculation can be performed to generate a comprehensive priority score. The comprehensive priority score can comprehensively reflect the urgency of the abnormal event. Then, multiple abnormal events can be sorted in descending order according to the comprehensive priority score. The higher the comprehensive priority score, the higher the priority of the abnormal event, and the more it needs to be handled first. Finally, the processing priority queue is output, which can provide clear guidance for subsequent abnormal handling work.
[0074] This application provides a method for determining a repair strategy, comprising: obtaining the abnormal propagation path of a current abnormal event, and the state change sequence and node characteristics of each service node in the abnormal propagation path; determining the composite similarity between the current abnormal event and each reference abnormal case in a case library based on the abnormal propagation path and the state change sequence of each service node, wherein the case library stores the reference propagation path of each reference abnormal case, the reference state change sequence of each service node in the reference propagation path, and the repair strategy corresponding to the reference abnormal case; determining the adaptability score of each repair strategy in the current abnormal event based on the node characteristics of each service node and the abnormal propagation path; and determining the repair strategy for the current abnormal event in the case library based on the adaptability score and the composite similarity. This method solves the problem that the repair success rate of repair strategies determined by existing technologies is low, making it difficult for the link to quickly return to normal. It enables rapid determination of the most suitable repair strategy for the current abnormal event, improves the repair success rate of abnormal events, and allows abnormal links to quickly return to normal.
[0075] Example 2: Based on the foregoing embodiments, this application provides a device for determining a repair strategy. The various modules and units included in the device can be implemented by a processor in a computer device; of course, they can also be implemented by specific logic circuits. In the implementation process, the processor can be a central processing unit (CPU), a microprocessor (MPU), a digital signal processor (DSP), or a field programmable gate array (FPGA), etc.
[0076] like Figure 2 As shown, a device for determining a repair strategy includes: a first acquisition module 1, a first determination module 2, a second determination module 3, and a third determination module 4.
[0077] The first acquisition module 1 is used to acquire the abnormal propagation path of the current abnormal event, as well as the state change sequence and node characteristics of each business node in the abnormal propagation path. The first determination module 2 is used to determine the composite similarity between the current abnormal event and each reference abnormal case in the case library based on the abnormal propagation path and the state change sequence of each business node, wherein the case library stores the reference propagation path of each reference abnormal case, the reference state change sequence of each business node in the reference propagation path, and the corresponding repair strategy for each reference abnormal case. The second determination module 3 is used to determine the adaptability score of each repair strategy in the current abnormal event based on the node characteristics of each business node and the abnormal propagation path. The third determination module 4 is used to determine the repair strategy for the current abnormal event in the case library based on the adaptability score and the composite similarity.
[0078] In some embodiments, the first determining module 2 is specifically configured to: determine the propagation similarity between the current abnormal event and each of the reference abnormal cases based on the abnormal propagation path and the reference propagation path; determine the abnormal similarity between the current abnormal event and each of the reference abnormal cases based on the state change sequence of each service node and the reference state change sequence of each service node in the reference propagation path; and determine the composite similarity between the current abnormal event and each of the reference abnormal cases based on the propagation similarity and the abnormal similarity.
[0079] In some embodiments, the first determining module 2 is further configured to: determine the path similarity between the current abnormal event and the reference abnormal case based on the abnormal propagation path and the reference propagation path; determine the root cause similarity between the current abnormal event and the reference abnormal case based on the current root cause node in the abnormal propagation path and the reference root cause node in the reference propagation path; and determine the propagation similarity between the current abnormal event and each of the reference abnormal cases based on the path similarity and the root cause similarity.
[0080] In some embodiments, the second determining module 3 is specifically used to: obtain the repair success rate of each repair strategy in historical repairs; determine the scenario matching degree of each repair strategy when applied to the current abnormal event based on the node characteristics of each business node; determine the rollback difficulty of each repair strategy after completing the repair of the current abnormal event based on the abnormal propagation path; and determine the adaptability score of each repair strategy in the current abnormal event based on the repair success rate, the scenario matching degree, and the rollback difficulty.
[0081] In some embodiments, the first acquisition module 1 is specifically configured to: acquire the anomaly occurrence node of the current abnormal event, a unique association identifier between the anomaly occurrence node and the business entity, and the full-link topology map of the business entity; determine all upstream nodes of the anomaly occurrence node in the full-link topology map based on the unique association identifier, wherein the upstream nodes are the upstream nodes of the business entity in the full-link topology map of the anomaly occurrence node; calculate the propagation contribution of each upstream node to the current abnormal event based on the historical call frequency and dependency strength between each upstream node and the anomaly occurrence node; determine the root cause node among each upstream node based on the propagation contribution; and form the anomaly propagation path based on the root cause node.
[0082] In some embodiments, the first acquisition module 1 is further configured to: acquire the time of occurrence of the current abnormal event; acquire the node data and behavioral baseline of each upstream node at the time of occurrence of the abnormality; calculate the state contribution of each upstream node to the current abnormal event based on the node data and the behavioral baseline; calculate the structural contribution of each upstream node to the current abnormal event based on the historical call frequency and dependency strength between the upstream node and the node where the abnormality occurred; and calculate the propagation contribution of each upstream node to the current abnormal event based on the structural contribution and the state contribution.
[0083] The modules in the aforementioned device for determining a repair strategy can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor within the device in hardware form, or stored in the memory of the processing device in software form, so that the processor can call and execute the operations corresponding to each module. It should be noted that the module division in this embodiment is illustrative and represents only a logical functional division; in actual implementation, other division methods may be used.
[0084] Example 3: Thirdly, this application provides a computer electronic production device, such as... Figure 3 As shown, it includes: at least one processor 100; and a memory 200 communicatively connected to the at least one processor 100; wherein the memory 200 stores instructions executable by the at least one processor 100, the instructions being executed by the at least one processor 100 to enable the at least one processor 100 to perform a method for determining a repair strategy in the above embodiments.
[0085] The memory and processor are connected via a bus, which can include any number of interconnecting buses and bridges, connecting various circuits of one or more processors and memories. The bus can also connect various other circuits, such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art and will not be described further herein. The bus interface provides an interface between the bus and the transceiver. The transceiver can be a single element or multiple elements, such as multiple receivers and transmitters, providing a unit for communicating with various other devices over a transmission medium. Data processed by the processor is transmitted over the wireless medium via an antenna, which further receives data and transmits it to the processor.
[0086] The processor manages the bus and general processing, and also provides various functions, including timing, peripheral interfaces, voltage regulation, power management, and other control functions. Memory is used to store data used by the processor during operation.
[0087] In addition, the computer device may include (but is not limited to) a data bus, an input / output (I / O) bus, a display, and input / output devices (e.g., keyboard, mouse, speakers, etc.).
[0088] The processor can communicate with external devices via the I / O bus through wired or wireless networks.
[0089] In one embodiment, the at least one computer-executable instruction may also be compiled into or comprise a software product / computer program product, wherein one or more computer-executable instructions are executed by a processor to perform the steps of the various functions and / or methods in the embodiments described herein.
[0090] Example 4: Fourthly, this application proposes a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method described in any of the first aspects.
[0091] Computer-readable storage media can be implemented by any type of volatile or non-volatile storage device or a combination thereof. Computer-readable storage media may include, but are not limited to, random access memory (RAM), read-only memory (ROM), flash memory, EPROM memory, EEPROM memory, registers, and computer storage media (e.g., hard disks, floppy disks, solid-state drives, removable disks, CD-ROMs, DVD-ROMs, Blu-ray discs, etc.).
[0092] Computer-readable storage media may also store at least one computer-executable program / instruction, such as computer-readable instructions. Computer-readable storage media include, but are not limited to, volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory. Computer-readable storage media may include, for example, read-only memory (ROM), hard disk, flash memory, etc. For example, a non-transitory computer-readable storage medium may be connected to a computing device such as a computer, and then, when the computing device executes the computer-readable instructions stored on the computer-readable storage medium, the various methods described above can be performed.
[0093] Example 5: Fifthly, this application proposes a computer program product, including a computer program / instructions, characterized in that, when the computer program is executed by a processor, it implements the steps of the method described in any one of the first aspects.
[0094] Those skilled in the art will understand that all or part of the steps in the methods of the above embodiments can be implemented by a program instructing related hardware. This program is stored in a storage medium and includes several instructions to cause a device (which may be a microcontroller, chip, etc.) or processor to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0095] It should be understood that the phrase "one embodiment" or "an embodiment" throughout the specification means that a specific feature, structure, or characteristic related to the embodiment is included in at least one embodiment of this application. Therefore, "in one embodiment" or "in an embodiment" appearing throughout the specification does not necessarily refer to the same embodiment. Furthermore, these specific features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. It should be understood that in the various embodiments of this application, the sequence numbers of the above-described processes do not imply a sequential order of execution; the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application. The sequence numbers of the above-described embodiments are merely descriptive and do not represent the superiority or inferiority of the embodiments.
[0096] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0097] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods, such as: multiple units or components can be combined, or integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the various components shown or discussed can be through some interfaces, and the indirect coupling or communication connection between devices or units can be electrical, mechanical, or other forms.
[0098] In the embodiments provided in this disclosure, it should be understood that the disclosed apparatus and methods can also be implemented in other ways. The apparatus embodiments described above are merely illustrative; for example, the flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of apparatus, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram and / or flowchart, and combinations of blocks in block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0099] Alternatively, if the integrated units described above are implemented as software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, or the parts that contribute to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a controller to execute all or part of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, ROMs, magnetic disks, or optical disks.
[0100] It should be noted that the terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein.
[0101] Those skilled in the art will understand that the above embodiments are specific embodiments for implementing this application, and in practical applications, various changes can be made to them in form and detail without departing from the spirit and scope of this application.
Claims
1. A method for determining a repair strategy, characterized in that, include: Obtain the anomaly propagation path of the current anomaly event, as well as the state change sequence and node characteristics of each business node in the anomaly propagation path; The composite similarity between the current abnormal event and each reference abnormal case in the case library is determined based on the abnormal propagation path and the state change sequence of each business node. The case library stores the reference propagation path of each reference abnormal case, the reference state change sequence of each business node in the reference propagation path, and the repair strategy corresponding to the reference abnormal case. The adaptability score of each repair strategy in the current abnormal event is determined based on the node characteristics of each business node and the abnormal propagation path. Based on the adaptability score and the composite similarity, a remediation strategy for the current anomalous event is determined in the case library.
2. The method according to claim 1, characterized in that, The step of determining the composite similarity between the current abnormal event and each reference abnormal case in the case library based on the abnormal propagation path and the state change sequence of each business node includes: Determine the propagation similarity between the current abnormal event and each of the reference abnormal cases based on the abnormal propagation path and the reference propagation path; The anomaly similarity between the current abnormal event and each of the reference abnormal cases is determined based on the state change sequence of each business node and the reference state change sequence of each business node in the reference propagation path. The composite similarity between the current abnormal event and each of the reference abnormal cases is determined based on the propagation similarity and the anomaly similarity.
3. The method according to claim 2, characterized in that, Determining the propagation similarity between the current abnormal event and each of the reference abnormal cases based on the abnormal propagation path and the reference propagation path includes: Determine the path similarity between the current abnormal event and the reference abnormal case based on the abnormal propagation path and the reference propagation path; The root cause similarity between the current anomalous event and the reference anomalous case is determined based on the current root cause node in the anomalous propagation path and the reference root cause node in the reference propagation path. The propagation similarity between the current anomalous event and each of the reference anomalous cases is determined based on the path similarity and the root cause similarity.
4. The method according to claim 1, characterized in that, The process of determining the adaptability score of each repair strategy in the current abnormal event based on the node characteristics of each business node and the abnormal propagation path includes: Obtain the repair success rate of each repair strategy in historical repairs; The scenario matching degree of each repair strategy when applied to the current abnormal event is determined based on the node characteristics of each business node. The difficulty of rolling back each repair strategy after the current abnormal event is completed is determined based on the abnormal propagation path. The adaptability score of each repair strategy in the current abnormal event is determined based on the repair success rate, the scenario matching degree, and the rollback difficulty.
5. The method according to claim 1, characterized in that, The step of obtaining the anomaly propagation path of the current anomaly event includes: Obtain the node where the current abnormal event occurred, the unique association identifier between the node and the business entity, and the full-link topology of the business entity; Based on the unique association identifier, all upstream nodes of the node where the anomaly occurred are determined in the full-link topology diagram, wherein the upstream nodes are the upstream nodes of the business entity in the full-link topology diagram of the node where the anomaly occurred. The contribution of each upstream node to the propagation of the current abnormal event is calculated based on the historical call frequency and dependency strength between each upstream node and the node where the anomaly occurred. The root cause node is determined among the upstream nodes based on the propagation contribution. The anomaly propagation path is formed based on the root cause node.
6. The method according to claim 5, characterized in that, The step of calculating the contribution of each upstream node to the propagation of the current abnormal event based on the historical call frequency and dependency strength between each upstream node and the node where the anomaly occurred includes: Obtain the time of occurrence of the current abnormal event; Obtain the node data and behavioral baseline of each upstream node at the time the anomaly occurred; Calculate the state contribution of each upstream node to the current abnormal event based on the node data and the behavioral baseline; The structural contribution of the current abnormal event is based on the historical call frequency and dependency strength between the upstream node and the node where the anomaly occurred. The contribution of each upstream node to the propagation of the current abnormal event is calculated based on the structural contribution and the state contribution.
7. A device for determining a repair strategy, characterized in that, include: The first acquisition module is used to acquire the abnormal propagation path of the current abnormal event, as well as the state change sequence and node characteristics of each business node in the abnormal propagation path. The first determining module determines the composite similarity between the current abnormal event and each reference abnormal case in the case library based on the abnormal propagation path and the state change sequence of each business node. The case library stores the reference propagation path of each reference abnormal case, the reference state change sequence of each business node in the reference propagation path, and the repair strategy corresponding to the reference abnormal case. The second determining module is used to determine the adaptability score of each repair strategy in the current abnormal event based on the node characteristics of each business node and the abnormal propagation path. The third determining module is used to determine a repair strategy for the current anomalous event from the case library based on the adaptability score and the composite similarity.
8. A computer electronic production equipment, characterized in that, The method includes a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the method according to any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the computer program performs the steps of the method according to any one of claims 1 to 6.
10. A computer program product, comprising a computer program, characterized in that, When executed by a processor, the computer program performs the steps of the method according to any one of claims 1 to 6.