Abnormal root cause determination method and device

By constructing a call path latency impact kernel for microservice systems and performing convolution processing, a cross-domain jitter hybrid kernel is generated, which solves the problem of accurate location of cross-domain anomaly propagation in cloud-edge collaborative hybrid architecture and achieves efficient and accurate location of anomaly root causes.

CN121750447APending Publication Date: 2026-03-27NEUSOFT CORP
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-30
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

In microservice systems with a cloud-edge collaborative hybrid architecture, existing methods cannot accurately pinpoint the root cause of cross-domain anomaly propagation, resulting in insufficient accuracy in location.

Method used

By obtaining the cause residual of upstream nodes and the effect residual of downstream abnormal nodes, a latency impact kernel of the call path is constructed. Then, a convolutional processing is performed using a latency distribution model between cloud data centers and edge nodes to generate a cross-domain jitter hybrid kernel, so as to accurately characterize the abnormal propagation effect caused by cross-domain long-tail latency.

Benefits of technology

It improves the accuracy of anomaly root cause localization, accurately locates the initial anomaly node, reduces computational complexity, and improves computational efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121750447A_ABST
    Figure CN121750447A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an abnormal root cause positioning method and device, and the method comprises the steps: obtaining a time delay influence kernel of a call path according to a cause residual error of an upstream node at a target moment and an effect residual error of a downstream abnormal node at the target moment; the time delay distribution model of the first link is used for carrying out convolution processing on the time delay influence kernel to obtain a first domain kernel, the time delay distribution model of the second link is used for carrying out convolution processing on the time delay influence kernel to obtain a second domain kernel, and the time delay distribution model of the third link is used for carrying out convolution processing on the time delay influence kernel to obtain a third domain kernel. And according to the first domain kernel, the second domain kernel and the third domain kernel, determining a cross-domain jitter hybrid kernel of the call path. The cross-domain jitter mixed kernel can accurately describe an abnormal propagation amplification effect caused by cross-domain long-tail time delay, and the root cause positioning accuracy is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of microservice system technology, and in particular to a method and apparatus for determining the root cause of anomalies. Background Technology

[0002] In cloud-edge hybrid microservice systems, the microservice call chain often needs to travel back and forth between the cloud data center and edge nodes. This means that any upstream node failure can be amplified and propagated to downstream critical Service Level Objective (SLO) nodes within a short period of time via communication. To address this type of cross-domain failure propagation, it is necessary to accurately pinpoint the root cause of the failure.

[0003] Currently, methods for locating the root cause of anomalies can be achieved using causal divergence or diffusion models. These methods treat anomalies as event-level propagation and infer the network structure to pinpoint the root cause. However, these methods suffer from insufficient accuracy in locating the root cause. For cross-domain anomaly propagation, accurately locating the root cause remains a challenging technical problem. Summary of the Invention

[0004] This application provides a method and apparatus for determining the root cause of anomalies, which is used to accurately locate the root cause of anomalies in cross-domain anomaly propagation.

[0005] In a first aspect, embodiments of this application provide a method for determining the root cause of an anomaly, applied to a microservice system, wherein the architecture of the microservice system is a cloud-edge collaborative hybrid architecture, the cloud-edge collaborative hybrid architecture including a cloud data center and edge nodes, and the method includes:

[0006] Based on the cause residual of the upstream node at the target time and the effect residual of the downstream abnormal node at the target time, obtain the latency impact kernel of the call path;

[0007] The call path is the call path from the upstream node to the downstream abnormal node, and the call path includes a first link, a second link, and a third link; the first link is a link within the cloud data center, the second link is a link between the edge nodes, and the third link is a cross-domain link between the cloud data center and the edge nodes; the cause residual is the difference between the indicator detection value of the upstream node at the target time and the expected value of the first indicator; the effect residual is the difference between the indicator detection value of the downstream abnormal node at the target time and the expected value of the second indicator, and the abnormal window of the upstream node and the abnormal window of the downstream abnormal node are the same; the latency impact kernel is a response function that quantifies the propagation of the upstream node's abnormality over time and its impact on the downstream abnormal node;

[0008] The latency impact kernel is convolved using the latency distribution model of the first link to obtain a first domain kernel; the latency impact kernel is convolved using the latency distribution model of the second link to obtain a second domain kernel; and the latency impact kernel is convolved using the latency distribution model of the third link to obtain a third domain kernel.

[0009] Based on the first domain kernel, the second domain kernel, and the third domain kernel, determine the cross-domain jitter hybrid kernel of the call path;

[0010] Based on the cross-domain jitter hybrid kernel, the root cause of the abnormality of the downstream abnormal node is determined.

[0011] Optionally, obtaining the latency impact kernel of the call path based on the cause residual of the upstream node at the target time and the effect residual of the downstream abnormal node at the target time includes:

[0012] Obtain the cross-correlation curves between the lag residuals and the causal residuals; the lag residuals are the effects residuals lagged by the lag. The residual obtained afterwards, the ;

[0013] Wherein, the dependent variable of the cross-correlation curve is the cross-correlation value between the lagged residual and the causal residual, and the independent variable of the cross-correlation curve is... ;

[0014] Obtain the maximum cross-correlation value and the corresponding baseline lag time from the cross-correlation curve;

[0015] After the baseline lag time, determine the target lag time corresponding to the first decrease of the cross-correlation value in the cross-correlation curve to the target correlation value;

[0016] The curve representing the target interval is extracted from the cross-correlation curve and used as the latency impact kernel of the calling path; the target interval is from 0 to the target lag time.

[0017] Optionally, if the delay impact kernel is a preliminary delay impact kernel, the method further includes:

[0018] The initial latency impact kernel of the call path is processed using regularized deconvolution to obtain a refined latency impact kernel;

[0019] The process of convolving the latency impact kernel with the latency distribution model of the first link to obtain a first domain kernel; convolving the latency impact kernel with the latency distribution model of the second link to obtain a second domain kernel; and convolving the latency impact kernel with the latency distribution model of the third link to obtain a third domain kernel includes:

[0020] The refined latency impact kernel is convolved using the latency distribution model of the first link to obtain the first domain kernel; the refined latency impact kernel is convolved using the latency distribution model of the second link to obtain the second domain kernel; and the refined latency impact kernel is convolved using the latency distribution model of the third link to obtain the third domain kernel.

[0021] Optionally, determining the cross-domain jitter hybrid core of the call path based on the first domain kernel, the second domain kernel, and the third domain kernel includes:

[0022] In the call path, a first time ratio of the first link in the complete call of the call path, a second time ratio of the second link in the complete call of the call path, and a third time ratio of the third link in the complete call of the call path are determined;

[0023] The sum of the first product, the second product, and the third product is taken as the cross-domain jitter hybrid kernel; wherein, the first product is the product of the first time ratio and the first domain kernel, the second product is the product of the second time ratio and the second domain kernel, and the third product is the product of the third time ratio and the third domain kernel.

[0024] Optionally, the methods for obtaining the delay distribution model of the first link, the delay distribution model of the second link, and the delay distribution model of the third link include:

[0025] From the distributed tracing logs, obtain the latency samples of the first link, the second link, and the third link;

[0026] Determine whether the number of delay samples of the first link is greater than or equal to a preset threshold. If yes, fit the delay samples of the first link with a Gaussian mixture model to obtain the delay distribution model of the first link. If no, fit the delay samples of the first link with a unimodal Gamma distribution to obtain the delay distribution model of the first link.

[0027] Determine whether the number of delay samples of the second link is greater than or equal to a preset threshold. If yes, fit the delay samples of the second link with a Gaussian mixture model to obtain the delay distribution model of the second link. If no, fit the delay samples of the second link with a unimodal Gamma distribution to obtain the delay distribution model of the second link.

[0028] Determine whether the number of delay samples of the third link is greater than or equal to a preset threshold. If yes, perform Gaussian mixture model fitting on the delay samples of the third link to obtain the delay distribution model of the third link. If no, use a unimodal Gamma distribution to fit the delay samples of the third link to obtain the delay distribution model of the third link.

[0029] Optionally, determining the cross-domain jitter hybrid core of the call path based on the first domain kernel, the second domain kernel, and the third domain kernel includes:

[0030] Based on the first domain kernel, the second domain kernel, and the third domain kernel, obtain the initial hybrid kernel for the call path;

[0031] Obtain the cumulative energy of the initial hybrid core; when the cumulative energy reaches a preset confidence level, set the cross-correlation value corresponding to the time after the current lag time in the initial hybrid core to zero, and obtain the cross-domain jitter hybrid core of the call path.

[0032] Optionally, the method for obtaining the causal residual and the effect residual includes:

[0033] Obtain the first indicator detection value of the upstream node at the target time and the second indicator detection value of the downstream abnormal node at the target time;

[0034] Based on the health baseline model of the upstream node, the expected value of the first indicator of the upstream node at the target time is obtained; based on the health baseline model of the downstream abnormal node, the expected value of the second indicator of the downstream abnormal node at the target time is obtained; the health baseline model of the node is used to describe the expected value of the node's indicator under a fault-free state.

[0035] Based on the expected value of the first indicator and the detected value of the first indicator, obtain the cause residual of the upstream node; based on the expected value of the second indicator and the detected value of the second indicator, obtain the effect residual of the downstream abnormal node.

[0036] Optionally, the method for obtaining the health baseline model of a node includes:

[0037] Obtain the health interval of the node; the health interval is a range of data whose indicator values ​​are less than or equal to the preset service level target SLO upper limit, and the health interval excludes the indicator values ​​corresponding to the fault window;

[0038] If the values ​​in the health interval are periodic and continuous, an autoregressive model is used to fit the values ​​in the health interval to obtain the health baseline model for that node.

[0039] If the base number of the values ​​in the health interval is less than or equal to a preset base number threshold and is a discrete value, Kalman filtering is used to fit the values ​​in the health interval to obtain the health baseline model of the node.

[0040] Optionally, methods for determining abnormal nodes include:

[0041] Determine the absolute value of the actual residual for each node in the microservice system;

[0042] Nodes whose actual residual absolute value is greater than or equal to a preset residual threshold are designated as instantaneous abnormal nodes.

[0043] Within the anomaly window, determine whether the absolute value of the actual residual of the instantaneous anomaly node continuously exceeds the preset residual threshold, or whether the average value of the absolute value of the actual residual of the instantaneous anomaly node exceeds the preset residual threshold;

[0044] If the absolute value of the actual residual of the instantaneous abnormal node continuously exceeds the preset residual threshold, or the average value exceeds the preset residual threshold, the instantaneous abnormal node is determined to be the abnormal node.

[0045] Secondly, embodiments of this application provide an anomaly root cause determination device applied to a microservice system, wherein the microservice system has a cloud-edge collaborative hybrid architecture, the cloud-edge collaborative hybrid architecture including a cloud data center and edge nodes, and the device includes:

[0046] The acquisition unit is used to acquire the latency impact kernel of the call path based on the cause residual of the upstream node at the target time and the effect residual of the downstream abnormal node at the target time.

[0047] The call path is the call path from the upstream node to the downstream abnormal node, and the call path includes a first link, a second link, and a third link; the first link is a link within the cloud data center, the second link is a link between the edge nodes, and the third link is a cross-domain link between the cloud data center and the edge nodes; the cause residual is the difference between the indicator detection value of the upstream node at the target time and the expected value of the first indicator; the effect residual is the difference between the indicator detection value of the downstream abnormal node at the target time and the expected value of the second indicator, and the abnormal window of the upstream node and the abnormal window of the downstream abnormal node are the same; the latency impact kernel is a response function that quantifies the propagation of the upstream node's abnormality over time and its impact on the downstream abnormal node;

[0048] The convolution processing unit is configured to perform convolution processing on the latency impact kernel using the latency distribution model of the first link to obtain a first domain kernel; perform convolution processing on the latency impact kernel using the latency distribution model of the second link to obtain a second domain kernel; and perform convolution processing on the latency impact kernel using the latency distribution model of the third link to obtain a third domain kernel.

[0049] A hybrid core determination unit is used to determine the cross-domain jitter hybrid core of the call path based on the first domain kernel, the second domain kernel, and the third domain kernel;

[0050] The root cause determination unit is used to determine the root cause of the abnormality of the downstream abnormal node based on the cross-domain jitter hybrid kernel.

[0051] Thirdly, embodiments of this application provide an electronic device, including:

[0052] Memory, used to store computer programs;

[0053] A processor for executing the computer program to implement the method as described in any one of the first aspects.

[0054] Fourthly, embodiments of this application provide a computer program that, when run on a computer, causes the computer to perform the method in any of the possible implementations of any of the above aspects.

[0055] Fifthly, embodiments of this application provide a computer-readable storage medium storing a computer program (also referred to as code or instructions) that, when run on a computer, causes the computer to perform the method in any of the possible implementations of any of the above aspects.

[0056] Sixthly, embodiments of this application provide a chip system including one or more processors for calling and executing instructions stored in memory, causing the methods in any of the above aspects or possible implementations to be executed. The chip system may be composed of chips or may include chips and other discrete devices.

[0057] This application provides an anomaly root cause localization method and apparatus, applied to a cloud-edge collaborative hybrid architecture microservice system, specifically executed by a root cause localization system. The method includes: obtaining the latency impact kernel of the call path based on the cause residual of the upstream node at the target time and the effect residual of the downstream abnormal node at the target time. The call path is the call path from the upstream node to the downstream abnormal node. In this application embodiment, the call path includes links within the cloud data center (i.e., the first link), links between edge nodes (i.e., the second link), and cross-domain links between the cloud data center and edge nodes (i.e., the third link). The anomaly window of the upstream node and the anomaly window of the downstream abnormal node are the same. The latency impact kernel is convolved using the latency distribution model of the first link to obtain a first domain kernel; the latency impact kernel is convolved using the latency distribution model of the second link to obtain a second domain kernel; and the latency impact kernel is convolved using the latency distribution model of the third link to obtain a third domain kernel. Based on the first domain kernel, the second domain kernel, and the third domain kernel, the cross-domain jitter hybrid kernel of the call path is determined. This cross-domain jitter hybrid kernel can accurately characterize the amplification effect of anomalies caused by cross-domain long-tail latency. Therefore, the root cause of anomalies determined based on this cross-domain jitter hybrid kernel helps to improve the accuracy of root cause localization. Attached Figure Description

[0058] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0059] Figure 1 This application provides a flowchart of an abnormal root cause determination method;

[0060] Figure 2 A flowchart illustrating a first method for obtaining causal residuals and effect residuals provided in this application embodiment;

[0061] Figure 3 A flowchart illustrating a method for a health baseline model and actual residuals of nodes provided in this application embodiment;

[0062] Figure 4 This application provides a flowchart of a method for obtaining a cross-domain jitter hybrid core;

[0063] Figure 5 A flowchart illustrating a method for locating the root cause of anomalies based on cross-domain jitter hybrid kernels, provided in this application embodiment;

[0064] Figure 6A flowchart illustrating a method for obtaining a candidate region of light cones, provided in an embodiment of this application;

[0065] Figure 7 This is a schematic diagram of an abnormal root cause determination device provided in an embodiment of this application. Detailed Implementation

[0066] To enable those skilled in the art to better understand the present application, the technical solutions in this embodiment will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0067] In the implementation of this application, the collection and processing of relevant data should strictly comply with the requirements of relevant laws and regulations, obtain the informed consent or separate consent of the personal information subject, and carry out subsequent data use and processing within the scope of laws and regulations and the authorization of the personal information subject.

[0068] First, the technical terms involved in the embodiments of this application will be explained.

[0069] (1) Light cone

[0070] In spacetime geometry, the light cone is used to delineate the region that an event can influence or be influenced by. The light cone includes the past light cone and the future light cone. The past light cone contains all historical events that can transmit information at the speed of light and act upon the event; the future light cone contains all future events that the event can influence within the constraints of information propagation. The region outside the light cone is causally isolated from the current event due to the limited propagation speed.

[0071] In this embodiment, the optical cone is specifically a quantile delay optical cone, which refers to an upstream node whose shortest path delay to a downstream abnormal node does not exceed the span of the abnormal window. The optical cone provided in this embodiment retains the causal reachability principle of the original physical concept and can be aligned with the actual network latency in microservice scenarios. Therefore, the optical cone can accurately locate the root cause of anomalies.

[0072] (2) Microservice system

[0073] A microservice system breaks down a traditional monolithic application into small, loosely coupled, and reusable service units (microservices) based on business functions. Each microservice focuses on a single business scenario and achieves cross-service communication through standardized interfaces to jointly complete complex business processes.

[0074] A microservice system comprises business microservices, communication middleware, and a monitoring and operations platform. Business microservices are the core functional units; for example, in an e-commerce system, business microservices might include order microservices, payment microservices, user management microservices, and data analysis microservices. The communication middleware is responsible for information transmission and request forwarding between the various microservices. The monitoring and operations platform is used to monitor the operational metrics of each microservice, enabling automatic service deployment, elastic scaling, and fault self-healing.

[0075] Current microservice systems often adopt a cloud-edge collaborative hybrid architecture. This architecture includes cloud data centers and edge nodes. Each microservice within a microservice system can be deployed in either the cloud data center or the edge node as needed. For example, taking an e-commerce microservice system, edge nodes can deploy an order microservice (referred to as the edge order microservice), a local communication proxy, and a local cache node, while the cloud data center can deploy a payment microservice (referred to as the cloud payment microservice), a message queue cluster (cloud message queue cluster), and a data analytics microservice.

[0076] The cloud-edge collaborative hybrid architecture allows microservice call chains to travel back and forth between the cloud data center and edge nodes. This means that any anomaly in an upstream node can be amplified and propagated to critical SLO nodes downstream within a short period of time through the communication middleware. For example, continuing with the e-commerce microservice system, after a user sends an order request, the corresponding microservice call chain is as follows: Edge Order Microservice → Local Communication Proxy → Cloud Message Queue Cluster → Cloud Payment Microservice → Cloud Message Queue Cluster → Edge Order Microservice → Critical SLO Node. If the local communication proxy at the edge node experiences a 10ms momentary delay due to network jitter, requests to the order microservice will accumulate on the local cache node, triggering the asynchronous RPC timeout retry mechanism and amplifying the request volume. The accumulated requests are synchronized to the cloud data center through the cloud message queue cluster, causing a surge in pressure on the cloud message queue cluster. This pressure then spreads to the cloud payment microservice, causing its response delay to exceed 50ms, for example, increasing to 200ms. This delay is then propagated back to the edge order microservice through the cloud message queue cluster, increasing order confirmation delays. Therefore, accurately locating the root cause is crucial to address this cross-domain anomaly propagation.

[0077] It should be noted that the cross-domain anomaly propagation provided in this application embodiment refers to the propagation of anomalies between cloud data centers and edge nodes.

[0078] Currently, the root cause of anomalies can be located using causal divergence or diffusion models. This method treats anomalies as events that propagate and then reverses the network structure to locate the root cause. However, this method suffers from insufficient accuracy in locating the root cause of anomalies.

[0079] The inventors discovered through analysis that the problem of insufficient accuracy in locating the root cause of anomalies using the causal divergence or diffusion model is that the method emphasizes the time or intensity of events and cannot accurately capture the amplification effect of anomalies caused by long-tail time delays across domains. This makes it impossible to trace the chain amplification path of anomalies, and consequently, it makes it impossible to locate the initial anomaly node that triggers the amplification effect, thus affecting the accuracy of anomaly root cause location.

[0080] In view of this, embodiments of this application provide an anomaly root cause localization method, which is applied to a microservice system with a cloud-edge collaborative hybrid architecture and can be executed by the root cause localization system.

[0081] The root cause localization system can perform the following operations: Based on the cause residual of the upstream node at the target time and the effect residual of the downstream abnormal node at the target time, obtain the latency impact kernel of the call path. The call path is the call path from the upstream node to the downstream abnormal node. In this embodiment, the call path includes links within the cloud data center (i.e., the first link), links between edge nodes (i.e., the second link), and cross-domain links between the cloud data center and edge nodes (i.e., the third link). The abnormal windows of the upstream node and the downstream abnormal node are the same. The latency impact kernel is convolved using the latency distribution model of the first link to obtain the first domain kernel; the latency impact kernel is convolved using the latency distribution model of the second link to obtain the second domain kernel; and the latency impact kernel is convolved using the latency distribution model of the third link to obtain the third domain kernel. Based on the first domain kernel, the second domain kernel, and the third domain kernel, determine the cross-domain jitter hybrid kernel of the call path. This cross-domain jitter hybrid kernel can accurately characterize the amplification effect of anomalies caused by cross-domain long-tail latency. Therefore, the root cause of anomalies determined based on this cross-domain jitter hybrid kernel helps to improve the accuracy of root cause localization.

[0082] In practical applications, a root cause analysis system can include a software system, which can be provided to the user as a software package for self-deployment, such as on a local physical server or in a private cloud. In some possible implementations, the load migration system can also be deployed in a public cloud and provided to the user as a cloud service. For example, a cloud service provider can offer a one-stop system service integrating the functions of the aforementioned core components.

[0083] The method for determining the root cause of anomalies provided in this application will be described below with reference to the accompanying drawings. It should be noted that this method is applied to a microservice system with a cloud-edge collaborative hybrid architecture, and the execution entity of this method will be illustrated using a root cause localization system as an example.

[0084] Appendix Figure 1 This application provides a flowchart of an anomaly root cause determination method, which includes the following steps:

[0085] S10: Based on the cause residual of the upstream node at the target time and the effect residual of the downstream abnormal node at the target time, obtain the latency impact kernel of the call path.

[0086] An upstream node refers to a node in a microservice call chain that precedes a downstream node in terms of business process and data transmission order. An anomaly in an upstream node may propagate through the microservice call chain, affecting the operational status of downstream nodes; therefore, it is a candidate for the root cause of an anomaly in a downstream node. In this embodiment, the upstream node of a downstream node can also be referred to as the cause node.

[0087] The target time refers to the sampling time at which a downstream anomaly is detected. For ease of description, the target time is referred to as time t.

[0088] The time-delay impact kernel is a function that quantifies the propagation of an upstream node anomaly over time and its influence on the downstream anomaly node's response. In this embodiment, the time-delay impact kernel describes the effect of the cause residual unit pulse input on the effect residual at different lag times. The time-delay impact kernel is used to capture the decay law of lag time and impact amplitude. Here, lag time refers to the time difference between the occurrence of an upstream node anomaly and the formation of an anomaly effect at the downstream anomaly node after a certain time delay.

[0089] The call path is the call path from the upstream node to the downstream abnormal node. In this embodiment, the call path includes a first link, a second link, and a third link. The first link is a link within the cloud data center, the second link is a link between the edge nodes, and the third link is a cross-domain link between the cloud data center and the edge nodes.

[0090] The causal residual is the difference between the upstream node's detected index value at the target time and the expected value of the first index. The causal residual measures the deviation between the upstream node's actual operating state and its expected normal state. The effect residual is the difference between the downstream abnormal node's detected index value at the target time and the expected value of the second index. The effect residual measures the deviation between the downstream abnormal node's actual operating state and its expected normal state. In this embodiment, both the upstream and downstream abnormal nodes are abnormal nodes, and their abnormal windows are the same, for example, both are abnormal nodes. .

[0091] It should be noted that, in this embodiment of the application, the indicator detection value of the node is specifically the detection value corresponding to the node's service indicator, which can be at least one of latency, error rate, or throughput.

[0092] To accurately obtain the cause residuals of upstream nodes and the effect residuals of downstream abnormal nodes, this application provides a method for obtaining these residuals. Figure 2A flowchart of a first method for obtaining causal residuals and effect residuals provided in this application embodiment is shown. The method includes the following steps:

[0093] S101: The root cause localization system obtains the health baseline model of the upstream node and the health baseline model of the downstream abnormal node.

[0094] A node health baseline model describes the expected values ​​of a node's metrics when there are no failures. Node metrics can be service metrics, such as request processing latency, throughput, misalignment rate, or number of failed API calls.

[0095] In this embodiment of the application, the root cause localization system can build a corresponding health baseline model for each node in the microservice system. The specific construction method includes:

[0096] Step 1: Obtain the health range of the node.

[0097] In this embodiment, the health interval of a node includes the values ​​of multiple indicators, each of which has a value less than or equal to a preset SLO upper limit. Furthermore, the health interval excludes indicator values ​​corresponding to fault windows. Specifically, the root cause localization system can use the monitoring logs from the nearest D days (up to time t) as the original dataset, filter out indicator values ​​exceeding the preset SLO upper limit, and remove fault windows registered by the monitoring and maintenance platform. The remaining continuous data segments constitute the health interval of the node.

[0098] It should be noted that D is a number greater than 0, for example, D is 10. The value of D can be adjusted as needed. The preset SLO upper limit is a pre-set upper limit value, which can be adjusted as needed. For example, the SLO upper limit for latency-related indicators is 50ms, and the SLO upper limit for error-related indicators can be set to 1%.

[0099] Step 2: If the values ​​in the health interval are periodic and continuous, an autoregressive model can be used to fit the values ​​in the health interval to obtain the health baseline model for that node. If the base number of the values ​​in the health interval is less than or equal to a preset base number threshold and is a discrete value, a Kalman filter can be used to fit the values ​​in the health interval to obtain the health baseline model for that node.

[0100] For example, metrics such as request processing latency and throughput exhibit periodicity and continuity. An autoregressive model can be used to fit the values ​​within the health interval to obtain the health baseline model for that node. Similarly, metrics such as misalignment rate or the number of failed API calls have values ​​within the health interval that are less than or equal to a preset threshold and are discrete. A Kalman filter model can be used to fit the values ​​within the health interval to obtain the health baseline model for that node.

[0101] In practical implementation, the autoregressive model can be an autoregressive model with a seasonal term, such as SARIMA, and the Kalman filter can be a Beta-Binomial Kalman filter. The preset threshold number is a positive integer and can be adjusted as needed; for example, the preset threshold number can be 20.

[0102] This application's embodiments employ an autoregressive model with a seasonal term for fitting periodic and continuous interval values, accurately capturing periodic trends and seasonal fluctuations. Therefore, the constructed health benchmark model accurately reflects the true changing patterns of the indicators. For interval values ​​with small and discrete base values, Kalman filtering is used for fitting, which helps solve the problem of fitting distortion to low-base data and improves the accuracy of characterizing the sparse fluctuation patterns of the indicators.

[0103] Furthermore, in this embodiment, the root cause localization system can re-estimate the most recent fixed-duration health segment at first intervals to obtain a baseline model updated over time. The first interval can be adjusted as needed, for example, to 30 minutes, and the fixed interval can also be adjusted as needed, for example, to 24 hours.

[0104] Furthermore, in this embodiment, the health baseline model of a node can be stored in a data warehouse for real-time invocation during the operation of the microservice system. In one implementation, the parameter vector, prediction error variance, distribution fitting time, and the model used for fitting of each node can be stored in the data warehouse. Thus, during operation, based on the node and its monitoring indicators, the corresponding fitting model and parameter vector are retrieved from the data warehouse, and the parameter vector is input into the model to obtain the health baseline model.

[0105] S102: The root cause localization system determines the expected value of the first indicator at the target time through the health baseline model of the upstream node, and determines the expected value of the second indicator through the health baseline model of the downstream abnormal node.

[0106] In practical use, the target time can be substituted into the health baseline model of the upstream node to obtain the expected value of the first indicator at the target time, and then substituted into the health baseline model of the downstream abnormal node to obtain the expected value of the second indicator.

[0107] In one specific implementation, if the parameter vector, prediction error variance, distribution fitting time, and fitting algorithm type of each node are stored in a data warehouse, the root cause localization system can first determine the monitoring indicator type of the node, call the corresponding health baseline model based on the monitoring indicator type of the node, and input the parameter vector into the model. Then the model can obtain the expected value of the indicator at the target time.

[0108] For example, if the monitoring metric for a node is request processing latency, a seasonal autoregressive model can be invoked. The parameter vectors of the nodes stored in the data warehouse can be input into the seasonal autoregressive model, and then the model can output the expected value of the metric at time t.

[0109] S103: The root cause localization system obtains the first indicator detection value of the upstream node at the target time and the second indicator detection value of the downstream abnormal node at the target time.

[0110] It should be noted that S101 and S103 can be executed simultaneously, or S103 can be executed first and then S101, or S101 can be executed first and then S103. This application embodiment does not limit the execution of S101.

[0111] S104, the root cause localization system obtains the cause residual of the upstream node based on the first indicator detection value and the first indicator expectation value, and determines the effect residual of the downstream abnormal node based on the second indicator detection value and the second indicator expectation value.

[0112] In one example, the causal residual is the difference between the detected value and the expected value of the first indicator, and the effect residual is the difference between the detected value and the expected value of the second indicator. For instance, if the indicator is request processing latency, and the detected value of the first indicator for upstream node u at time t is 43ms, and the detected value of the second indicator for downstream abnormal node v at time t is 200ms, the expected value of the first indicator is 35ms, and the expected value of the second indicator is 140ms, then the causal residual is 8ms, and the effect residual is 60ms.

[0113] In another example, the causal residual is the standardized value of the prediction error variance of the upstream node divided by a first difference. Here, the first difference is the difference between the detected value and the expected value of the first indicator. The effect residual is the standardized value of the prediction error variance of the downstream anomalous node divided by a second difference. The second difference is the difference between the detected value and the expected value of the second indicator. This method ensures that the residuals of different nodes and different indicators are on a uniform scale, facilitating comparison of candidate thresholds.

[0114] Furthermore, embodiments of this application provide a method for obtaining the latency impact kernel of a call path. This method extracts the latency impact kernel through time-series signal analysis of residuals. Specifically, the root cause localization system can obtain the cross-correlation curves corresponding to the lag residuals and the cause residuals; the lag residual is the lag time of the effect residual. The resulting residual. For example, the residual of the downstream abnormal node v is... The corresponding lagged residual The reason is that the residual is the cross-correlation curve obtained from the residual corresponding to the upstream node u. As shown in the following formula (1):

[0115] (1)

[0116] Where N represents the actual number of valid summation points. This can be understood as the cross-correlation curve... The dependent variable is the cross-correlation value between the lagged residuals and the causal residuals, and the independent variable of the cross-correlation curve is the lag time. Obtain the maximum cross-correlation value and the corresponding baseline lag time from the cross-correlation curve. After the baseline lag time, determine the target lag time corresponding to the first decrease in the cross-correlation value in the cross-correlation curve to the target correlation value. ; Extract the target interval curve from the cross-correlation curve as the latency impact kernel of the call path; the target interval is [ This approach ensures that the acquired time-delay impact kernels demonstrate the correlation strength of different lag times between cause and effect, and ensures that only the periods of significant impact are covered.

[0117] Furthermore, this application also provides a method for determining abnormal nodes, used to accurately locate abnormal nodes. This method includes: determining the absolute value of the actual residuals of each node in the microservice system; identifying nodes whose absolute residuals are greater than or equal to a preset residual threshold as instantaneous abnormal nodes; determining whether, within an abnormal window, the absolute value of the actual residuals of the instantaneous abnormal nodes continuously exceeds the preset residual threshold, or whether the average value of the absolute residuals of the instantaneous abnormal nodes exceeds the preset residual threshold; if the absolute value of the actual residuals of the instantaneous abnormal nodes continuously exceeds the preset residual threshold, or the average value exceeds the preset residual threshold, then the instantaneous abnormal node is determined to be an abnormal node. This method, compared to directly specifying abnormal nodes, can improve the accuracy of abnormal node location. Moreover, this application only requires calculating the latency impact kernel of the upstream and downstream abnormal nodes, thus significantly reducing computational complexity and improving computational efficiency.

[0118] It should be noted that the preset residual threshold provided in this application embodiment can be adjusted as needed, and this application embodiment is not limited thereto.

[0119] S20, the latency impact kernel is convolved using the latency distribution model of the first link to obtain a first domain kernel; the latency impact kernel is convolved using the latency distribution model of the second link to obtain a second domain kernel; the latency impact kernel is convolved using the latency distribution model of the third link to obtain a third domain kernel.

[0120] In this embodiment, the link latency distribution model refers to the link latency probability density distribution model, which is used to quantify the latency fluctuation pattern and long-tail characteristics of the link. In this embodiment, the root cause localization system can obtain latency samples of the first link, the second link, and the third link from the distributed tracing log. Through probability distribution fitting, the latency samples of the first link, the second link, and the third link are fitted respectively to obtain the latency distribution models of the first link, the second link, and the third link.

[0121] Furthermore, to address the fitting accuracy issue caused by variations in the number of link delay samples, the root cause localization system can also determine the link delay distribution model based on whether the number of link delay samples is greater than or equal to a preset threshold. If so, a Gaussian mixture model is applied to the link delay samples to obtain the link delay distribution model, thus adapting to the multi-peak characteristics of link delay. If not, a unimodal Gamma distribution is used to fit the link delay samples to obtain the link delay distribution model, avoiding overfitting of the Gaussian mixture model due to insufficient samples. The preset threshold can be adjusted as needed, for example, a preset threshold of 500 samples.

[0122] Specifically, the root cause localization system can determine whether the number of delay samples in the first link is greater than or equal to a preset threshold. If so, it performs a Gaussian mixture model fitting on the delay samples of the first link to obtain the delay distribution model of the first link; otherwise, it uses a unimodal Gamma distribution to fit the delay samples of the first link to obtain the delay distribution model of the first link. Similarly, it can determine whether the number of delay samples in the second link is greater than or equal to a preset threshold. If so, it performs a Gaussian mixture model fitting on the delay samples of the second link to obtain the delay distribution model of the second link; otherwise, it uses a unimodal Gamma distribution to fit the delay samples of the second link to obtain the delay distribution model of the second link. Finally, it can determine whether the number of delay samples in the third link is greater than or equal to a preset threshold. If so, it performs a Gaussian mixture model fitting on the delay samples of the third link to obtain the delay distribution model of the third link; otherwise, it uses a unimodal Gamma distribution to fit the delay samples of the third link to obtain the delay distribution model of the third link.

[0123] It should be noted that the Gaussian mixture model provided in this application embodiment can be a 2-3 component Gaussian mixture model, or other Gaussian models, and this application embodiment is not limited thereto.

[0124] The latency impact kernel is described below. In this embodiment, the latency impact kernel obtained in step S10 includes a lot of noise, which interferes with the generation of the domain kernel. Therefore, to address the interference of noise in the latency impact kernel (hereinafter referred to as the preliminary latency impact kernel) on the generation of the domain kernel, this application can optimize the preliminary latency impact kernel by denoising it using regularized deconvolution, and then perform domain-specific convolution. For ease of explanation, the call path below is taken as the call path from upstream node v to downstream abnormal node v, and the preliminary latency impact kernel of this call path is... Regularized deconvolution affects the initial time delay kernel The noise reduction and optimization process will be explained below:

[0125] residuals of cause and residual effect Perform convolution matrix processing separately, due to residuals The corresponding convolution matrix is ​​the X matrix, and the effect residual is... The corresponding convolution matrix is ​​y, which satisfies formula (2):

[0126] (2)

[0127] in, This represents observation noise. After matrixing, the convolution operation involving sliding multiplication and summation can be transformed into ordinary matrix multiplication for subsequent processing.

[0128] Then, a first-order difference matrix and a smoothed least squares objective function are introduced. The first-order difference matrix... for:

[0129]

[0130] The objective function for smooth least squares is shown in formula (3):

[0131] (3)

[0132] in, These are fixed parameters that can be adjusted as needed, for example... The value is 0.01. The first term requires convolution. Try to reproduce The second term suppresses the jagged edges between adjacent kernel coefficients. We put... As the starting point for optimization, the solution is obtained through the conjugate gradient iteration method.

[0133] Received Scanning from the peak to the right: If the h value is lower than the target value for f consecutive sampling points, it is considered that the effective energy has decayed. The tail is truncated at this position to obtain the refined time delay influence kernel. f is a positive integer, for example, f=3. The target value is related to the peak value, for example, target value = 15% × peak value. Nuclear energy is also calculated simultaneously. As shown in formula (4):

[0134] (4)

[0135] in, To refine the impact of time delay on the core The effective time range, also known as the support length. Refined delay affects the core. It has been smoothed, denoised, and energy calibrated.

[0136] Subsequently, the root cause localization system can use the delay distribution model of the first link to perform convolution processing on the refined delay impact kernel to obtain the first domain kernel; use the delay distribution model of the second link to perform convolution processing on the refined delay impact kernel to obtain the second domain kernel; and use the delay distribution model of the third link to perform convolution processing on the refined delay impact kernel to obtain the third domain kernel.

[0137] It should be noted that the domain kernel can also be determined in other ways in the embodiments of this application, and the embodiments of this application are not limited thereto.

[0138] S30, determine the cross-domain jitter hybrid core of the call path based on the first domain kernel, the second domain kernel and the third domain kernel.

[0139] Specifically, the root cause localization system can integrate the first domain kernel, the second domain kernel, and the third domain kernel to determine the cross-domain jitter hybrid kernel of the call path.

[0140] Furthermore, the root cause localization system can also generate a cross-domain jitter hybrid kernel by weighted fusion of three intra-domain classes based on link time proportions. This allows the cross-domain jitter hybrid kernel to accurately adapt to heterogeneous cloud-edge links and cover the amplification effect of abnormal propagation caused by cross-domain long tails. Specifically, the root cause localization system first determines the first time ratio of the first link in the complete call of the call path, the second time ratio of the second link in the complete call of the call path, and the third time ratio of the third link in the complete call of the call path. Then, the sum of the first product, the second product, and the third product is used as the cross-domain jitter hybrid kernel. Here, the first product is the product of the first time ratio and the first domain kernel, the second product is the product of the second time ratio and the second domain kernel, and the third product is the product of the third time ratio and the third domain kernel.

[0141] Furthermore, the root cause localization system can obtain the initial hybrid kernel of the call path based on the first domain kernel, the second domain kernel, and the third domain kernel; obtain the cumulative energy of the initial hybrid kernel; and when the cumulative energy reaches a preset confidence level, set the cross-correlation value corresponding to the time after the current lag time in the initial hybrid kernel to zero, thus obtaining the cross-domain jitter hybrid kernel of the call path. For example, the confidence level is... Thus, the cross-domain jitter hybrid kernel is obtained while retaining the shape of the main cross-correlation peak, and adaptively extends to cover 95% of the propagation energy long tail.

[0142] S40, based on the cross-domain jitter hybrid kernel, determine the root cause of the abnormality of the downstream abnormal node.

[0143] The cross-domain jitter hybrid core includes the core latency of abnormal propagation, the attenuation trend of impact intensity, and the coverage range of long-tail latency. Therefore, the embodiments of this application can determine the root cause of abnormality in downstream abnormal nodes based on the cross-domain jitter hybrid core.

[0144] For example, the root cause localization system first determines the cross-domain jitter mixing kernel of possible upstream nodes to identify candidate nodes. Then, it obtains the abnormal sequence of each candidate node in the period before the downstream abnormal window, convolves the abnormal sequence with the corresponding cross-domain jitter mixing kernel to obtain the response sequence of the candidate node's abnormality propagating to the downstream abnormal node, calculates the similarity between the response sequence and the abnormal sequence of the actual downstream abnormal node, and sorts them according to the similarity to obtain the root cause of the abnormality.

[0145] In summary, the anomaly root cause determination method provided in this application obtains the latency impact kernel of the call path based on the cause residual of the upstream node at the target time and the effect residual of the downstream abnormal node at the target time. The call path is the call path from the upstream node to the downstream abnormal node. In this application embodiment, the call path includes links within the cloud data center (i.e., the first link), links between edge nodes (i.e., the second link), and cross-domain links between the cloud data center and edge nodes (i.e., the third link). The anomaly windows of the upstream node and the downstream abnormal node are the same. The latency impact kernel is convolved using the latency distribution model of the first link to obtain the first domain kernel; the latency impact kernel is convolved using the latency distribution model of the second link to obtain the second domain kernel; and the latency impact kernel is convolved using the latency distribution model of the third link to obtain the third domain kernel. Based on the first domain kernel, the second domain kernel, and the third domain kernel, the cross-domain jitter hybrid kernel of the call path is determined. This cross-domain jitter hybrid kernel can accurately characterize the amplification effect of anomalies caused by cross-domain long-tail latency. Therefore, the root cause of anomalies determined based on this cross-domain jitter hybrid kernel helps to improve the accuracy of root cause localization.

[0146] The following, in conjunction with specific embodiments, illustrates... Figure 3 ~Appendix Figure 5 This application introduces a method for determining the root cause of anomalies, based on embodiments of the present application. (Appendix) Figure 3 A flowchart illustrating a method for a health baseline model and actual residuals of nodes, provided in an embodiment of this application. (Attached) Figure 4 This application provides a flowchart of a method for obtaining a cross-domain jitter hybrid core, as illustrated in the embodiments of this application. Figure 5 This document provides a flowchart of a method for locating the root cause of anomalies based on cross-domain jitter hybrid kernels, as illustrated in an embodiment of this application. It should be noted that in this embodiment, the target time is time t, and the health baseline model is a sliding baseline model, updated over time, as an example. Furthermore, to improve positioning accuracy, the residuals are normalized to unify the processing units.

[0147] like Figure 3 As shown, the method includes steps S310 to S340.

[0148] S310: The root cause localization system obtains the health range of a node.

[0149] In this embodiment, the health interval of a node includes the values ​​of multiple indicators, each of which has a value less than or equal to a preset SLO upper limit. Furthermore, the health interval excludes indicator values ​​corresponding to fault windows. In this embodiment, the root cause localization system uses the monitoring logs from the nearest D days to time t as the original dataset. It filters out indicator values ​​in the original dataset that exceed the preset SLO upper limit and removes fault windows registered by the monitoring and maintenance platform. The remaining continuous data segments constitute the health interval of the node.

[0150] It should be noted that D is a number greater than 0, for example, D is 10. The value of D can be adjusted as needed. The preset SLO upper limit is a pre-set upper limit value, which can be adjusted as needed. For example, the SLO upper limit for latency-related indicators is 50ms, and the SLO upper limit for error-related indicators can be set to 1%.

[0151] S320: Baseline model fitting.

[0152] If the values ​​in the health interval are periodic and continuous, an autoregressive model can be used to fit the values ​​in the health interval to obtain the health baseline model for that node. If the base value of the values ​​in the health interval is less than or equal to a preset base value threshold and is a discrete value, a Kalman filter can be used to fit the values ​​in the health interval to obtain the health baseline model for that node.

[0153] For example, metrics such as request processing latency and throughput exhibit periodicity and continuity. An autoregressive model can be used to fit the values ​​within the health interval to obtain the health baseline model for that node. Similarly, metrics such as misalignment rate or the number of failed API calls have values ​​within the health interval that are less than or equal to a preset threshold and are discrete. A Kalman filter model can be used to fit the values ​​within the health interval to obtain the health baseline model for that node.

[0154] In practical implementation, the autoregressive model can be an autoregressive model with a seasonal term, such as SARIMA, and the Kalman filter can be a Beta-Binomial Kalman filter. The preset threshold number is a positive integer and can be adjusted as needed; for example, the preset threshold number can be 20.

[0155] This application's embodiments employ an autoregressive model with a seasonal term for fitting periodic and continuous interval values, accurately capturing periodic trends and seasonal fluctuations. Therefore, the constructed health benchmark model accurately reflects the true changing patterns of the indicators. For interval values ​​with small and discrete base values, Kalman filtering is used for fitting, which helps solve the problem of fitting distortion to low-base data and improves the accuracy of characterizing the sparse fluctuation patterns of the indicators.

[0156] In one example, the root cause localization system specifically outputs the parameter vector and prediction error variance corresponding to the health baseline model for each node. The parameter vector is the core representation factor of the health baseline model, containing all the key parameters of the model and directly determining its shape and prediction accuracy. The values ​​of the parameter vector are obtained by training on historical operational data under the node's health state and are crucial for adapting the model to the node's health characteristics. The prediction error variance is a validity metric for the health baseline model, reflecting the degree of deviation between the model's predicted values ​​and the actual observed values ​​under the node's health state. It measures the model's fitting accuracy to the node's health characteristics and provides a threshold basis for anomaly detection.

[0157] S330: Root cause localization model sliding reestimation.

[0158] The root cause localization system re-estimates the health status by taking the most recent new health segment of a preset fixed duration at the first interval. This forms a sliding health benchmark model. The first duration can be adjusted as needed, for example, to 30 minutes, and the fixed duration can also be adjusted as needed, for example, to 24 hours.

[0159] In this embodiment of the application, information corresponding to the health baseline model of each node can be stored. For example, the node... The following information is stored in a data warehouse for real-time forecasting during runtime:

[0160] Where v is the identifier of node v, used to uniquely identify the node. For baseline fitting model, The baseline fitting time, Let v be the parameter vector of node v. Let V be the prediction error variance for node v. Thus, during operation, based on the node and its monitoring metrics, the corresponding fitting model and parameter vectors are retrieved from the data warehouse, and the parameter vectors are input into the model to obtain the healthy baseline model.

[0161] S340, the root cause localization system determines the actual residuals of nodes based on the node's health baseline model.

[0162] In one specific application, after the microservice system enters the runtime phase, sampling is performed at fixed intervals. The system continuously receives real-time monitoring values ​​from each node. At time t, the root cause localization system provides the expected value of the node based on its health baseline model. For example, based on the health baseline model of node v, it provides the expected value of node v. Root cause localization systems use the difference between a node's actual observed value and its expected value as the actual residual at time t. For example, the actual observed value of node v... With this expected value Subtracting the two yields the actual residual at time t. As in formula (5):

[0163] (5)

[0164] Prediction error variance Standardize the actual residual at time t to obtain the standardized residual. As in formula (6):

[0165] (6)

[0166] This ensures that the residuals of different nodes and indicators are on a uniform scale, facilitating subsequent threshold comparisons. Furthermore, for ease of use, the root cause localization system will also include <v, t, , > Store the residual cube to form an anomaly sequence that is updated over time.

[0167] Furthermore, in order to automatically identify the real abnormal segments in the continuous monitoring stream, a dual detection strategy is applied to the standardized residuals, which specifically includes the following: a point-level anomaly detection strategy and a window-level aggregation strategy.

[0168] Specifically, the point-level anomaly detection strategy involves quantifying transient anomalies by utilizing the degree of deviation from the normal baseline. Specifically, control limits are set. .when When the sampling time is t, it is considered that node v has a transient anomaly. For numbers greater than 0, adjustments can be made as needed, for example... The value is 3.

[0169] For example, suppose the health baseline model of an e-commerce edge caching service has been trained, and set... =3, the real-time acquisition delay during operation is 210ms; the expected delay of the healthy baseline model is 180ms, the variance of the prediction error during the healthy period is 100, the residual is 30ms, and the standardized residual is 3. Since the standardized residual does not exceed 3, this indicator is considered to be within the normal range.

[0170] Considering that sporadic noise may cause single-point exceedances, this application embodiment filters sporadic noise by exceeding the threshold in continuous intervals. Specifically, it targets continuous intervals whose length exceeds a preset length (e.g., a preset length of 3 sampling periods). , Continue to exceed or its average value exceeds In this interval The window has been identified as abnormal.

[0171] Detected anomaly windows are added to the event stream. Each event is recorded in the event stream with the following information: node, start time of the anomaly window, end time of the anomaly window, monitoring index type, and maximum actual error. These events provide temporal boundaries for subsequent light cone candidate domains and limit the observation range that needs to be interpreted in sparse inversion, achieving a crucial transition from raw monitoring streams to structured anomaly information.

[0172] Appendix Figure 4 This application provides a flowchart of a method for obtaining a cross-domain jitter hybrid core, as shown in the embodiments. Figure 4 As shown, this section includes S410~S450:

[0173] S410: The root cause localization system extracts the initial time delay effect kernel from the original residual signal.

[0174] The time-delay impact kernel is a function that quantifies the propagation of an upstream node anomaly over time and its influence on the downstream anomaly node's response. In this embodiment, the time-delay impact kernel describes the effect of the cause residual unit pulse input on the effect residual at different lag times. The time-delay impact kernel is used to capture the decay law of lag time and impact amplitude. Here, lag time refers to the time difference between the occurrence of an upstream node anomaly and the formation of an anomaly effect at the downstream anomaly node after a certain time delay.

[0175] For ease of description, we will use the upstream node u as the cause node and the downstream abnormal node v as the result node as an example. The standardized residual is obtained by standardizing the actual residual (called the causal residual) of the upstream node u. The standardized residual is obtained by standardizing the actual residual (effect residual) of the downstream abnormal nodes. See [link to documentation] for details on how to obtain it. Figure 3 As shown.

[0176] In this embodiment of the application, the root cause localization system for and First, perform a first-order difference and normalize the result to zero mean and unit variance to obtain the first target sequence. .in, and It is a zero-mean, approximately stationary sequence. This ensures that subsequent cross-correlation analysis focuses on short-term fluctuations rather than long-term trends.

[0177] Among them, the upstream and downstream abnormal nodes share a common abnormal time window. .

[0178] S420, the root cause localization system determines the cross-correlation curve.

[0179] Specifically, the root cause localization system can obtain the cross-correlation curves corresponding to the lag residuals and the causal residuals; the lag residual is the lag time of the effect residual. The resulting residual. For example, The lagged residual is In the non-negative lag interval The cross-correlation curves are calculated as shown in formulas (7) to (8):

[0180] (7)

[0181] N (8)

[0182] L0 represents the effective time range of the actual residual. Lag for the second target sequence The corresponding residuals. The cross-correlation curve shows the correlation strength of "cause-lead-effect" at different time lags. Specifically, the cross-correlation curve... The dependent variable is the cross-correlation value between the lagged residuals and the causal residuals, and the independent variable of the cross-correlation curve is the lag time. .

[0183] S430, the root cause localization system extracts the main peak of the cross-correlation function and performs kernel normalization to obtain the preliminary time delay influence kernel.

[0184] Root cause localization system in The region is used to obtain the baseline lag time at which the global maximum peak of the cross-correlation curve is located. ;by Centered on, the cross-correlation curves at The initial time delay impact kernel is obtained by truncating the interval and normalizing it according to the peak value. As shown in formula (9):

[0185] (9)

[0186] in, The target lag time corresponding to the first decrease of the cross-correlation value in the cross-correlation curve to the target correlation value. For example, the lag time corresponding to when the cross-correlation value first drops to 10% of its peak. Normalization ensures that the numerical scale of the kernel is comparable; truncation ensures that its support length only covers the period of significant influence.

[0187] Furthermore, if the global maximum peak occurs after a negative hysteresis (i.e., If the cross-correlation curve shows no significant peak, it is concluded that the two sequences lack unidirectional causal indications, and kernel estimation of the signal pair is stopped. Thus, a preliminary time-delay-influence kernel is obtained. and the corresponding benchmark lag time Target lag time And peak amplitude value.

[0188] S440, the root cause localization system uses regularized deconvolution to process the initial delay impact kernel to obtain the refined delay impact kernel.

[0189] For ease of explanation, the call path below refers to the call path from upstream node v to downstream abnormal node v, and the initial latency impact core of this call path is... Regularized deconvolution affects the initial time delay kernel The noise reduction and optimization process will be explained. In the embodiments of this application, a first target sequence is used. Second target sequence This will be used as a benchmark for explanation.

[0190] The specific processing method includes the following steps:

[0191] Step A1: For and Perform convolution matrix conversion.

[0192] In the error window In the middle, let the support length be... .right and Perform convolution matrix processing. In this embodiment of the application, the first target sequence is... Constructing Toeplitz convolution matrices ,in ); for the second target sequence Constructing the Toeplitz convolution matrix And extract the first line to get y.

[0193] For example, the first target sequence The corresponding X matrix is ​​as follows:

[0194]

[0195] Second target sequence The corresponding y-matrix is ​​as follows:

[0196]

[0197] in, .

[0198] in It is the initial delay that needs to be corrected that affects the kernel. This represents observation noise. After matrixing, the convolution operation involving sliding multiplication and summation can be transformed into ordinary matrix multiplication for subsequent processing.

[0199] Step A2: Least squares with smoothing regularization

[0200] We introduce a first-order difference matrix and a smoothed least squares objective function. The first-order difference matrix D is:

[0201]

[0202] The objective function for smooth least squares is:

[0203]

[0204] in, These are fixed parameters that can be adjusted as needed, for example... The value is 0.01. The first term requires convolution. Try to reproduce The second term suppresses the jagged edges between adjacent kernel coefficients. We put... Normalized As the starting point for optimization, the solution is obtained through the conjugate gradient iteration method.

[0205] Step A3: Post-processing

[0206] Root cause localization system Scanning from the peak to the right: If the h value is lower than the target value for f consecutive sampling points, it is considered that the effective energy has decayed. The tail is truncated at this position to obtain the refined time delay influence kernel. f is a positive integer. For example, f=3. The target value is related to the peak value. For example, the target value = 15% × the peak value.

[0207] Furthermore, the root cause localization system, upon obtaining... First normalize by the highest point, so that Then, scanning from the peak to the right: if the kernel value is 5% lower than the peak value for three consecutive sampling points, it is considered that the effective energy has decayed, and the tail is truncated at that position to obtain the final length. . To refine the impact of time delay on the core The effective time range, also known as the support length.

[0208] Root cause localization system calculates nuclear energy As in formula (10):

[0209] (10)

[0210] The kernel is affected by the refined time delay. The details are as follows:

[0211]

[0212] Among them, the refinement of latency affects the core. It has been smoothed, denoised, and energy calibrated.

[0213] S450, Root Cause Localization System acquires cross-domain jitter hybrid core.

[0214] Refined nucleus This only represents the average propagation shape; in a cloud-edge collaborative hybrid architecture, the call path will traverse three network segments: cloud-internal, edge-internal, and cloud-edge, each with different jitter. By... The kernel is convolved with three types of delay distributions, then linearly fused according to the link proportion, and the tail is pruned using a cumulative energy threshold (e.g., 95% energy threshold) to obtain a cross-domain jitter hybrid kernel. .

[0215] First, we will introduce the latency distribution model of the link.

[0216] Root cause localization systems can extract three types of latency samples from distributed tracing logs. , , .in, This is a delay sample for the first link. This is a delay sample for the second link. This is a time delay sample for the third link. The number of time delay samples in each category is determined. For time delay samples with a sample size greater than or equal to a preset threshold, a 2–3 component Gaussian mixture fitting is used. For time delay samples with a sample size less than the preset threshold, a unimodal Gamma fitting is used. The resulting link time delay distribution model, i.e., the density function, is derived. As shown in formulas (11) to (12):

[0217] (11)

[0218] (12)

[0219] In this embodiment of the application, the root cause localization system will Convolution is performed with each of the three types of delay distributions, and the kernels of the three types of domains are obtained as shown in formula (13):

[0220] (13)

[0221] Next, the root cause localization system uses tracing logs to calculate the time percentage of the three types of links in the complete call path. ),in, This represents the percentage of time corresponding to the first link. This represents the second time segment corresponding to the second link. This represents the third time segment corresponding to the third link. For example... =0.3:0.3:0.4.

[0222] Cross-domain jitter hybrid core As in formula (14):

[0223] (14)

[0224] Then Uniform scaling enables .

[0225] Furthermore, the root cause localization system can also calculate the cumulative energy curve. As in formula (15):

[0226] (15)

[0227] Determine the minimum make and will The tail of the core is set to zero; the final support length for .in The sampling interval is denoted as .

[0228] Furthermore, in the embodiments of this application, the root cause localization system can also detect cross-domain jitter hybrid kernels. The specific storage format is as follows: .

[0229] In this embodiment of the application, the cross-domain jitter hybrid core The cross-correlation peak shape is preserved, and a long tail adaptively covers 95% of the propagation energy. In the embodiments of this application, a cross-domain dithering hybrid kernel is used. The path propagation delay can be correlated with the propagation model determined by the candidate domain of the light cone. This approach avoids missing the effects of slow cross-domain propagation and also controls the scale of subsequent computations.

[0230] Appendix Figure 5 This document provides a flowchart of a method for locating the root cause of anomalies based on a cross-domain jitter hybrid kernel, as illustrated in an embodiment of this application. Figure 5 As shown, this step includes S510~S540.

[0231] This application embodiment can also determine the candidate domain of the optical cone where runtime anomalies occur based on a cross-domain jitter hybrid model. A convolutional dictionary is constructed within the candidate domain. Sparse inversion is used to process the convolutional dictionary to obtain the root cause nodes of the anomalies. By compressing the candidate range from all upstream nodes to the local feasible region through the optical cone candidate domain, and then using sparse inversion to process the transitive kernel dictionary determined based on the optical cone candidate domain, the root cause nodes of the anomalies are obtained. This method avoids the candidate explosion problem caused by full graph traversal and improves search speed. Furthermore, by retaining upstream nodes that can influence the anomaly nodes within the quantile delay, the search space is compressed from the source, reducing inverse causal nodes and improving the accuracy of anomaly root cause localization.

[0232] S510, the root cause localization system determines the candidate optical cone domain corresponding to the downstream abnormal node.

[0233] Nodes in the optical cone candidate domain are upstream nodes whose shortest path quantile delay to downstream anomalous nodes does not exceed the span of the anomalous window. Upstream nodes in the optical cone candidate domain are considered to be within the range where they can influence downstream anomalous nodes.

[0234] This application provides a method for obtaining a candidate region of light cones. This method calculates the quantile propagation delay from an upstream node to an abnormal downstream node and then prunes it according to an adaptive threshold. The following is a detailed description in conjunction with the appendix. Figure 6 Explanation. (Attached) Figure 6 A flowchart of a method for obtaining a candidate region of light cones is provided in this application embodiment. The method includes the following:

[0235] S6100, obtain the quantization propagation delay from any upstream node to the abnormal downstream node and the corresponding optical cone front time.

[0236] In the application embodiment, the quantile propagation delay from the upstream node to the abnormal downstream node refers to the maximum time required for abnormal information to travel from the upstream node to the downstream abnormal node under a preset confidence level (e.g., 95% confidence level). If the quantile propagation delay does not exceed the abnormal window span, it indicates that the abnormality of the upstream node is sufficient to affect the downstream abnormal node. If the quantile propagation delay exceeds the abnormal window span, it indicates that the upstream node is unreachable in terms of timing and does not belong to the optical cone candidate domain.

[0237] The embodiments of this application obtain the quantization propagation delay through the following steps:

[0238] Step B1: Enumerate topology paths.

[0239] Obtain the service dependency graph from the microservice registry or distributed tracing system of the microservice system. In this context, V represents a node, signifying an independent functional unit within the microservice system. This can be a service node in a cloud data center, a service instance on an edge node, or a key component such as middleware (e.g., RPC services or message queues) and databases. E represents an edge, signifying the call relationship or data transfer path between nodes.

[0240] Using the downstream abnormal node v as the sink, traverse the service dependency graph in reverse. By filtering out all upstream nodes u within the upstream k layers, we obtain all directed paths (referred to as topological paths) that do not contain loops and have a length less than or equal to k:

[0241] Where k is a positive integer, for example, k=6, and e is the calling edge.

[0242] Step B2: Extract edge-level quantile delay.

[0243] For each call edge e in the topology path, based on its link attributes, the delay quantiles at a specified confidence level are extracted from the pre-constructed delay distribution and used as the edge-level quantile delay for that call edge. For example, the delay distribution is... Confidence level is Then the edge-level quantile delay of the called edge is .

[0244] in, The probability of meeting the latency target can be adjusted as needed, for example... . For quantile calculation, for delayed distribution After sorting by probability, select the one with the highest cumulative probability. The latency value. This method, compared to a fixed latency threshold, can adapt to the latency fluctuation characteristics of the link.

[0245] In this embodiment, the pre-constructed latency distribution consists of three types of baseline latency distributions used by the cross-domain jitter hybrid core, including the baseline latency distribution for cross-type links, the baseline latency distribution for cloud-type links, and... The baseline delay distribution for the type of link. For example, the baseline delay distribution for a cross-type link. If 95% of the latency is ≤100ms, then the edge-level latency is 100ms. This means that there is a 95% chance that the propagation latency of the cross-domain call edge e will not exceed 100ms.

[0246] It should be noted that the link attribute of the calling edge is... .

[0247] Step B3: Accumulate the edge-level quantile delays to obtain the quantile propagation delays.

[0248] Since the propagation delays of each calling edge e within the same path p in the topological path are independent, the quantile propagation delay of path p is the sum of the edge-level quantile delays, as shown in formula (16):

[0249] (16)

[0250] in, Let P be the propagation delay of the path P.

[0251] It's understandable that the more network domains a path traverses, the more the long-tail characteristic accumulates. The long-tail characteristic refers to the fact that in cloud, edge, and cross-domain links, some requests experience significantly higher latency than the average link latency due to factors such as network congestion, route jumps, and fluctuations in edge node resources. Although the proportion of these "slow requests" is low, it is not negligible, resulting in a "tail" shape in the high-latency range of the latency distribution curve. Therefore, the path percentile latency provided in this application's embodiment more accurately reflects the propagation characteristics of cross-domain paths, which can improve the accuracy of causal reachability determination.

[0252] When an upstream node u triggers an anomaly, the latest time from which the anomaly affects the downstream anomalous node v is called the optical cone front time from upstream node u to downstream anomalous node v. For example, if path p is from upstream node u to downstream anomalous node v, and the starting point of the anomaly window is time t0, the optical cone front time along path p is:

[0253] Let t0 be the light cone leading edge time of path p. This time represents the latest time from when the upstream node u triggers the anomaly at time t0 to when it affects the downstream anomalous node v.

[0254] Furthermore, in this embodiment, each enumerated path can record its corresponding upstream node, path, path quantile propagation delay, and optical cone front time to obtain a candidate list. For example, the enumerated path is the path from upstream node u to downstream abnormal node v, and the corresponding quantile propagation delay is... The time of the leading edge of the light cone is The triplet is obtained as Add the triple to the candidate list.

[0255] It is understandable that the candidate list includes each enumerated path, as well as the upstream node, path, path quantile propagation delay, and optical cone front time corresponding to each enumerated path.

[0256] S6110 uses adaptive pruning with quantile thresholds to obtain the candidate domain of the light cone.

[0257] For the upstream paths in the topology, candidate optical cone regions are obtained by eliminating temporally impossible paths, merging redundant paths, and removing extreme noise paths. The specific acquisition method includes the following steps:

[0258] Step C1: Remove upstream nodes corresponding to impossible time-series paths.

[0259] Specifically, if the optical cone leading edge time of a certain path is later than the termination time t1 of the abnormal window, for example, the optical cone leading edge time of path P... The path is a temporally impossible path, so the path P is removed from the candidate list.

[0260] Step C2: Merge the shortest paths.

[0261] Specifically, for the same upstream node u, there may be multiple paths to reach the downstream abnormal node v. In this application, only the path corresponding to the shortest quantile delay is retained in the embodiment, denoted as:

[0262]

[0263] This represents the shortest quantile delay corresponding to the upstream node u.

[0264] In this way, each candidate node retains only one representative path. This approach can avoid multiple paths from the same upstream node participating in subsequent calculations, reduce redundancy, and also help reduce the complexity of constructing the subsequent kernel dictionary.

[0265] Step C3: Remove extreme trailing nodes.

[0266] An extreme tail node refers to an upstream node whose quantile propagation delay to the downstream abnormal node v significantly deviates from the normal delay distribution of most paths, and belongs to a low-probability propagation delay anomaly. In the embodiments of this application, the quantile propagation delay from the upstream node u to the downstream abnormal node v is greater than or equal to the target threshold, and the upstream node u is an extreme tail node.

[0267] The target threshold and the average quantile propagation delay of other candidate nodes after excluding upstream node u and standard deviation There is a positive correlation, for example, the target threshold is If the propagation delay of the upstream node u is... satisfy:

[0268]

[0269] If the upstream node is considered an extreme trailing node, it will be removed from the candidate list. Removing extreme trailing nodes from the candidate list can prevent them from affecting subsequent root cause localization and further improve localization accuracy.

[0270] Step C4: Obtain the candidate region of the light cone.

[0271] The candidate list after removing upstream nodes corresponding to impossible temporal paths and extreme trailing nodes constitutes the candidate region of the light cone. .

[0272] It is understandable that this optical cone candidate domain reduces the solution complexity of candidate root cause anomaly localization while avoiding the missed detection of potential root causes that are time-reachable.

[0273] S520: Construct a convolution dictionary and solve for the target activation vector using sparse inversion.

[0274] The target activation vector includes only the upstream nodes essential for reconstructing the residual signals of downstream anomalous nodes, along with the weakest anomalous activation intensity of each upstream node. It can be understood that the target activation vector minimizes the number of upstream nodes, and by utilizing the weakest anomalous activation intensity of each upstream node, the anomalous image of a single upstream node can be avoided. Using the target activation vector, the residual signals of downstream anomalous nodes can be reconstructed to the greatest extent possible through propagation kernel convolution stacking.

[0275] This application provides a method for obtaining a target activation vector, including the following steps:

[0276] Step D1: Align the propagation kernel.

[0277] Cross-domain jitter hybrid kernel of any upstream node u This describes the change in the intensity of the unit anomalous excitation at node u over time τ. As a shock response model without a time reference, it is unclear how long it takes for the excitation to reach the downstream anomalous node v after it originates from the upstream node u.

[0278] The path propagation delay from upstream node u to downstream anomalous node v in the candidate region of the optical cone is less than or equal to the span of the anomalous window. Without alignment, the cross-domain jittering hybrid kernel... The peak value may occur earlier than the path propagation delay from upstream node u to downstream abnormal node v, or later than the path propagation delay from upstream node u to downstream abnormal node v, i.e., exceeding the allowed time of the optical cone, which leads to misjudgment of the excitation time in subsequent sparse inversion and destroys causal consistency.

[0279] In view of this, the embodiments of this application first align cross-domain jitter hybrid cores. Time is used to avoid a mismatch between the propagation kernel time and the actual propagation delay. The specific alignment method is as follows:

[0280] Cross-domain jitter hybrid kernel for upstream node u in the candidate domain of the optical cone The overall propagation delay of its shortest path is shifted to the right. 1 sampling point, to obtain the alignment propagation kernel As shown in formula (17):

[0281] (17)

[0282] This is equivalent to delaying the abnormal excitation effect of the upstream node u. Then it takes effect again. For example, Cross-domain jitter hybrid core Peak at =5ms, and the peak appears at 15ms after alignment. That is, the upstream node u triggers the anomaly at t=0 and is most affected at t=15ms, which is consistent with the definition of light cone.

[0283] Step D2: Obtain multiple shift cores triggered by each upstream node at different times within the abnormal window.

[0284] This application embodiment solves the problem that the propagation kernel cannot cover multiple time-lapse stimuli by expanding the alignment propagation kernel of a single upstream node u into multiple shift kernels triggered at different times within the abnormal window.

[0285] If the alignment propagation kernel of upstream node u The core length is The abnormal window is By sliding along the abnormal window, you can obtain... The dictionary atoms are shown in formulas (18) to (19):

[0286] (18)

[0287] (19)

[0288] in, For the sampling interval, each dictionary atom represents the upstream node u at time. The timing response is triggered only once. This method ensures that the effective effect of each shift core falls entirely within the anomaly window and does not exceed the window boundaries.

[0289] Step D3: Construct the convolution dictionary.

[0290] To obtain the upstream node u, at time... When an anomaly is triggered, the superimposed responses can best fit the residual of the downstream anomaly node v, thus matching the upstream node u at time t. shift nuclei Flatten into column vectors This transforms the time-series propagation problem into a linear algebra problem that can be solved efficiently.

[0291] Stack the column vectors of the same upstream node u in ascending time order, and concatenate them vertically between different upstream nodes to obtain the convolution dictionary D as follows:

[0292]

[0293] in, As an observation dimension, It is understandable that each column in the convolutional dictionary D... Represents node u at time t. The theoretical response of a unit shock to the residual at the downstream anomalous node v.

[0294] Furthermore, the convolution dictionary D and its columns can be indexed to... The mapping is then used. Then, the convolution dictionary D and the causal residual are used... Forming linear equations:

[0295]

[0296] The excitation coefficient is used to locate the smallest and most reliable root cause node and trigger time of the anomaly. In the embodiments of this application, it can be obtained through sparse regularization optimization. The value of .

[0297] In the embodiments of this application, the target residual refers to the standardized residual sequence of downstream abnormal nodes within the abnormal window, which is the abnormal signal after normal fluctuations have been removed.

[0298] Step D4: Obtain the target incentive vector based on the sparse incentive optimization model.

[0299] The following uses the abnormal window as an example. The sampling interval for the indicator detection values ​​is [missing information]. The residual effect is , Let's take an example to illustrate. From the linear equation composed of the convolution dictionary D and the effect residual r, we obtain the sparsest and non-negative activation coefficients. This minimizes the convolution reconstruction error. The specific method for obtaining this is as follows:

[0300] The target activation vector includes only the upstream nodes essential for reconstructing the residual signals of downstream anomalous nodes, along with the weakest anomalous activation intensity of each upstream node. It can be understood that the target activation vector minimizes the number of upstream nodes, and by utilizing the weakest anomalous activation intensity of each upstream node, the anomalous image of a single upstream node can be avoided. The target activation vector can be superimposed with a propagation kernel convolution to maximize the reconstruction of the residual effect of downstream anomalous nodes. In this embodiment, each element of the target activation vector is greater than or equal to 0, and the vector strength value of the residual vector satisfying the condition that the reconstructed residual and the effect residual are minimized; the reconstructed residual is the product of the target activation vector and the convolution dictionary.

[0301] In one example, the upstream nodes are multiple upstream nodes within the optical cone candidate domain of the downstream anomalous nodes. Nodes within the optical cone candidate domain are upstream nodes whose path propagation delay to the downstream anomalous node does not exceed the span of the anomalous window. The root cause localization system can utilize a pre-constructed loss function, combined with the effect residual and convolution dictionary, to determine the target activation vector. The target activation vector is a global activation vector, including sub-activation vectors of multiple upstream nodes. The loss function includes a reconstruction error term, an L1 coefficient regularization term, and a node-level group sparse regularization term. The reconstruction error term is the squared L2 norm of the difference vector between the effect residual and the residual to be reconstructed. The residual to be reconstructed is the product of the global activation vector to be solved and the convolution dictionary. The L1 sparse regularization term is the product of the L1 norm of the global activation vector to be solved and the first weight. The node-level group sparse regularization term is the product of the sum of the L2 norms of the sub-activation vectors of multiple upstream nodes and the second weight.

[0302] For example, if the effect residual is The pre-defined functions constructed are as follows:

[0303]

[0304] Among them, the reconstruction error term Used to make the selected excitation superimposed fit as closely as possible to the observed residual. Sparse regularization At the column (dictionary atom) level, coefficient sparsity is encouraged, prompting only a small number of dictionary atoms to participate in the reconstruction. First weight. Node group sparsity regularization. :in Treat all excitation coefficients of the same upstream node u as a group; if the entire group To reduce to zero means to remove the node entirely – further compressing the root cause set at the node level. This is the second weight. Non-negativity constraint. This indicates that abnormal excitation can only be amplified in the positive direction, and there is no reverse cancellation.

[0305] For each upstream node u, based on the target activation vector Obtain the overall incentive intensity and the moment when the strongest incentive occurs. The specific method for obtaining this information is as follows:

[0306]

[0307]

[0308] in, The overall excitation intensity representing node u can be considered as a confidence score; This represents the moment when the strongest incentive occurs at that node.

[0309] Sort the overall incentive intensity according to its intensity, for example, from low to high or from high to low, to obtain a root cause priority list. For example, the root cause priority list sorted from high to low is as follows:

[0310]

[0311] Based on the root cause priority list, the ranking of root cause nodes and their corresponding incentive times and intensities can be obtained, thus yielding the target incentive vector. Therefore, the root cause localization system completes the entire chain of inference, from the abnormal residual to which node and when the current fault is most likely to be triggered.

[0312] S530 verifies the target excitation vector through replay verification.

[0313] To ensure that the sparse inversion results have sufficient explanatory power for the effect residuals, embodiments of this application also include an exception window. Three quantifiable metrics are calculated: reconstruction error, peak alignment, and energy coverage. Based on these metrics, the sparse inversion results are assessed to determine whether the interpretation is adequate.

[0314] In practical implementation, the reconstruction error, peak alignment, and energy coverage corresponding to the target excitation vector can be determined. If the reconstruction error is greater than or equal to the first threshold, the peak alignment is greater than or equal to the second threshold, and the energy coverage is greater than or equal to the preset third threshold, the target excitation vector is considered to have passed verification. If any indicator fails to meet the standard, the restrictions can be relaxed step by step, for example, by expanding the light cone quantile, lowering the regularization weight, and re-executing the sparse excitation optimization model to obtain the target excitation vector until it passes or reaches the maximum iteration. The embodiments of this application can use comprehensive indicators such as reconstruction error, peak alignment, and energy coverage for evaluation, which can maintain lower error, higher peak synchronization, and more complete energy restoration when reconstructing the residual waveform of the target node. This makes the characterization of abnormal patterns more realistic and the root cause localization results more accurate.

[0315] S540 outputs root cause ranking and confidence level.

[0316] For optical cone candidate domain any upstream node in to comprehensively incentivize its intensity Linear scaling to the 0–1 interval yields the confidence score.

[0317]

[0318] If there is only one node, then set =1. This represents the minimum combined excitation intensity of multiple upstream nodes in the candidate domain of the optical cone. It represents the maximum value of the combined excitation intensity of multiple upstream nodes in the candidate domain of the light cone.

[0319] according to Sort the root causes from highest to lowest confidence score, and output the root cause nodes corresponding to the top m confidence scores, along with their rankings and the confidence scores for each node. m is a positive integer, for example, m=5. Therefore, based on the output root cause rankings and confidence scores, the root cause of the anomaly can be located.

[0320] In summary, the embodiments of this application can be evaluated using comprehensive indicators such as reconstruction error, peak alignment, and energy coverage. When reconstructing the residual waveform of the target node, it can maintain low error, high peak synchronization, and relatively complete energy restoration, resulting in more realistic characterization of abnormal patterns and more accurate root cause localization. The cross-domain jitter hybrid core integrates the multimodal latency distribution of the three links—intra-cloud, intra-edge, and cloud-edge interconnection—into the propagation model, avoiding long-tail loss caused by a single Gaussian assumption. This design effectively covers complex network jitter scenarios, reduces the risk of mis-pruning or miscalculation of long-tail paths, and improves robustness in cloud-edge collaborative environments. The entire inference link relies only on FFT and linear algebra operations, without relying on GPUs or deep learning inference frameworks. In typical deployment environments, a complete root cause analysis can be completed within seconds or even sub-seconds, meeting the timeliness requirements of online operation and maintenance while significantly reducing hardware and operation and maintenance investment.

[0321] Regarding the appendix Figure 1 ~Appendix Figure 6 The present application provides an abnormal root cause determination device in addition to the method for determining abnormal root causes shown.

[0322] Appendix Figure 7 An anomaly root cause determination device provided in this application embodiment is applied to a microservice system. The microservice system has a cloud-edge collaborative hybrid architecture, which includes a cloud data center and edge nodes. The device 700 includes:

[0323] The acquisition unit 701 is used to acquire the latency impact kernel of the call path based on the cause residual of the upstream node at the target time and the effect residual of the downstream abnormal node at the target time.

[0324] The call path is the call path from the upstream node to the downstream abnormal node, and the call path includes a first link, a second link, and a third link; the first link is a link within the cloud data center, the second link is a link between the edge nodes, and the third link is a cross-domain link between the cloud data center and the edge nodes; the cause residual is the difference between the indicator detection value of the upstream node at the target time and the expected value of the first indicator; the effect residual is the difference between the indicator detection value of the downstream abnormal node at the target time and the expected value of the second indicator, and the abnormal window of the upstream node and the abnormal window of the downstream abnormal node are the same; the latency impact kernel is a response function that quantifies the propagation of the upstream node's abnormality over time and its impact on the downstream abnormal node;

[0325] The convolution processing unit 702 is used to perform convolution processing on the delay influence kernel using the delay distribution model of the first link to obtain a first domain kernel; to perform convolution processing on the delay influence kernel using the delay distribution model of the second link to obtain a second domain kernel; and to perform convolution processing on the delay influence kernel using the delay distribution model of the third link to obtain a third domain kernel.

[0326] The hybrid core determination unit 703 is used to determine the cross-domain jitter hybrid core of the call path based on the first domain kernel, the second domain kernel, and the third domain kernel;

[0327] The root cause determination unit 704 is used to determine the root cause of the abnormality of the downstream abnormal node based on the cross-domain jitter hybrid kernel.

[0328] Optionally, obtaining the latency impact kernel of the call path based on the cause residual of the upstream node at the target time and the effect residual of the downstream abnormal node at the target time includes:

[0329] Obtain the cross-correlation curves between the lag residuals and the causal residuals; the lag residuals are the effects residuals lagged by the lag. The residual obtained afterwards, the ;

[0330] Wherein, the dependent variable of the cross-correlation curve is the cross-correlation value between the lagged residual and the causal residual, and the independent variable of the cross-correlation curve is... ;

[0331] Obtain the maximum cross-correlation value and the corresponding baseline lag time from the cross-correlation curve;

[0332] After the baseline lag time, determine the target lag time corresponding to the first decrease of the cross-correlation value in the cross-correlation curve to the target correlation value;

[0333] The curve representing the target interval is extracted from the cross-correlation curve and used as the latency impact kernel of the calling path; the target interval is from 0 to the target lag time.

[0334] Optionally, if the latency impact kernel is a preliminary latency impact kernel, the convolution processing unit 702 is further configured to: process the preliminary latency impact kernel of the call path using regularized deconvolution to obtain a refined latency impact kernel;

[0335] The process of convolving the latency impact kernel with the latency distribution model of the first link to obtain a first domain kernel; convolving the latency impact kernel with the latency distribution model of the second link to obtain a second domain kernel; and convolving the latency impact kernel with the latency distribution model of the third link to obtain a third domain kernel includes:

[0336] The refined latency impact kernel is convolved using the latency distribution model of the first link to obtain the first domain kernel; the refined latency impact kernel is convolved using the latency distribution model of the second link to obtain the second domain kernel; and the refined latency impact kernel is convolved using the latency distribution model of the third link to obtain the third domain kernel.

[0337] Optionally, determining the cross-domain jitter hybrid core of the call path based on the first domain kernel, the second domain kernel, and the third domain kernel includes:

[0338] In the call path, a first time ratio of the first link in the complete call of the call path, a second time ratio of the second link in the complete call of the call path, and a third time ratio of the third link in the complete call of the call path are determined;

[0339] The sum of the first product, the second product, and the third product is taken as the cross-domain jitter hybrid kernel; wherein, the first product is the product of the first time ratio and the first domain kernel, the second product is the product of the second time ratio and the second domain kernel, and the third product is the product of the third time ratio and the third domain kernel.

[0340] Optionally, the methods for obtaining the delay distribution model of the first link, the delay distribution model of the second link, and the delay distribution model of the third link include:

[0341] From the distributed tracing logs, obtain the latency samples of the first link, the second link, and the third link;

[0342] Determine whether the number of delay samples of the first link is greater than or equal to a preset threshold. If yes, fit the delay samples of the first link with a Gaussian mixture model to obtain the delay distribution model of the first link. If no, fit the delay samples of the first link with a unimodal Gamma distribution to obtain the delay distribution model of the first link.

[0343] Determine whether the number of delay samples of the second link is greater than or equal to a preset threshold. If yes, fit the delay samples of the second link with a Gaussian mixture model to obtain the delay distribution model of the second link. If no, fit the delay samples of the second link with a unimodal Gamma distribution to obtain the delay distribution model of the second link.

[0344] Determine whether the number of delay samples of the third link is greater than or equal to a preset threshold. If yes, perform Gaussian mixture model fitting on the delay samples of the third link to obtain the delay distribution model of the third link. If no, use a unimodal Gamma distribution to fit the delay samples of the third link to obtain the delay distribution model of the third link.

[0345] Optionally, determining the cross-domain jitter hybrid core of the call path based on the first domain kernel, the second domain kernel, and the third domain kernel includes:

[0346] Based on the first domain kernel, the second domain kernel, and the third domain kernel, obtain the initial hybrid kernel for the call path;

[0347] Obtain the cumulative energy of the initial hybrid core; when the cumulative energy reaches a preset confidence level, set the cross-correlation value corresponding to the time after the current lag time in the initial hybrid core to zero, and obtain the cross-domain jitter hybrid core of the call path.

[0348] Optionally, the method for obtaining the causal residual and the effect residual includes:

[0349] Obtain the first indicator detection value of the upstream node at the target time and the second indicator detection value of the downstream abnormal node at the target time;

[0350] Based on the health baseline model of the upstream node, the expected value of the first indicator of the upstream node at the target time is obtained; based on the health baseline model of the downstream abnormal node, the expected value of the second indicator of the downstream abnormal node at the target time is obtained; the health baseline model of the node is used to describe the expected value of the node's indicator under a fault-free state.

[0351] Based on the expected value of the first indicator and the detected value of the first indicator, obtain the cause residual of the upstream node; based on the expected value of the second indicator and the detected value of the second indicator, obtain the effect residual of the downstream abnormal node.

[0352] Optionally, the method for obtaining the health baseline model of a node includes:

[0353] Obtain the health interval of the node; the health interval is a range of data whose indicator values ​​are less than or equal to the preset service level target SLO upper limit, and the health interval excludes the indicator values ​​corresponding to the fault window;

[0354] If the values ​​in the health interval are periodic and continuous, an autoregressive model is used to fit the values ​​in the health interval to obtain the health baseline model for that node.

[0355] If the base number of the values ​​in the health interval is less than or equal to a preset base number threshold and is a discrete value, Kalman filtering is used to fit the values ​​in the health interval to obtain the health baseline model of the node.

[0356] Optionally, methods for determining abnormal nodes include:

[0357] Determine the absolute value of the actual residual for each node in the microservice system;

[0358] Nodes whose actual residual absolute value is greater than or equal to a preset residual threshold are designated as instantaneous abnormal nodes.

[0359] Within the anomaly window, determine whether the absolute value of the actual residual of the instantaneous anomaly node continuously exceeds the preset residual threshold, or whether the average value of the absolute value of the actual residual of the instantaneous anomaly node exceeds the preset residual threshold;

[0360] If the absolute value of the actual residual of the instantaneous abnormal node continuously exceeds the preset residual threshold, or the average value exceeds the preset residual threshold, the instantaneous abnormal node is determined to be the abnormal node.

[0361] In summary, the anomaly root cause determination device provided in this application can obtain the latency impact kernel of the call path based on the cause residual of the upstream node at the target time and the effect residual of the downstream abnormal node at the target time. The call path is the call path from the upstream node to the downstream abnormal node. In this application embodiment, the call path includes links within the cloud data center (i.e., the first link), links between edge nodes (i.e., the second link), and cross-domain links between the cloud data center and edge nodes (i.e., the third link). The anomaly windows of the upstream node and the downstream abnormal node are the same. The latency impact kernel is convolved using the latency distribution model of the first link to obtain the first domain kernel; the latency impact kernel is convolved using the latency distribution model of the second link to obtain the second domain kernel; and the latency impact kernel is convolved using the latency distribution model of the third link to obtain the third domain kernel. Based on the first domain kernel, the second domain kernel, and the third domain kernel, the cross-domain jitter hybrid kernel of the call path is determined. This cross-domain jitter hybrid kernel can accurately characterize the amplification effect of anomalies caused by cross-domain long-tail latency. Therefore, the root cause of anomalies determined based on this cross-domain jitter hybrid kernel helps to improve the accuracy of root cause localization.

[0362] According to the method provided in the embodiments of this application, this application also provides a chip system, which includes one or more processors for calling and executing instructions stored in memory, thereby causing the method described in the embodiments of this application to be executed. The chip system may be composed of chips or may include chips and other discrete devices.

[0363] The chip system may include input circuits or interfaces for transmitting information or data, and output circuits or interfaces for receiving information or data.

[0364] According to the method provided in the embodiments of this application, this application also provides a computer program product, which includes: computer program code, which, when run on a computer, causes the computer to execute the various steps or processes executed by the network device or terminal device in any of the foregoing method embodiments.

[0365] According to the method provided in the embodiments of this application, this application also provides a computer-readable storage medium storing program code, which, when run on a computer, causes the computer to execute the various steps or processes executed by the network device or terminal device in any of the foregoing method embodiments.

[0366] The computer-readable storage medium may be the aforementioned volatile memory or non-volatile memory, or it may include both volatile memory and non-volatile memory.

[0367] In the embodiments of this application, the terms and English abbreviations are exemplary examples given for ease of description and should not be construed as limiting the application in any way. This application does not preclude the possibility of defining other terms that can achieve the same or similar functions in existing or future agreements.

[0368] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions. When these computer instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated.

[0369] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

Claims

1. A method for determining the root cause of anomalies, characterized in that, Applied to a microservice system, wherein the microservice system has a cloud-edge collaborative hybrid architecture, the cloud-edge collaborative hybrid architecture includes a cloud data center and edge nodes, and the method includes: Based on the cause residual of the upstream node at the target time and the effect residual of the downstream abnormal node at the target time, obtain the latency impact kernel of the call path; The call path is the call path from the upstream node to the downstream abnormal node, and the call path includes a first link, a second link, and a third link; the first link is a link within the cloud data center, the second link is a link between the edge nodes, and the third link is a cross-domain link between the cloud data center and the edge nodes; the cause residual is the difference between the indicator detection value of the upstream node at the target time and the expected value of the first indicator; the effect residual is the difference between the indicator detection value of the downstream abnormal node at the target time and the expected value of the second indicator, and the abnormal window of the upstream node and the abnormal window of the downstream abnormal node are the same; the latency impact kernel is a response function that quantifies the propagation of the upstream node's abnormality over time and its impact on the downstream abnormal node; The latency impact kernel is convolved using the latency distribution model of the first link to obtain a first domain kernel; the latency impact kernel is convolved using the latency distribution model of the second link to obtain a second domain kernel; and the latency impact kernel is convolved using the latency distribution model of the third link to obtain a third domain kernel. Based on the first domain kernel, the second domain kernel, and the third domain kernel, determine the cross-domain jitter hybrid kernel of the call path; Based on the cross-domain jitter hybrid kernel, the root cause of the abnormality of the downstream abnormal node is determined.

2. The method according to claim 1, characterized in that, The step of obtaining the latency impact kernel of the call path based on the cause residual of the upstream node at the target time and the effect residual of the downstream abnormal node at the target time includes: Obtain the cross-correlation curves between the lag residuals and the causal residuals; the lag residuals are the effects residuals lagged by the lag. The residual obtained afterwards, the ; Wherein, the dependent variable of the cross-correlation curve is the cross-correlation value between the lagged residual and the causal residual, and the independent variable of the cross-correlation curve is... ; Obtain the maximum cross-correlation value and the corresponding baseline lag time from the cross-correlation curve; After the baseline lag time, determine the target lag time corresponding to the first decrease of the cross-correlation value in the cross-correlation curve to the target correlation value; The curve representing the target interval is extracted from the cross-correlation curve and used as the latency impact kernel of the calling path; the target interval is from 0 to the target lag time.

3. The method according to claim 1 or 2, characterized in that, If the delay impact kernel is a preliminary delay impact kernel, the method further includes: The initial latency impact kernel of the call path is processed using regularized deconvolution to obtain a refined latency impact kernel; The process of convolving the latency impact kernel with the latency distribution model of the first link to obtain a first domain kernel; convolving the latency impact kernel with the latency distribution model of the second link to obtain a second domain kernel; and convolving the latency impact kernel with the latency distribution model of the third link to obtain a third domain kernel includes: The refined latency impact kernel is convolved using the latency distribution model of the first link to obtain the first domain kernel; the refined latency impact kernel is convolved using the latency distribution model of the second link to obtain the second domain kernel; and the refined latency impact kernel is convolved using the latency distribution model of the third link to obtain the third domain kernel.

4. The method according to claim 1, characterized in that, The step of determining the cross-domain jitter hybrid core of the call path based on the first domain kernel, the second domain kernel, and the third domain kernel includes: In the call path, a first time ratio of the first link in the complete call of the call path, a second time ratio of the second link in the complete call of the call path, and a third time ratio of the third link in the complete call of the call path are determined; The sum of the first product, the second product, and the third product is taken as the cross-domain jitter hybrid kernel; wherein, the first product is the product of the first time ratio and the first domain kernel, the second product is the product of the second time ratio and the second domain kernel, and the third product is the product of the third time ratio and the third domain kernel.

5. The method according to claim 1, characterized in that, The methods for obtaining the delay distribution model of the first link, the delay distribution model of the second link, and the delay distribution model of the third link include: From the distributed tracing logs, obtain the latency samples of the first link, the second link, and the third link; Determine whether the number of delay samples of the first link is greater than or equal to a preset threshold. If yes, fit the delay samples of the first link with a Gaussian mixture model to obtain the delay distribution model of the first link. If no, fit the delay samples of the first link with a unimodal Gamma distribution to obtain the delay distribution model of the first link. Determine whether the number of delay samples of the second link is greater than or equal to a preset threshold. If yes, fit the delay samples of the second link with a Gaussian mixture model to obtain the delay distribution model of the second link. If no, fit the delay samples of the second link with a unimodal Gamma distribution to obtain the delay distribution model of the second link. Determine whether the number of delay samples of the third link is greater than or equal to a preset threshold. If yes, perform Gaussian mixture model fitting on the delay samples of the third link to obtain the delay distribution model of the third link. If no, use a unimodal Gamma distribution to fit the delay samples of the third link to obtain the delay distribution model of the third link.

6. The method according to claim 1, characterized in that, The step of determining the cross-domain jitter hybrid core of the call path based on the first domain kernel, the second domain kernel, and the third domain kernel includes: Based on the first domain kernel, the second domain kernel, and the third domain kernel, obtain the initial hybrid kernel for the call path; Obtain the cumulative energy of the initial hybrid core; when the cumulative energy reaches a preset confidence level, set the cross-correlation value corresponding to the time after the current lag time in the initial hybrid core to zero, and obtain the cross-domain jitter hybrid core of the call path.

7. The method according to claim 1, characterized in that, The methods for obtaining the causal residual and the effect residual include: Obtain the first indicator detection value of the upstream node at the target time and the second indicator detection value of the downstream abnormal node at the target time; Based on the health baseline model of the upstream node, the expected value of the first indicator of the upstream node at the target time is obtained; based on the health baseline model of the downstream abnormal node, the expected value of the second indicator of the downstream abnormal node at the target time is obtained; the health baseline model of the node is used to describe the expected value of the node's indicator under a fault-free state. Based on the expected value of the first indicator and the detected value of the first indicator, obtain the cause residual of the upstream node; based on the expected value of the second indicator and the detected value of the second indicator, obtain the effect residual of the downstream abnormal node.

8. The method according to claim 7, characterized in that, Methods for obtaining the health baseline model of a node include: Obtain the health interval of the node; the health interval is a range of data whose indicator values ​​are less than or equal to the preset service level target SLO upper limit, and the health interval excludes the indicator values ​​corresponding to the fault window; If the values ​​in the health interval are periodic and continuous, an autoregressive model is used to fit the values ​​in the health interval to obtain the health baseline model for that node. If the base number of the values ​​in the health interval is less than or equal to a preset base number threshold and is a discrete value, Kalman filtering is used to fit the values ​​in the health interval to obtain the health baseline model of the node.

9. The method according to claim 1, characterized in that, Methods for identifying anomalous nodes include: Determine the absolute value of the actual residual for each node in the microservice system; Nodes whose actual residual absolute value is greater than or equal to a preset residual threshold are designated as instantaneous abnormal nodes. Within the anomaly window, determine whether the absolute value of the actual residual of the instantaneous anomaly node continuously exceeds the preset residual threshold, or whether the average value of the absolute value of the actual residual of the instantaneous anomaly node exceeds the preset residual threshold; If the absolute value of the actual residual of the instantaneous abnormal node continuously exceeds the preset residual threshold, or the average value exceeds the preset residual threshold, the instantaneous abnormal node is determined to be the abnormal node.

10. An apparatus for determining the root cause of an anomaly, characterized in that, Applied to a microservice system, wherein the microservice system has a cloud-edge collaborative hybrid architecture, the cloud-edge collaborative hybrid architecture includes a cloud data center and edge nodes, and the device includes: The acquisition unit is used to acquire the latency impact kernel of the call path based on the cause residual of the upstream node at the target time and the effect residual of the downstream abnormal node at the target time. The call path is the call path from the upstream node to the downstream abnormal node, and the call path includes a first link, a second link, and a third link; the first link is a link within the cloud data center, the second link is a link between the edge nodes, and the third link is a cross-domain link between the cloud data center and the edge nodes; the cause residual is the difference between the indicator detection value of the upstream node at the target time and the expected value of the first indicator; the effect residual is the difference between the indicator detection value of the downstream abnormal node at the target time and the expected value of the second indicator, and the abnormal window of the upstream node and the abnormal window of the downstream abnormal node are the same; the latency impact kernel is a response function that quantifies the propagation of the upstream node's abnormality over time and its impact on the downstream abnormal node; The convolution processing unit is configured to perform convolution processing on the latency impact kernel using the latency distribution model of the first link to obtain a first domain kernel; perform convolution processing on the latency impact kernel using the latency distribution model of the second link to obtain a second domain kernel; and perform convolution processing on the latency impact kernel using the latency distribution model of the third link to obtain a third domain kernel. A hybrid core determination unit is used to determine the cross-domain jitter hybrid core of the call path based on the first domain kernel, the second domain kernel, and the third domain kernel; The root cause determination unit is used to determine the root cause of the abnormality of the downstream abnormal node based on the cross-domain jitter hybrid kernel.