Root cause positioning method, device and equipment

By using multi-round iterative loops and streaming writes to perform real-time root cause localization of the target system, the problem of high localization latency in existing technologies is solved, and potential problems can be identified and responded to quickly, thus improving the efficiency of operation and maintenance management.

CN120973559APending Publication Date: 2025-11-18HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410887661.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-05-17
Filing Date
2024-07-02
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Existing root cause localization solutions require system fault reporting to trigger, resulting in high localization latency, affecting fault recovery time, and failing to identify potential problems.

Method used

The system employs a multi-round iterative approach to perform real-time detection of the target system's data during streaming. It utilizes an event detection algorithm to convert the data into an event sequence and actively locates the root cause using a root cause localization model. The model parameters are updated using a sliding time window to reduce reliance on full historical data.

Benefits of technology

It enables the identification of potential problems, shortens the latency of root cause localization, improves the convenience and flexibility of operation and maintenance management, and reduces the amount of data required.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120973559A_ABST
    Figure CN120973559A_ABST
Patent Text Reader

Abstract

The invention discloses a root cause positioning method, device and equipment, and relates to the field of computers, and the method comprises the following steps: carrying out the following operations on to-be-detected data of a target system in a multi-round loop iteration mode: obtaining the to-be-detected data of the round; a to-be-detected event sequence of the round is determined, the to-be-detected event sequence comprises a to-be-detected event sequence of the previous round and an event determined by the round, and the event determined by the round is generated based on an event detection algorithm and to-be-detected data of the round; performing root cause positioning on the to-be-detected event sequence of the current round, and determining a root cause positioning result of the current round; when a set iteration ending condition is met, outputting a root cause positioning result of the round; wherein the set iteration ending condition comprises that the root cause positioning result indicates that the target system is in an abnormal state, and the root cause positioning result can comprise one or more potential root causes which cause the target system to generate a to-be-detected event sequence of the round. According to the method, active streaming root cause positioning can be realized, so that potential problems are identified, and positioning time delay is shortened.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application claims priority to the Chinese Patent Application No. 202410619724.8, filed on May 17, 2024, and entitled "A Detection System", the content of which is incorporated herein by reference in its entirety. TECHNICAL FIELD

[0002] The present application relates to the technical field of computer, and particularly relates to a root cause positioning method, device and equipment. BACKGROUND

[0003] Root cause positioning technology is a technology for positioning the root cause of an exception when the target system has the exception. For example, in the field of cloud services, when an exception of high service delay occurs, the root cause positioning technology can be used to locate that an XX node of an XX instance has an out of memory (OOM). After the root cause is located, the operation and maintenance can take expansion operation to repair the problem.

[0004] Current root cause positioning schemes mostly adopt a mode of mining based on customer fault reporting, that is, a passive way of querying and analyzing business data and analyzing root causes step by step according to service business exceptions of the upper layer. For example, in an existing root cause positioning scheme, after a user reports a fault, full historical data is pulled from a database and sent into a Bayesian network for inference to obtain a root cause. Since this scheme needs system fault reporting to trigger root cause positioning, the positioning delay is high, which affects the fault recovery time, and potential problems cannot be identified. SUMMARY

[0005] The present application provides a root cause positioning method, device and equipment, which can identify potential problems and shorten the positioning delay.

[0006] In a first aspect, the present application provides a root cause positioning method, which can be used for root cause positioning of a target system. The method comprises the following steps.

[0007] For the target system of the stream write, the following operations are performed on the target system of the to-be-detected data by using a multi-round loop iteration manner: obtaining the to-be-detected data of the current round; updating the to-be-detected event sequence of the current round based on the to-be-detected data of the current round and the to-be-detected event sequence of the previous round, the to-be-detected event sequence including the to-be-detected event sequence of the previous round (which can be events determined in each historical round of the current loop iteration) and events determined in the current round, the events determined in the current round being generated based on an event detection algorithm of the target system and the to-be-detected data of the current round, the event detection algorithm being used to convert the to-be-detected data of the target system into events; after obtaining the to-be-detected event sequence of the current round, performing root cause positioning on the to-be-detected event sequence of the current round to determine a root cause positioning result of the current round, the root cause positioning result being used to indicate that the state of the target system is a normal state or an abnormal state.

[0008] When the set iteration end condition is reached, the root cause result of the last round is output; wherein the set iteration end condition includes that the root cause positioning result indicates that the target system is in an abnormal state; when the target system is indicated to be in an abnormal state, the root cause positioning result can include one or more potential root causes that cause the target system to generate the to-be-detected event sequence of the current round.

[0009] Through the above design, each time the to-be-detected data of the target system is written, a round of loop iteration is triggered for the to-be-detected data: the to-be-detected data is detected in real time using the event detection algorithm registered by the target system to convert the to-be-detected data into events. The to-be-detected event sequence of the current round is composed based on the event sequence of the previous round and the events converted from the to-be-detected data of the current round, i.e., the event sequence composed of events occurring in the target system in sequence. Then, the root cause positioning is performed on the to-be-detected event sequence of the current round to obtain the root cause positioning result. When the root cause positioning result indicates that the target system is in an abnormal state, the root cause positioning result also gives one or more potential root causes that cause the target system to be in an abnormal state. It can be seen that this method uses the stream write to-be-detected data to actively perform root cause positioning, which can reduce the data demand of root cause positioning, identify potential problems, and shorten the root cause positioning delay.

[0010] In a possible design, the event detection algorithm associated with the target system is determined based on a rule configured for the target system by a user, the rule including a specified event sequence and a root cause corresponding to the specified event sequence.

[0011] Through the above design, the convenience and flexibility of user operation and maintenance management of the target system are provided.

[0012] In one possible design, the method further includes: providing an application programming interface (API) for user configuration rules, the API including multiple fields, including an event sequence field and a root cause field; receiving API content submitted by the user, the API content including the multiple fields, event information input by the user for the event sequence field, and root cause information input by the user for the root cause field, wherein the event information is used to describe the multiple events included in the set event sequence and the order of the multiple events, and the root cause information is used to describe the root cause corresponding to the specified event sequence.

[0013] Through the above design, the API allows users to configure the rules of the target system, achieving convenience, flexibility, and compatibility with different software programs, thus enhancing its applicability.

[0014] In one possible design, the method further includes: providing a configuration interface for users to configure rules for the target system; and obtaining the user-configured rules from the configuration interface.

[0015] The above design provides users with a more convenient and intuitive way to configure the rules of the target system through a configuration interface.

[0016] In one possible design, root cause localization is performed on the current round of the event sequence to be detected, including: determining the root cause localization result based on the current round of the event sequence to be detected and the root cause localization model; the root cause localization model is used to predict the root cause based on the event sequence.

[0017] The root cause localization model includes the following parameters: state transition probability matrix and observation matrix;

[0018] The state transition probability matrix includes the transition probabilities between any two root causes associated with the target system, where root cause e i To the root cause e k The transition probability represents the current root cause e. i Then, at the next point in time, it will become the root cause e. k The probability of;

[0019] The observation matrix includes the transition probabilities between any root cause associated with the target system and any event associated with the target system, where the root cause e i To the event o m The transition probability represents the probability that the current root cause is known to be root cause e. i Then the current corresponding event is event o. m The probability of.

[0020] Through the above design, a root cause localization algorithm is constructed based on the state transition probability matrix and the observation matrix. In this way, the potential root cause can be determined based on the event sequence to be detected, without having to pull the full historical data from the database, and proactive root cause localization can be achieved.

[0021] In one possible design, the events associated with the target system include: events occurring in the target system recorded within a sliding time window of fixed length, and / or events contained in the rules configured for the target system; the root causes associated with the target system include root causes corresponding to the event sequences belonging to the target system recorded within the sliding time window, and / or root causes contained in the rules configured for the target system.

[0022] With the above design, when the data changes in characteristics over time, the events and root causes of the target system are recorded based on a sliding time window. Events and root causes from earlier times can be eliminated over time. In this way, the model parameters can be actively updated and learned as the data changes without the need for manual intervention to relearn.

[0023] In one possible design, the root cause localization result is the topk Viterbi path obtained by searching using the Viterbi decoding algorithm based on the sequence of events to be detected and the root cause localization model. The topk Viterbi path is used to indicate the top k root causes that generate the sequence of events to be detected, arranged from high to low probability, where k is a positive integer.

[0024] Secondly, this application also provides a root cause localization device, which has the function of implementing the behavior in the method example of the first aspect described above. The beneficial effects can be found in the description of the first aspect and any possible implementation of the first aspect, and will not be repeated here. The function can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above functions. In one possible design, the structure of the computing device includes a streaming detection module and a root cause analysis module; optionally, it also includes a rule registration module and a model update module.

[0025] The root cause localization device uses a multi-round iterative approach to perform the following operations on the data to be detected that is streamed into the target system:

[0026] The streaming detection module is used to acquire the data to be detected in the current round. The streaming detection module is also used to determine the sequence of events to be detected in the current round. The sequence of events to be detected in the current round includes the sequence of events to be detected in the previous round and the events determined in the current round. The events determined in the current round are generated based on the event detection algorithm associated with the target system and the data to be detected in the current round. The event detection algorithm is used to convert the data to be detected into events.

[0027] The root cause analysis module is used to locate the root cause of the event sequence to be detected in the current round and determine the root cause location result of the current round. The root cause analysis module is also used to output the root cause location result of the current round when the set iteration termination condition is met. The set iteration termination condition includes the root cause location result of the current round indicating that the target system is in an abnormal state. The root cause location result includes one or more potential root causes that caused the target system to generate the event sequence to be detected in the current round.

[0028] In one possible design, the event detection algorithm associated with the target system is determined based on rules submitted by the user and configured for the target system. The rules include specified event sequences and the root causes corresponding to the specified event sequences.

[0029] In one possible design, the rule registration module provides an application programming interface (API) for users to configure rules. The API includes multiple fields, including an event sequence field and a root cause field. It receives API content submitted by the user, which includes multiple fields, event information input for the event sequence field, and root cause information input for the root cause field. The event information describes the multiple events included in the set event sequence and their order, while the root cause information describes the root cause corresponding to the specified event sequence.

[0030] In one possible design, a rule registration module is used to provide a configuration interface for users to configure rules for the target system; and to retrieve user-configured rules from the configuration interface.

[0031] In one possible design, the root cause analysis module, when performing root cause localization on the current round of the event sequence to be detected, is specifically used to: determine the root cause localization result based on the current round of the event sequence to be detected and the root cause localization model; the root cause localization model is used to predict the root cause based on the event sequence.

[0032] In one possible design, the root cause localization model includes the following parameters: state transition probability matrix and observation matrix;

[0033] The state transition probability matrix includes the transition probabilities between any two root causes associated with the target system, where root cause e i To the root cause e k The transition probability represents the current root cause e. i Then, at the next point in time, it will become the root cause e. k The probability of;

[0034] The observation matrix includes the transition probabilities between any root cause associated with the target system and any event associated with the target system, where the root cause e i To the event o m The transition probability represents the probability that the current root cause is known to be root cause e. iThen the current corresponding event is event o. m The probability of.

[0035] In one possible design, the events associated with the target system include: events occurring in the target system recorded within a sliding time window of fixed length, and / or events contained in the rules configured for the target system; the root causes associated with the target system include root causes corresponding to the event sequences belonging to the target system recorded within the sliding time window, and / or root causes contained in the rules configured for the target system.

[0036] In one possible design, the root cause localization result is the topk Viterbi path obtained by searching using the Viterbi decoding algorithm based on the sequence of events to be detected and the root cause localization model. The topk Viterbi path is used to indicate the top k root causes that generate the sequence of events to be detected, arranged from high to low probability, where k is a positive integer.

[0037] Thirdly, this application also provides a computing device cluster, which includes at least one computing device. This at least one computing device has the functionality to implement the behavior described in the method example of the first aspect above. The beneficial effects are described in the first aspect and will not be repeated here. Each computing device includes a processor and a memory. The processor is configured to support the computing device in executing the method described in the first aspect or any possible design of the first aspect. The memory is coupled to the processor and stores the necessary program instructions and data of the computing device. The computing device also includes a communication interface for communicating with other devices, such as acquiring the data to be detected from a target system.

[0038] Fourthly, this application also provides a computer-readable storage medium storing instructions that, when executed on a computer, cause the computer to perform the method described in the first aspect or any possible design of the first aspect.

[0039] Fifthly, this application also provides a computer program product containing instructions that, when run on a computer, cause the computer to perform the method described in the first aspect or any possible design of the first aspect.

[0040] Sixthly, this application also provides a computer chip connected to a memory, the chip being used to read and execute a software program stored in the memory, and to execute the method described in the first aspect or any possible design of the first aspect.

[0041] For the beneficial effects of aspects two through six, please refer to the beneficial effects of aspect one, which will not be repeated here. Attached Figure Description

[0042] Figure 1This is a schematic diagram illustrating a possible application scenario provided by an embodiment of this application;

[0043] Figure 2 A schematic diagram of the architecture of a root cause localization device provided in an embodiment of this application;

[0044] Figure 3 A flowchart illustrating a rule registration method provided in an embodiment of this application;

[0045] Figure 4 A flowchart illustrating a root cause localization method provided in an embodiment of this application;

[0046] Figure 5 A schematic diagram of a flow cytometry detection process provided in an embodiment of this application;

[0047] Figure 6 This is a schematic diagram of a root cause localization process provided in an embodiment of this application;

[0048] Figure 7 This is a schematic diagram of the structure of a computing device provided in an embodiment of this application;

[0049] Figure 8 This is a schematic diagram of the structure of a computing device cluster provided in an embodiment of this application;

[0050] Figure 9 This is a schematic diagram of another computing device cluster provided in an embodiment of this application. Detailed Implementation

[0051] This application provides a root cause localization device (or root cause localization system) 20, which can be applied to various application scenarios, such as vehicle networking, cloud computing, cloud services, etc., and is not specifically limited thereto. The root cause localization device 20 is used to actively locate the root causes of a target system 10, identify potential root cause problems in the operation of the target system 10, and reduce the root cause localization latency.

[0052] The target system 10 can be a physical hardware device or a virtual device. Physical hardware devices include, but are not limited to, servers, terminal devices, vehicles, etc., while virtual devices include, but are not limited to, virtual machines, containers, etc. Terminal devices can be portable terminals, such as mobile phones, tablets, laptops, wearable devices (such as smartwatches), etc. Exemplary embodiments of the aforementioned terminal devices include, but are not limited to, portable terminal devices running iOS, Android, Microsoft, or other operating systems. It should also be understood that in some other embodiments of this application, the terminal devices described above may not be portable terminal devices, but rather desktop computers, in-vehicle terminals, smart screens, televisions, etc. In terms of quantity, the target system 10 may include one or more physical hardware devices, or one or more virtual devices, or at least one physical hardware device and at least one virtual device. For example, in one instance, the target system 10 is a cluster of multiple devices, such as a cloud, a computer cluster, or a vehicle cluster.

[0053] Figure 1 This is a schematic diagram illustrating an application scenario using the Internet of Vehicles (IoV) as an example. It should be understood that... Figure 1 The scenario shown is merely an example; the scenarios applicable to the embodiments of this application may vary. Figure 1 It may have more or fewer devices. For example, target system 10 may consist of only one vehicle. In another application scenario, root cause localization device 20 may acquire or receive operational data of target system 10 from a third-party database. This embodiment does not limit this.

[0054] The target system 10 generates operational data during operation, for example, in Figure 1 When the target system 10 is a vehicle, the operational data may include sensor-collected metrics such as speed and whether a collision has occurred. The root cause localization device 20 can proactively locate the root cause of any abnormality in the target system 10 based on this operational data. The root cause can be understood as the primary reason for the abnormal operation of the target system 10. For example, in a vehicle-to-everything (V2X) scenario, the root cause of a vehicle's abnormal operation might be a car accident. After localization, the backend maintenance team can promptly contact the vehicle owner to confirm the severity of the problem and provide timely assistance.

[0055] In this embodiment, the root cause localization device 20 uses the running data of the target system 10 written in a streaming manner to perform streaming root cause localization. Compared with the existing method, the root cause localization device 20 can actively initiate root cause localization, thereby shortening the root cause localization latency.

[0056] Figure 2 This is a schematic diagram of the structure of a root cause localization device 20 provided in an embodiment of this application.

[0057] like Figure 2As shown, the root cause localization device 20 includes a rule registration module 201, a streaming detection module 202, a root cause analysis module 203, a caching module 204, and a model update module 205.

[0058] The rule registration module 201 is used to register the rules that users are interested in into the root cause localization device 20.

[0059] The users here can be the operation and maintenance administrators or users of the target system 10, without specific limitations. Rules include event sequences and their corresponding root causes. An event sequence includes multiple events arranged in an ordered manner. In some cases, an event sequence may also include only one event, or in other words, a rule may include one event and its corresponding root cause. The root cause of an event sequence refers to the root cause that generated the event sequence.

[0060] The rule registration module 201 is specifically used to register user-focused rules into the root cause localization device 20. The registration operation includes two aspects: First, parsing the rules to obtain event information in the event sequence, generating multiple event detection algorithms based on the event information, and registering the event detection algorithms in the streaming detection module 202 for streaming detection. Second, registering the event sequence and root cause in the rules into the root cause analysis module 203.

[0061] The streaming detection module 202 is used to receive a data stream containing the operating data of the target system 10 and perform streaming processing on the operating data of the target system 10. The streaming processing may include: converting the input data stream into an event sequence based on an event detection algorithm, and inputting the event sequence to the root cause analysis module 203.

[0062] Root cause analysis module 203 is used to perform streaming root cause localization. Streaming root cause localization involves streaming the input event sequence into root causes.

[0063] The caching module 204 is used to cache the intermediate processes of streaming detection and / or root cause analysis. Specifically, it may include one or more of the following data: anomaly detection cache, current event sequence cache, historical event cache, and model parameter cache.

[0064] The system includes: an anomaly detection cache for storing intermediate detection values ​​from the streaming detection module 202; a current event sequence cache for storing event sequences to be detected (from the root cause analysis module 203); a historical event cache for storing historical event sequences of the target system 10; and a model parameter cache for storing parameters of the event detection algorithm and / or the root cause localization algorithm.

[0065] The model update module 205 is used to calculate and update the parameters of the event detection algorithm and / or the root cause localization algorithm using the features of the data cached in the cache module 204.

[0066] It should be noted that, Figure 2 The structure of the root cause localization device 20 shown is only an example. The root cause localization device 20 applicable to this embodiment may have a relative Figure 2 More or fewer modules, for example, the root cause localization device 20 in actual application may not include the cache module 204, or the model update module 205, etc., and this embodiment does not limit this.

[0067] The following is applied to Figure 2 Taking the root cause localization device 20 as an example, this embodiment introduces the streaming root cause localization method. The root cause localization method of this embodiment includes two process flows: first, a rule registration process (see...). Figure 3 (Flowchart shown). II. Flow Cytometry Root Cause Localization Method (see...) Figure 4 (The process is shown below). These will be described in detail below. For ease of understanding, the following explanation will use target system 10 as an example, representing a single vehicle.

[0068] Figure 3 This is a flowchart illustrating a rule registration method provided in this embodiment. Figure 3 As shown, the method may include the following steps:

[0069] Step 301: Obtain the rules submitted by the user.

[0070] In one implementation, the rule registration module 201 provides an application programming interface (API) for users to configure rules, which users can call to submit rules. Accordingly, the rule registration module 201 receives the API content submitted by the user, parses the API content, and completes the rule registration (see steps 302-304).

[0071] For example, the API includes multiple fields, which represent different information included in the rules. The API content submitted by the user includes these multiple fields and the content entered by the user for each field. For example, using... Figure 3 Taking the API content shown as an example, the API content includes:

[0072] (1) Field 1 and the content entered for Field 1.

[0073] Figure 3 In this context, field 1 is the incident stream. The input for field 1 includes event information for each event within the incident stream, such as... Figure 3 In this context, event information may include an event identifier and an event order, where the event identifier is used to uniquely identify an event. Figure 3The content within the double quotes represents the event identifier, and `order` indicates the order of the events. (Analysis) Figure 3 The content of field 1 shown shows the event sequence: "overspeed" → "collision" → "speed shift" → "car still".

[0074] (2) Field 2 and the content entered for Field 2.

[0075] Field 2 is the group by identifier, which indicates the group to which the rule belongs. Its function is to indicate which group the user configured the rule for. For example... Figure 3 The grouping in the code is "Vehicle ID (car id)". The Vehicle ID can be the VIN (Vehicle Identifier) ​​of a vehicle entered by the user. The VIN is used to uniquely identify a vehicle. This indicates that the rule submitted by the user through this API is configured for this vehicle. It should be understood that this grouping identifier is only an example, and the specific content included in the grouping identifier is not limited. For example, the grouping identifier can also include the VIN and the indicator identifier. The indicator identifier is used to uniquely identify an indicator, which will be introduced later and will not be elaborated here.

[0076] (3) Field 3 and the content entered for field 3.

[0077] Field 3 represents the root cause, where the root cause represented by field 3 corresponds to the event sequence represented by field 1. For example... Figure 3 The root cause of the event sequence in the text is "serious car accident". This root cause is the event sequence: "overspeed" → "collision" → "speed shift" → "car still".

[0078] It should be noted that the API content described above is merely an example. In actual applications, the number of fields included in the API, the content of each field, and the arrangement of multiple fields are not limited to this. Figure 3 This example is not limited to the above. It should also be noted that configuring rules via API is merely an example; this embodiment also supports other methods. For instance, in another implementation, the rule registration module 201 can provide a configuration interface for users to configure rules. Users configure rules through this interface, and the rule registration module 201 retrieves the user-configured rules from the interface and completes rule registration. This embodiment does not limit the method of user rule configuration; any method that allows for rule configuration is applicable to this embodiment.

[0079] Step 302: Parse the rules to obtain event information in the event sequence.

[0080] After obtaining the API content submitted by the user, the rule registration module 201 parses the API content to obtain the event information in the event sequence. As mentioned earlier, the event information includes the event identifier and the event order.

[0081] Step 303: Generate one or more event detection algorithms based on each event information, and register the one or more event detection algorithms in the streaming detection module 202.

[0082] Continue with Figure 3 Taking the API content shown as an example, after parsing the event information, the rule registration module 201 generates one or more event detection algorithms based on the event information. For example, simply put, "speeding" can be converted into a speeding detection algorithm: such as detecting whether the speed exceeds a high threshold (e.g., 100 mph). "Collision" can be converted into a sudden rise anomaly detection algorithm. Since vibration sensors report the amplitude of vibration, the value is low when no vibration occurs, and the amplitude increases after vibration occurs; therefore, a sudden rise is used to detect whether a collision has occurred. "Speed ​​drop" can be converted into a speed drop anomaly detection algorithm. "Stopping" can be converted into a stopping detection algorithm: such as detecting whether the speed is below a low threshold (e.g., 0.5 mph). In one possible implementation, the one or more event detection algorithms corresponding to an event can be preset; for example, binding user-configurable events with corresponding event detection algorithms and writing this binding relationship into a library file. After parsing the event information for each event in the event sequence, the corresponding event detection algorithm is obtained from the library file based on the event information, thus completing the conversion. Alternatively, the event detection algorithm corresponding to the event can be obtained from a third party. Or, the event detection algorithm can be generated through other methods, which are not specifically limited.

[0083] Next, the rule registration module 201 registers the converted event detection algorithm in the streaming detection module 202. For example, the rule registration module 201 can send the group identifier and event detection algorithm from the API content to the streaming detection module 202. The streaming detection module 202 saves the binding relationship between the group identifier and the event detection algorithm. Subsequently, the streaming detection module 202 can find the event detection algorithm based on the group identifier.

[0084] It should be understood that the event detection algorithm listed above is only a simple example, and this embodiment does not limit the algorithm content and implementation method of the event detection algorithm. For example, the event detection algorithm in this embodiment may also include a real-time updated algorithm model.

[0085] Step 304: Register the event sequence and root cause in the rule in the root cause analysis module 203.

[0086] In one embodiment, the root cause analysis module 203 includes a Hidden Markov Model (HMM). In this embodiment, the HMM includes a state transition probability matrix and an observation matrix, which will be described below and will not be repeated here. Accordingly, step 304, registering the event sequence and root cause in the rules to the root cause analysis module 203, includes updating the state transition probability matrix and the observation matrix based on the user-configured rules.

[0087] This completes one rule registration. It should be understood that a user can configure rules for the same target system more than 10 times. Each submitted rule may include one or more event sequences and root cause pairs. Each event sequence and root cause pair includes an event sequence (or event) and one or more root causes corresponding to that event sequence (or event).

[0088] It should be noted that there is no strict timing requirement between steps 304 and 303. Steps 304 and 303 can be executed simultaneously, or step 304 can be executed first and then step 303.

[0089] Subsequently, the flow cytometry detection module 202 and the root cause analysis module 203 complete the flow cytometry root cause localization for the target system 10 according to the rules registered for the target system 10, as follows: Figure 4 Let me introduce it.

[0090] Figure 4 This is a flowchart illustrating a flow cytometry root cause localization method provided in this embodiment. Figure 4 As shown, the method may include the following steps:

[0091] Step 401: The streaming detection module 202 receives the data to be detected from the target system 10.

[0092] In one implementation, the streaming detection module 202 receives the data to be detected from the target system 10, which can be real-time operating data or monitoring data of the target system. Monitoring data refers to the indicator data collected by monitoring the target system 10. Streaming can refer to the real-time operating data or monitoring data of the target system 10 being written to the streaming detection module 202 in the form of a data stream. The streaming detection module 202 performs streaming conversion on the received data stream.

[0093] The data to be detected in the target system 10 includes, but is not limited to, the following types of data:

[0094] (1) Time series data, such as indicator data sorted by acquisition time, such as velocity series data, including the velocity of the target system 10 at different time points within a certain period. Another example is vibration series data, which includes the vibration values ​​of the target system 10 at different times within a certain period, and the vibration values ​​can indicate whether the target system 10 has experienced a collision. Indicator data may also include CPU utilization, memory utilization, etc., without specific limitations.

[0095] (2) Log data.

[0096] Log data includes metrics data, events, errors, alarms, and other information collected during the runtime of the target system 10.

[0097] (3) Event data.

[0098] Each type of data to be detected contains metadata information, which is used to uniquely identify its characteristics. The streaming detection module 202 can query the event detection algorithm used to transform the data to be detected based on the metadata information.

[0099] In one implementation, the metadata information can be a group identifier, or in other words, a group identifier can be metadata information. For example, in a vehicle-to-everything (V2X) scenario, when the target system 10 is a vehicle, the metadata information can be a vehicle identification number (VIN) and / or an indicator identifier. As another example, in a cloud computing scenario, when the target system 10 is a cloud service system, the metadata information can include, but is not limited to, one or more of the following: region identifier, availability zone identifier, instance identifier, node identifier, indicator identifier, etc. For example, metadata information includes XXregion+XXAZ+XXinstance+XXnode+XXindicators. The streaming detection module 202 obtains the event detection algorithm bound to the metadata information based on the metadata information contained in the data to be detected.

[0100] Step 402: The streaming detection module 202 performs streaming processing on the received data to be detected to generate an event sequence.

[0101] Streaming processing includes: using the event detection algorithm registered in the target system 10 to sequentially convert the input data to be detected into events. Events occurring sequentially in the same target system 10 form an event sequence.

[0102] As mentioned above, the streaming detection module 202 can obtain the event detection algorithm bound to the metadata information contained in the data to be detected, thereby performing streaming processing.

[0103] The streaming processing steps performed by the streaming detection module 202 may differ depending on the type of data to be detected. For example, for time series data, the streaming detection module 202 can directly use an event detection algorithm to convert the time series data into an event sequence. However, for log data, the streaming detection module 202 needs to preprocess the log data (such as field extraction, format conversion, etc., the specifics of which are not limited) to obtain the target data to be detected, such as the indicator data contained in the log data, and then use an event detection algorithm to convert the target data to be detected into events.

[0104] The following section uses time series data as an example to illustrate how to perform streaming processing.

[0105] In one example, assuming the target system 10 is vehicle 130, the metadata information contained in the data to be detected for target system 10 includes the vehicle identification number (VIN) and speed index of vehicle 130; that is, the data to be detected includes the real-time speed sequence of vehicle 130. The streaming detection module 202 uses this metadata information to obtain the event detection algorithm corresponding to the speed index of vehicle 130. It should be noted that one index identifier may correspond to one or more event detection algorithms. The streaming detection module 202 uses all the event detection algorithms corresponding to the obtained speed index to detect each input speed value. If an event detection algorithm is satisfied, the speed value is converted into the corresponding event. In another example, in addition to the speed index, the data to be detected for target system 10 also includes a vibration index. Therefore, one or more event detection algorithms corresponding to the vibration index of vehicle 130 are obtained based on the VIN and the vibration index.

[0106] For example, see Figure 4As shown, assuming the data stream sequence input to the streaming detection module 202 includes speed value 1, vibration value 1, speed value 2, speed value 3, ..., for the first input speed value 1, the streaming detection module 202 iterates through all event detection algorithms corresponding to the speed index of the target system 10. Whichever event detection algorithm speed value 1 matches, it is converted into the corresponding event. For example, if speed value 1 matches the "speeding" event detection algorithm, then an "speeding" event is generated based on speed value 1. Next, the input vibration value 1 is detected. The streaming detection module 202 iterates through all event detection algorithms corresponding to the vibration index of the target system 10. Similarly, if vibration value 1 matches the "sudden rise anomaly detection" algorithm, then a "collision" event is generated based on vibration value 1. Assuming speed value 3 matches the "sudden drop anomaly detection" algorithm, a "speed drop" event is generated based on speed value 3. Similarly, a "stopping" event is generated based on speed value 4. And so on. It is worth noting that the order of events depends on the order of the corresponding data to be detected. In the example above, according to the data sorting, the corresponding events are sorted as follows: "speeding" event is sorted as 1, "collision" event is sorted as 2, "sudden speed drop" event is sorted as 3, "stopping" event is sorted as 4, and the subsequent generated events are sorted by 1 in turn.

[0107] It should be noted that the data to be tested may come from different types of data sources, rather than a single data source. For example, see [link to relevant documentation]. Figure 5 The velocity values ​​come from the velocity monitoring sequence, while the vibration values ​​may come from sensor logs. The events generated from multiple data sources to be detected, whether from the same or different data sources, can be sorted according to the generation time or acquisition time of the data, or according to the time of writing to the streaming detection module 202; the specific order is not limited.

[0108] Furthermore, in order to speed up the detection process and reduce the amount of data queries for each event detection, the streaming detection module 202 can use a real-time updated detection model. During the streaming detection process, it needs to interact with the cache module 204 to update the detection model or detection algorithm with real-time updated algorithm parameters in order to perform the detection.

[0109] Step 403: The flow cytometry detection module 202 inputs the generated event sequence into the root cause analysis module 203.

[0110] Step 404: The root cause analysis module 203 performs streaming root cause localization on the input event sequence.

[0111] It is worth noting that during the streaming conversion process, the streaming detection module 202 sequentially inputs the converted events into the root cause analysis module 203 in the form of a stream. The root cause analysis module 203 updates the event sequence based on the input events and performs root cause localization on the updated event sequence, thereby performing streaming root cause localization.

[0112] For example, combining Figure 4 The streaming detection module 202 inputs the converted events into the root cause analysis module 203 in a streaming format. For example, the root cause analysis module 203 receives event 1 (e.g., speeding event), event 2 (e.g., collision event), event 3 (e.g., sudden drop event), and event 4 (e.g., parking event) sequentially. Upon receiving event 1, root cause localization is performed on the event sequence 1 to be detected (including event 1). Upon receiving event 2, root cause localization is performed on the event sequence 2 to be detected (in the order of event 1 and event 2). Upon receiving event 3, root cause localization is performed on the event sequence 3 to be detected (in the order of event 1, event 2, and event 3). Upon receiving event 4, root cause localization is performed on the event sequence 4 to be detected (in the order of event 1, event 2, event 3, and event 4). Instead of performing root cause localization only on event sequence 4 after receiving events 1, 2, 3, and 4.

[0113] Subsequently received events can either combine with previous event sequences to form new event sequences, or a new event sequence can be generated starting with a subsequent event. This can also be determined based on different conditions. For example, based on the interval time, if the interval between two received events exceeds a preset value, the cumulative update of the previous event sequence ends, meaning the event sequence update begins with the newly received event. Alternatively, if the updated event sequence matches a user-defined rule, the cumulative update stops, and the newly received event becomes the starting point for the event sequence update. This embodiment does not impose limitations on this.

[0114] The root cause localization method used in this embodiment will be described in detail below.

[0115] First, as described above, the rule registration module 201 registers the event sequences and root causes in the rules to the root cause analysis module 203. For ease of explanation, the registered event sequences and root causes are denoted as data labels below. For example, event sequence a and root cause a are denoted as data label 1, and event sequence b and root cause b are denoted as data label 2.

[0116] Based on this, the root cause localization method adopted by the root cause analysis module 203 may include the following steps:

[0117] Step 404a: Obtain algorithm parameters for root cause localization of the event sequence to be detected in the target system. The algorithm parameters store the historical event states of the target system and can predict the root cause corresponding to the event sequence to be detected based on the historical event states of the target system.

[0118] In one example, the algorithm parameters used in this embodiment include a state transition probability matrix and an observation matrix. The state transition probability matrix stores the transition probabilities between any two root causes associated with the target system, denoted as […]. The observation matrix stores the transition probabilities between root causes and events associated with the target system, denoted as B. j,k =P(o k |e j ).

[0119] Let all events associated with the target system be the observation set O = {o1, o2, ..., o...} M All roots are due to the state set E = {e1, e2, ..., e}. N}. Where ei is any root cause in the state set E, o k Let N be any event in the observation set O. Here, N and M are both positive integers. N is the number of possible states, and M is the number of possible observations.

[0120] For example, root cause e i To the root cause e j The transition probability represents the probability that the current root cause is known to be root cause e. i Then the root cause e occurs at the next time point. j The probability of i taking values ​​from 1 to N, and the probability of j taking values ​​from 1 to N.

[0121] Another example is root cause e i To the event o k The transition probability represents the probability that the current root cause is known to be root cause e. i Then the current corresponding event is event o. k The probability of i is 1 to N, and the probability of k is 1 to M.

[0122] Taking the observation set as an example, in one implementation, the caching module 204 maintains a window of fixed time length (denoted as the cache time window) to cache the event sequence of the target system 10. The observation set contains all events occurring in the target system 10 recorded within the cache time window. For example, the historical event sequence cache of the caching module 204 caches events generated by the target system 10 at different times. The length of the cache time window is a fixed duration T, and the cache time window is a sliding window, with the starting point being the occurrence time of the most recently written event. When a new event is written to the caching module 204, the cache time window slides to the occurrence time of the new event. When an old event is more than a fixed duration T away from the current time, i.e., an old event that has slid out of the cache time window, it can be deleted from the caching module 204. In short, when the occurrence time of an event is more than a fixed duration T away from the current time, the event will be deleted from the cache time window.

[0123] Optionally, the starting point of the cache time window can also be the current time. This embodiment does not limit this. Since the data of the target system changes in characteristics over time, and the rules are also updated in real time as needed, this embodiment records all events within the current period along with the cache time window to determine the observation set, and updates the algorithm parameters based on the determined observation set. In this way, the root cause localization algorithm can actively update and learn as the data and rules are updated, without the need for manual intervention to relearn.

[0124] Accordingly, in this embodiment, the state set includes the root causes corresponding to the event sequences within the aforementioned observation set.

[0125] Step 404b: Match the current event sequence to be detected with the data label to obtain the matched event sequence root cause pair (referred to as historical event sequence root cause pair). Each event sequence root cause pair includes the event sequence and the root cause.

[0126] A match is formed if the event sequence is exactly the same as the event sequence in the data label. For example, data label 1 includes event sequence a and root cause a. If the event sequence to be detected is event sequence a, then event sequence a matches data label 1. The historical event root cause pair includes event sequence a and root cause a.

[0127] It is worth noting that different root causes can lead to the same event sequence. In other words, the same event sequence may correspond to multiple root causes, meaning multiple different root causes can correspond to the same event sequence. For example, the root cause analysis module 203 has two registered data labels: data label 1 (event sequence a and root cause a) and data label 3 (event sequence a and root cause c). Data label 1 and data label 3 include the same event sequence, but their root causes are different. Therefore, the event sequence to be detected may match multiple data labels. For instance, referring to the previous example, if the event sequence to be detected is event sequence a, then the matched data labels include data label 1 and data label 3.

[0128] Step 404c: Update the algorithm parameters of the obtained root cause localization algorithm based on the root cause pairs of the matched historical event sequences, calculate the updated algorithm parameters, and use the updated algorithm parameters to update the root cause localization algorithm.

[0129] In this way, the algorithm parameters of the root cause localization algorithm can be proactively updated as data and rules are updated, without the need for manual intervention and relearning.

[0130] Step 404d: Based on the event sequence to be detected and the updated root cause localization algorithm, determine the root cause sequence corresponding to the event sequence.

[0131] In this embodiment, the root cause localization algorithm can be implemented by a machine learning model. The process may include: after the event sequence to be detected is input into the root cause localization module 203, the updated algorithm parameters are obtained and written into the trained root cause localization model to obtain the updated root cause localization model; the event sequence to be detected is input into the updated root cause localization model to obtain the root cause localization result output by the root cause localization model.

[0132] For example, see Figure 6 Given a Hidden Markov Model (HMM) for root cause localization, including a state transition probability matrix and an observation matrix, when the event sequence to be detected is input into the root cause analysis module 203, the updated state transition probability matrix and observation matrix are obtained. Based on the event sequence to be detected and the updated HMM, the Viterbi decoding algorithm is used to dynamically search for possible generators of the event sequence (e1, e2, e3, ..., e...). n The topk Viterbi path of the root cause sequence (o1, o2, o3, ..., o n The specific formula is as follows:

[0133]

[0134] Where P represents probability, and n represents the number of event elements contained in the event sequence to be detected.

[0135] Root cause localization results are used to indicate whether the state of the target system 10 is normal or abnormal, in combination with Figure 6 Understandably, in one possible approach, the root cause localization result is "normal," indicating that the target system 10 is in a normal state. Alternatively, the root cause localization result may include the topk Viterbi path, i.e., the root cause sequence (o1, o2, o3, ..., o...). n In other words, the root cause sequence includes the top K root causes (denoted as topk root cause results) sorted by root cause probability from highest to lowest. Root cause probability refers to the probability that a cause might lead to this sequence of events. Figure 6 In this context, k can be a positive integer no greater than 3. For example, in one instance, the root cause sequence of the event sequence to be detected includes (serious car accident 0.9, battery failure 0.05, brake failure 0.04).

[0136] It should be noted that the method flow of steps 404a to 404d above is only an example. This embodiment also supports other root cause localization methods. For example, after the event sequence to be detected is input into the root cause localization module 203, the root cause of the event sequence is first performed. After the root cause localization, the algorithm parameters of the root cause localization algorithm are updated. For example, after the root cause localization, the event sequence is stored in a cache time window, and the event and label data are compared. If they match, an event sequence and root cause sequence pair are generated, and the state transition matrix A is updated according to the data. i,j and observation matrix B j,k This prepares for the next root cause analysis. This embodiment does not limit this approach.

[0137] Step 405: Root cause analysis module 203 outputs the root cause sequence.

[0138] For example, if the "serious traffic accident" tag exists in the topk root cause results, root cause analysis information is generated and output. This root cause analysis information includes vehicle information, vehicle event sequence, and root cause. The root cause analysis module 203 can return the root cause analysis information to the user, who can then locate the problem based on the root cause analysis information returned by the root cause location device 20 and take corresponding maintenance actions. It should be noted that step 405 is an optional step and is not a mandatory step.

[0139] With the above design, for the target system 10, the root cause localization device 20 uses real-time running data written in a streaming manner to perform streaming root cause localization, without having to pull the full historical data of the target system 10 from the database for root cause localization, thereby reducing the data requirement for root cause localization. Furthermore, since it can actively initiate root cause localization, it can identify potential problems and shorten the root cause localization latency.

[0140] In summary, a specific scenario-based method implementation is provided. Continuing with the vehicle-to-everything (V2X) scenario as an example, the target vehicle in the following text can be replaced with the target system. The method includes:

[0141] For the target vehicle's data to be detected during streaming, the following operations are performed on the target vehicle's data using a multi-round iterative approach:

[0142] Obtain the data to be tested in this round;

[0143] The current round's event sequence is updated based on the current round's data to be detected and the previous round's event sequence. Specifically, the current round's event sequence includes the previous round's event sequence and the events generated in this round, with the events generated in this round following the previous round's event sequence. The events generated in this round are generated based on the event detection algorithm already registered for the target vehicle and the current round's data to be detected. Typically, the previous round's event sequence includes events generated in each historical round. It is worth noting that "each historical round" here refers to historical rounds belonging to the same multi-round cyclic iteration, not all cyclic iterations, which will be explained later and will not be repeated here.

[0144] Root cause localization is performed on the sequence of events to be detected in this round to determine the root cause localization result of this round; wherein, the root cause localization result of this round is used to indicate whether the state of the target vehicle is normal or abnormal.

[0145] When the set iteration termination condition is met, the root cause localization result for this round is output. The set iteration termination condition includes a root cause localization result indicating that the target vehicle is in an abnormal state. In this case, the root cause localization result may include one or more potential root causes that led the target vehicle to generate the detected event sequence for this round, such as the aforementioned topk root cause sequence. Additionally, the set iteration termination condition may also include other conditions, such as reaching a set number of iterations or reaching a set iteration time, etc., without specific limitations.

[0146] For example, combining Figure 4Understanding: The streaming write operation involves writing monitoring data for the target vehicle. Here, the written monitoring data refers to the vehicle monitoring data to be detected in the streaming detection module 202. Whenever vehicle monitoring data is written, the streaming detection module 202 triggers a round of iterative operations. For example, if the streaming monitoring data consists of the target vehicle's speed monitoring sequence and vibration value data extracted from vibration sensor logs, including: monitoring data 1, monitoring data 2, monitoring data 3, monitoring data 4, monitoring data 5, monitoring data 6…, then the first round of iterative operations is triggered for monitoring data 1, including: using the target vehicle's speed detection algorithm to detect monitoring data 1 (assuming it's speed), generating event 1, inputting event 1 into the root cause localization module 203 for root cause localization, and obtaining the root cause localization result. Assuming the root cause localization result is normal; a second round of iterative iteration is triggered for monitoring data 2, including: using the target vehicle's speed detection algorithm to detect monitoring data 2 (assuming it's speed), generating event 2; based on the previous round's event sequence (event 1) and the event 2 generated in this round, obtaining the current round's event sequence (event 1 and event 2); inputting the current round's event sequence into the root cause localization module 203 for root cause localization, obtaining the root cause localization result, assuming the root cause localization result is normal; a third round of iterative iteration is triggered for monitoring data 3, including: using the target vehicle's anomaly detection algorithm to detect the target vehicle's speed, generating event 2; based on the previous round's event sequence (event 1) and the generated event 2 ...2); and a third round of iterative iteration is triggered for monitoring data 3, including: using the target vehicle's anomaly detection algorithm to detect the target vehicle's speed, generating event 2; and a third round of iterative iteration is triggered for monitoring data 3, including: using the target vehicle's anomaly detection algorithm to detect the target vehicle's speed, generating event 2; and a third round of iterative iteration is triggered for monitoring data 3, including: using the target vehicle's anomaly detection algorithm to detect the target vehicle's speed, generating event 2; and a third round of iterative iteration is triggered for monitoring data 3, including: using the target vehicle's anoma The constant rise detection algorithm detects monitoring data 3 (assuming it is a vibration value) and generates event 3. Based on the event sequence of the previous round (event 1, event 2) and the event 3 generated in this round, the event sequence of this round (event 1, event 2, event 3) is obtained. The event sequence of this round is input into the root cause localization module 203 for root cause localization, and the root cause localization result is obtained. It is assumed that the root cause localization result is normal. For monitoring data 4, a fourth round of loop iteration is triggered, including: using the target vehicle speed detection algorithm to detect monitoring data 4 (assuming it is a speed value) and generate event 4, based on the previous round's... The event sequence (event 1, event 2, event 3) and event 4 generated in this round are used to obtain the event sequence (event 1, event 2, event 3, event 4) for this round. The event sequence for this round is input into the root cause localization module 203 for root cause localization to obtain the root cause localization result. If the root cause localization result includes the topk root cause sequence, it indicates that the target vehicle is in an abnormal state. At this point, the set iteration termination condition is met, and the root cause analysis information is output. The root cause analysis information may include the topk root cause sequence, and optionally may also include one or more pieces of information such as the target vehicle information and the event sequence of the target vehicle determined in this round.

[0147] This completes one round of multi-cycle iteration.

[0148] It is worth noting that, in one implementation, after a complete multi-round iteration is achieved, the next multi-round iteration can still be restarted. A new iteration can be started as long as there is data to be detected. For example, the first round of iteration is triggered again for monitoring data 5 in the example above, including: using the target vehicle speed detection algorithm to detect monitoring data 5 (assuming it is speed), generating event 1', inputting event 1' into the root cause localization module 203 for root cause localization, and obtaining the root cause localization result, assuming the root cause localization result is normal; the second round of iteration is triggered for monitoring data 6, including: using the target vehicle speed detection algorithm to detect monitoring data 6 (assuming it is speed), generating event 2', based on the event sequence of the previous round (event 1') and the event 2' generated in this round, obtaining the event sequence of this round (event 1' and event 2'), inputting the event sequence of this round into the root cause localization module 203 for root cause localization, and obtaining the root cause localization result, assuming the root cause localization result is normal, and so on.

[0149] In another implementation, after the set iteration termination condition is met, iteration can continue. For example, the fifth round of loop iteration is triggered for the monitoring data 5 in the above example, including: using the target vehicle speed detection algorithm to detect the monitoring data 5 (assuming it is speed), generating event 5, and obtaining the event sequence of the current round (event 1, event 2, event 3, event 4) based on the event sequence of the previous round (event 1, event 2, event 3, event 4, event 5) and the event 5 generated in this round, and inputting the event sequence of the current round into the root cause localization module 203 for root cause localization to obtain the root cause localization result, and so on.

[0150] It should be noted that this embodiment can also set the number of events included in the event sequence that needs to be root cause located. When the number of events included in the event sequence of this round reaches the set number of events, root cause location will be performed on the event sequence of this round. For example, when the set number of events is N, root cause location is not performed in the first N-1 rounds, and root cause location is performed in the Nth round.

[0151] For example, the implementation of the streaming detection module 202 in the root cause localization device 20 will be described below. Similarly, the implementation of the root cause localization module 203, the rule registration module 201, and the model update module 205 can refer to the implementation of the streaming detection module 202.

[0152] When implemented in software, the streaming detection module 202 can be an application or code block running on a computer device. The computer device can be at least one of a physical host, virtual machine, container, or other computing device. Furthermore, there can be one or more computer devices. For example, the streaming detection module 202 can be an application running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers running the application can be distributed within the same availability zone (AZ) or in different AZs. Similarly, the multiple hosts / virtual machines / containers running the application can be distributed within the same region or in different regions. Typically, a region can include multiple AZs.

[0153] Similarly, multiple hosts / virtual machines / containers used to run the application can be distributed within the same Virtual Private Cloud (VPC) or across multiple VPCs. Typically, a region can include multiple VPCs, and a VPC can include multiple Availability Zones (AZs).

[0154] When implemented in hardware, the streaming detection module 202 may include at least one computing device, such as a server. Alternatively, the streaming detection module 202 may be implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be implemented using a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), generic array logic (GAL), or any combination thereof.

[0155] The streaming detection module 202 includes multiple computing devices that can be distributed within the same Availability Zone (AZ) or in different AZs. Similarly, the streaming detection module 202 can be distributed within the same region or in different regions. Likewise, the streaming detection module 202 can be distributed within the same Virtual Private Cloud (VPC) or in multiple VPCs. These multiple computing devices can be any combination of computing devices such as servers, ASICs, PLDs, CPLDs, FPGAs, and GALs.

[0156] It should be noted that the module division in this embodiment is illustrative and only represents a logical functional division. In actual implementation, there may be other division methods. The functional modules in this embodiment can be integrated into one module, or each module can exist physically separately, or two or more modules can be integrated into one module. For example, the streaming detection module 202 and the root cause localization module 203 can be integrated into one module, or the rule registration module 201 and the model update module 205 can be the same module. The integrated units described above can be implemented in hardware or as software functional units.

[0157] This application also provides a computing device 700. For example... Figure 7 As shown, the computing device 700 includes a bus 702, a processor 704, a memory 706, and a communication interface 708. The processor 704, the memory 706, and the communication interface 708 communicate with each other via the bus 702. The computing device 700 can be a server or a terminal device. It should be understood that this application does not limit the number of processors and memories in the computing device 700.

[0158] The 702 bus can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of representation, Figure 7 The bus 702 may be represented by a single line, but this does not mean that there is only one bus or one type of bus. The bus 702 may include a path for transmitting information between various components of the computing device 700 (e.g., memory 706, processor 704, communication interface 708).

[0159] Processor 704 may include any one or more processors such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).

[0160] The memory 706 may include volatile memory, such as random access memory (RAM). The processor 704 may also include non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid state drive (SSD).

[0161] The memory 706 stores executable program code, and the processor 704 executes the executable program code to implement the functions of the aforementioned streaming detection module 202 and root cause localization module 203, respectively. Optionally, it can also implement the functions of the rule registration module 201 and the model update module 205, thereby realizing the root cause localization method provided in this embodiment. That is, the memory 706 stores instructions for the root cause localization device 20 to execute the root cause localization method provided in this application.

[0162] The communication interface 708 uses transceiver modules, such as, but not limited to, network interface cards and transceivers, to enable communication between the computing device 700 and other devices or communication networks.

[0163] This application also provides a computing device cluster. The computing device cluster includes at least one computing device. The computing device may be a server. In some embodiments, the computing device may also be a desktop computer, a laptop computer, or a smartphone, or other terminal device.

[0164] like Figure 8 As shown, the computing device cluster includes at least one computing device 700. The memory 706 of one or more computing devices 700 in the computing device cluster may store the same instructions for executing the root cause localization method.

[0165] In some possible implementations, the memory 706 of one or more computing devices 700 in the computing device cluster may also store partial instructions for executing the root cause localization method. In other words, a combination of one or more computing devices 700 can jointly execute the instructions for executing the root cause localization method.

[0166] It should be noted that the memory 706 in different computing devices 700 within the computing device cluster can store different instructions, each used to execute a portion of the functions of the data access device. That is, the instructions stored in the memory 706 of different computing devices 700 can implement the functions of one or more of the aforementioned streaming detection module 202, root cause localization module 203, rule registration module 201, and model update module 205.

[0167] In some possible implementations, one or more computing devices in a computing device cluster can be connected via a network. This network can be a wide area network (WAN) or a local area network (LAN), etc. Figure 9 One possible implementation is shown. For example... Figure 9 As shown, the two computing devices 700A and 700B are connected via a network. Specifically, they are connected to the network through the communication interfaces in each computing device. In this possible implementation, the memory 706 in computing device 700A stores instructions for performing the functions of the streaming detection module 202 and the rule registration module 201. Simultaneously, the memory 706 in computing device 700B stores instructions for performing the functions of the root cause localization module 203 and the model update module 205.

[0168] It should be understood that Figure 9 The functions of the computing device 700A shown can also be performed by multiple computing devices 700. Similarly, the functions of the computing device 700B can also be performed by multiple computing devices 700.

[0169] This application also provides another computing device cluster. The connection relationships between the computing devices in this computing device cluster can be similarly referred to... Figure 8 and Figure 9 The connection method of the computing device cluster is different in that the memory 706 of one or more computing devices 700 in the computing device cluster can store the same instructions for executing the root cause localization method.

[0170] In some possible implementations, the memory 706 of one or more computing devices 700 in the computing device cluster may also store partial instructions for executing the root cause localization method. In other words, a combination of one or more computing devices 700 can jointly execute the instructions for executing the root cause localization method.

[0171] This application also provides a computer program product containing instructions. The computer program product may be a software or program product containing instructions, capable of running on a computing device or stored on any usable medium. When the computer program product is run on at least one computing device, it causes the at least one computing device to perform a root cause localization method, or instructs the computing device to perform a root cause localization method.

[0172] This application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that a computing device can store, or a data storage device such as a data center containing one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive). The computer-readable storage medium includes instructions that instruct the computing device to perform a root cause localization method, or instruct the computing device to perform a root cause localization method.

[0173] Optionally, the computer execution instructions in the embodiments of this application may also be referred to as application code, and the embodiments of this application do not specifically limit this.

[0174] Those skilled in the art will understand that the various numerical designations, such as "first," "second," etc., used in this application are merely for descriptive convenience and are not intended to limit the scope of the embodiments of this application, nor do they indicate a sequential order. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one" refers to one or more. "At least two" refers to two or more. "At least one," "any one," or similar expressions refer to any combination of these items, including any combination of single or multiple items. For example, at least one of a, b, or c can represent: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or multiple. "Multiple" refers to two or more, and other quantifiers are similar. Furthermore, for elements appearing in the singular forms "a," "an," and "the," unless the context explicitly specifies otherwise, they do not imply "one or only one," but rather "one or more." For example, "a device" implies one or more such devices.

[0175] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state disk (SSD)).

[0176] The various illustrative logic units and circuits described in the embodiments of this application can be implemented or operate the described functions using a general-purpose processor, digital signal processor, application-specific integrated circuit (ASIC), field-programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof. The general-purpose processor can be a microprocessor; alternatively, it can also be any conventional processor, controller, microcontroller, or state machine. The processor can also be implemented using a combination of computing devices, such as a digital signal processor and a microprocessor, multiple microprocessors, one or more microprocessors combined with a digital signal processor core, or any other similar configuration.

[0177] The steps of the methods or algorithms described in the embodiments of this application can be directly embedded in hardware, software units executed by a processor, or a combination of both. The software units can be stored in RAM, flash memory, ROM, EPROM, EEPROM, registers, hard disks, removable disks, CD-ROMs, or any other form of storage medium in the art. Exemplarily, the storage medium can be connected to the processor so that the processor can read information from and write information to the storage medium. Optionally, the storage medium can also be integrated into the processor. The processor and storage medium can be housed in an ASIC.

[0178] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0179] Although this application has been described in conjunction with specific features and embodiments, it is obvious that various modifications and combinations can be made thereto without departing from the spirit and scope of this application. Accordingly, this specification and drawings are merely illustrative descriptions of the application as defined by the appended claims, and are considered to cover any and all modifications, variations, combinations, or equivalents within the scope of this application. Clearly, those skilled in the art can make various alterations and modifications to this application without departing from its scope. Thus, if such modifications and modifications fall within the scope of the claims and their equivalents, this application is also intended to include such modifications and modifications.

Claims

1. A root cause localization method, characterized in that, include: Perform the following operations on the target system's test data using a multi-round iterative approach: Obtain the data to be tested in this round; The sequence of events to be detected in the current round is determined. The sequence of events to be detected in the current round includes the sequence of events to be detected in the previous round and the events determined in the current round. The events determined in the current round are generated based on the event detection algorithm associated with the target system and the event data to be detected in the current round. The event detection algorithm is used to convert the event data to be detected into events. Root cause localization is performed on the sequence of events to be detected in this round to determine the root cause localization results for this round; When the set iteration termination condition is met, the root cause localization result of this round is output. The set iteration termination condition includes the root cause localization result of this round indicating that the target system is in an abnormal state. The root cause localization result includes one or more potential root causes that cause the target system to generate the sequence of events to be detected in this round.

2. The method as described in claim 1, characterized in that, The event detection algorithm associated with the target system is determined based on rules submitted by the user for the configuration of the target system. The rules include a specified event sequence and the root cause corresponding to the specified event sequence.

3. The method as described in claim 2, characterized in that, The method further includes: An application programming interface (API) is provided for users to configure the rules. The API includes multiple fields, including an event sequence field and a root cause field. The system receives API content submitted by the user. The API content includes the multiple fields, event information input for the event sequence field, and root cause information input for the root cause field. The event information is used to describe the multiple events included in the set event sequence and the order of the multiple events. The root cause information is used to describe the root cause corresponding to the specified event sequence.

4. The method as described in claim 2, characterized in that, The method further includes: A configuration interface is provided, which is used by the user to configure the rules of the target system; The rules configured by the user are obtained from the configuration interface.

5. The method according to any one of claims 1-4, characterized in that, Root cause localization of the event sequence to be detected in this round includes: Based on the event sequence to be detected in this round and the root cause localization model, the root cause localization result is determined; the root cause localization model is used to predict the root cause based on the event sequence.

6. The method as described in claim 5, characterized in that, The root cause localization model includes the following parameters: state transition probability matrix and observation matrix; The state transition probability matrix includes the transition probabilities between any two root causes associated with the target system, where root cause e i To the root cause e k The transition probability represents the current root cause e. i Then it will become the root cause e at the next time point. k The probability of; The observation matrix includes the transition probabilities between any root cause associated with the target system and any event associated with the target system, where the root cause e i To the event o m The transition probability represents the probability that the current root cause is known to be root cause e. i Then the current corresponding event is event o. m The probability of.

7. The method as described in claim 6, characterized in that, The events associated with the target system include: events that occur in the target system within a sliding time window of a fixed length, and / or events contained in the rules configured for the target system; The root causes associated with the target system include root causes corresponding to event sequences belonging to the target system recorded within the sliding time window, and / or root causes contained in the rules configured for the target system.

8. The method according to any one of claims 1-7, characterized in that, The root cause localization result is a topk Viterbi path obtained by searching using the Viterbi decoding algorithm based on the event sequence to be detected and the root cause localization model. The topk Viterbi path is used to indicate the top k root causes with the probability of generating the event sequence to be detected arranged from high to low, where k is a positive integer.

9. A root cause localization device, characterized in that, include: Perform the following operations on the target system's test data using a multi-round iterative approach: The flow cytometry module is used to acquire the data to be detected in the current round; The streaming detection module is further configured to determine the sequence of events to be detected in the current round. The sequence of events to be detected in the current round includes the sequence of events to be detected in the previous round and the events determined in the current round. The events determined in the current round are generated based on the event detection algorithm associated with the target system and the data to be detected in the current round. The event detection algorithm is configured to convert the data to be detected into events. The root cause analysis module is used to locate the root cause of the event sequence to be detected in this round and determine the root cause location result of this round. The root cause analysis module is further configured to output the root cause localization result of the current iteration when a set iteration termination condition is met; the set iteration termination condition includes the root cause localization result of the current iteration indicating that the target system is in an abnormal state, and the root cause localization result includes one or more potential root causes that cause the target system to generate the sequence of events to be detected in the current iteration.

10. The apparatus as claimed in claim 9, characterized in that, The event detection algorithm associated with the target system is determined based on rules submitted by the user for the configuration of the target system. The rules include a specified event sequence and the root cause corresponding to the specified event sequence.

11. The apparatus as claimed in claim 10, characterized in that, The device also includes a rule registration module; The rule registration module is used to provide an application programming interface (API) for users to configure the rules. The API includes multiple fields, including an event sequence field and a root cause field. The API receives API content submitted by the user. The API content includes the multiple fields, event information input for the event sequence field, and root cause information input for the root cause field. The event information is used to describe the multiple events included in the set event sequence and the order of the multiple events. The root cause information is used to describe the root cause corresponding to the specified event sequence.

12. The apparatus as claimed in claim 10, characterized in that, The device also includes a rule registration module; The rule registration module is used to provide a configuration interface for users to configure rules for the target system; and to obtain the rules configured by the user from the configuration interface.

13. The apparatus according to any one of claims 9-12, characterized in that, The root cause analysis module, when performing root cause localization on the sequence of events to be detected in this round, is specifically used for: Based on the event sequence to be detected in this round and the root cause localization model, the root cause localization result is determined; the root cause localization model is used to predict the root cause based on the event sequence.

14. The apparatus as claimed in claim 13, characterized in that, The root cause localization model includes the following parameters: state transition probability matrix and observation matrix; The state transition probability matrix includes the transition probabilities between any two root causes associated with the target system, where root cause e i To the root cause e k The transition probability represents the current root cause e. i Then it will become the root cause e at the next time point. k The probability of; The observation matrix includes the transition probabilities between any root cause associated with the target system and any event associated with the target system, where the root cause e i To the event o m The transition probability represents the probability that the current root cause is known to be root cause e. i Then the current corresponding event is event o. m The probability of.

15. The apparatus as claimed in claim 14, characterized in that, The events associated with the target system include: events that occur in the target system within a sliding time window of a fixed length, and / or events contained in the rules configured for the target system; The root causes associated with the target system include root causes corresponding to event sequences belonging to the target system recorded within the sliding time window, and / or root causes contained in the rules configured for the target system.

16. The apparatus according to any one of claims 9-15, characterized in that, The root cause localization result is a topk Viterbi path obtained by searching using the Viterbi decoding algorithm based on the event sequence to be detected and the root cause localization model. The topk Viterbi path is used to indicate the top k root causes with the probability of generating the event sequence to be detected arranged from high to low, where k is a positive integer.

17. A computing device cluster, characterized in that, It includes at least one computing device, each computing device including a processor and memory; The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device to cause the cluster of computing devices to perform the method as described in any one of claims 1 to 8.

18. A computer program product containing instructions, characterized in that, When the instruction is executed by the computing device cluster, the computing device cluster causes the computing device cluster to perform the method as described in any one of claims 1 to 8.

19. A computer-readable storage medium, characterized in that, Includes computer program instructions, which, when executed by a cluster of computing devices, perform the method as described in any one of claims 1 to 8.