A method and system for anomaly localization

A system-level approach combining synchronous and asynchronous location components addresses the challenge of rapid fault localization in monitoring systems by providing immediate and stored analysis results, enhancing the monitoring system's fault identification capabilities.

CN114490150BActive Publication Date: 2025-07-15DUXIAOMAN TECH (BEIJING) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111611880.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-27
Publication Date
2025-07-15
Estimated Expiration
2041-12-27

AI Technical Summary

Technical Problem

The existing monitoring system lacks effective auxiliary information when locating faults, making it difficult to quickly and accurately locate the causes of abnormalities.

Method used

Using a combination of synchronous positioning components and asynchronous positioning components, synchronous positioning results and asynchronous positioning results are obtained through synchronous positioning operations and asynchronous positioning operations, abnormal positioning information is generated, and query indication information is provided so that users can further query asynchronous positioning results.

Benefits of technology

The fault positioning capability of the monitoring system is improved, and users can quickly and accurately understand the abnormal details, improving positioning efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114490150B_ABST
    Figure CN114490150B_ABST
Patent Text Reader

Abstract

The present application provides a method and a system for anomaly localization. The method includes: if the current anomaly detection result indicates an anomaly alarm, the alarm module sends a localization request message to the localization module, where the localization module includes at least one synchronous localization component and at least one asynchronous localization component; the localization module receives the localization request message and performs synchronous localization operations and asynchronous localization operations according to the localization request message; the alarm module generates anomaly localization information based on the synchronous localization result provided by the localization module and sends the anomaly localization information to the target user, where the anomaly localization information includes the synchronous localization result and query indication information corresponding to the anomaly detection result. The present application solves the problem of difficult fault localization after a system anomaly in the monitoring field, and can improve the auxiliary localization ability of the entire monitoring system through a general system architecture for anomaly localization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular, to a technical solution for anomaly localization. Background Art

[0002] For any information system, anomalies may occur. Quickly, accurately, and effectively locating anomalies is the key to promptly eliminating anomalies and repairing the information system. In addition, timely discovering the causes of anomalies and warning of possible anomaly results helps information system administrators realize the severity of various anomaly causes and eliminate potential anomaly hazards in a timely manner. With the development of network technology, anomaly localization in monitoring systems has become a major focus in the field of intelligent operation and maintenance. When an anomaly is monitored, identifying the element that is most likely the root cause of the anomaly can facilitate further repair and loss prevention. Existing technologies usually obtain metric data to prompt an alarm when an anomaly alarm occurs. For example, when the monitoring system detects an anomaly alarm, it obtains metric data such as the number of errors, alarm objects, anomaly time, product line, and anomaly module and provides them to the alarm recipient. Summary of the Invention

[0003] The purpose of this application is to provide a technical solution for anomaly localization.

[0004] According to an embodiment of this application, a method for anomaly localization is provided, which includes:

[0005] If the current anomaly detection result indicates an anomaly alarm, the alarm module sends a localization request message to the localization module, where the localization module includes at least one synchronous localization component and at least one asynchronous localization component;

[0006] The localization module receives the localization request message and, according to the localization request message, performs a synchronous localization operation and an asynchronous localization operation. The synchronous localization operation includes traversing and executing the at least one synchronous localization component to obtain a synchronous localization result and providing the synchronous localization result to the alarm module. The asynchronous localization operation includes constructing a task sending queue based on the at least one asynchronous localization component, consuming the task sending queue through an asynchronous thread to execute the at least one asynchronous localization component to obtain an asynchronous localization result, and writing the asynchronous localization result into a database;

[0007] The alarm module generates anomaly localization information according to the synchronous localization result and sends the anomaly localization information to the target user, where the anomaly localization information includes the synchronous localization result and query indication information corresponding to the anomaly detection result.

[0008] According to another embodiment of the present application, a system for anomaly localization is provided, wherein the system includes an alarm module and a localization module, and the localization module includes at least one synchronous localization component and at least one asynchronous localization component;

[0009] Among them, the alarm module is configured to: if the current anomaly detection result indicates an anomaly alarm, send a localization request message to the localization module, and generate anomaly localization information based on the synchronous localization result provided by the localization module, and send the anomaly localization information to the target user, where the anomaly localization information includes the synchronous localization result and query indication information corresponding to the anomaly detection result;

[0010] Among them, the localization module is configured to: receive the localization request message, and perform a synchronous localization operation and an asynchronous localization operation according to the localization request message, where the synchronous localization operation includes traversing and executing the at least one synchronous localization component to obtain the synchronous localization result, and providing the synchronous localization result to the alarm module, and the asynchronous localization operation includes constructing a task sending queue based on the at least one asynchronous localization component, consuming the task sending queue through an asynchronous thread to execute the at least one asynchronous localization component to obtain an asynchronous localization result, and writing the asynchronous localization result into a database.

[0011] According to another embodiment of the present application, a computer device is further provided, wherein the computer device includes: a memory for storing one or more programs; one or more processors connected to the memory, and when the one or more programs are executed by the one or more processors, the one or more processors execute the method for anomaly localization according to the present application.

[0012] According to another embodiment of the present application, a computer-readable storage medium is further provided, on which a computer program is stored, and the computer program can be executed by a processor to execute the method for anomaly localization according to the present application.

[0013] Compared with the prior art, the present application has the following advantages: when an anomaly alarm is detected, it can synchronously request the localization module to perform a synchronous localization operation and an asynchronous localization operation, and then can reach the user the anomaly localization information generated based on the synchronous localization result obtained by traversing all synchronous localization components, and since the anomaly localization information includes query indication information for querying the asynchronous localization result, the user can query the asynchronous localization result based on the query indication information to further understand the alarm details, so as to facilitate quickly and accurately completing the fault localization; the present application solves the problem of difficult fault localization after system anomalies in the monitoring field, and can improve the auxiliary localization ability of the entire monitoring system through a general system architecture for anomaly localization. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Other features, objects, and advantages of the present application will become more apparent from the following detailed description of non - limiting embodiments read in conjunction with the accompanying drawings:

[0015] Figure 1 A flowchart showing the process of a method for anomaly localization according to an embodiment of the present application;

[0016] Figure 2 A flowchart showing the process of a method for anomaly localization according to an example of the present application;

[0017] Figure 3 A structural diagram showing the structure of a system for anomaly localization according to an example of the present application;

[0018] Figure 4 An exemplary system that can be used to implement the various embodiments described in the present application is shown.

[0019] Like or similar reference numerals in the drawings represent like or similar components. Detailed implementation manners

[0020] Before discussing the exemplary embodiments in more detail, it should be noted that some exemplary embodiments are described as processes or methods depicted as flowcharts. Although the flowcharts describe the operations as sequential processes, many of the operations can be implemented in parallel, concurrently, or simultaneously. In addition, the order of the operations can be rearranged. The process can be terminated when its operations are completed, but it can also have additional steps not included in the drawings. The process can correspond to a method, function, procedure, subroutine, sub - program, etc.

[0021] As used herein, the term "device" refers to an intelligent electronic device that can perform a predetermined processing process such as numerical calculation and / or logical calculation by running a predetermined program or instruction. It may include a processor and a memory, and the processor executes the program instructions pre - stored in the memory to perform the predetermined processing process, or a dedicated integrated circuit (ASIC), a field - programmable gate array (FPGA), a digital signal processor (DSP), etc. performs the predetermined processing process, or a combination of the above two is used to achieve it.

[0022] The technical solution of this application is mainly implemented by computer devices. Among them, the computer devices include network devices and user devices. The network devices include, but are not limited to, a single network server, a server group composed of multiple network servers, or a cloud composed of a large number of computers or network servers based on Cloud Computing. Among them, Cloud Computing is a type of distributed computing, which is a super virtual computer composed of a group of loosely coupled computer sets. The user devices include, but are not limited to, PC machines, tablet computers, smart phones, IPTVs, PDAs, wearable devices, etc. Among them, the computer devices can run independently to implement this application, or can be connected to the network and implement this application through interaction with other computer devices in the network. Among them, the network where the computer devices are located includes, but is not limited to, the Internet, wide area network, metropolitan area network, local area network, VPN network, wireless ad hoc network (Ad Hoc network), etc.

[0023] It should be noted that the above computer devices are only examples. Other existing or future possible computer devices that are applicable to this application should also be included within the protection scope of this application and are hereby incorporated by reference.

[0024] The methods discussed later in this article (some of which are illustrated by flowcharts) can be implemented by hardware, software, firmware, middleware, microcode, hardware description language, or any combination thereof. When implemented with software, firmware, middleware, or microcode, the program code or code segments for implementing the necessary tasks can be stored in a machine or computer-readable medium (such as a storage medium). One or more processors can implement the necessary tasks.

[0025] The specific structures and functional details disclosed herein are merely representative and are for the purpose of describing the exemplary embodiments of this application. However, this application can be specifically implemented in many alternative forms and should not be construed as being limited only to the embodiments set forth herein.

[0026] It should be understood that although terms such as "first", "second", etc. may be used herein to describe various units, these units should not be limited by these terms. These terms are only used to distinguish one unit from another. For example, without departing from the scope of the exemplary embodiments, the first unit can be referred to as the second unit, and similarly the second unit can be referred to as the first unit. The term "and / or" used herein includes any and all combinations of one or more of the listed associated items.

[0027] The terms used herein are for the purpose of describing particular embodiments only and are not intended to limit the exemplary embodiments. Unless the context clearly dictates otherwise, the singular forms "a", "an" used herein are also intended to include the plural. It should also be understood that the terms "comprises" and / or "comprising" specify the presence of the stated features, integers, steps, operations, units and / or components, and do not preclude the presence or addition of one or more other features, integers, steps, operations, units, components and / or combinations thereof.

[0028] It should also be noted that in some alternative implementations, the functions / actions mentioned may occur in a different order than that indicated in the figures. For example, depending on the functions / actions involved, two consecutively shown figures may actually be executed substantially simultaneously or sometimes in the reverse order.

[0029] The applicant has found that when the prior art monitoring system performs fault location, it lacks information that can help with auxiliary location. For example, after the alarm information of the prior art reaches the user, the user can only know that some indicators of a certain module have problems at a certain time point, and it is very difficult to quickly perform fault location based on this alone. In response to the above technical problems, the present application proposes a general location capability process for system-level planning. The main purpose is to obtain effective auxiliary information (including the synchronous location results obtained based on the synchronous location components and the asynchronous location results obtained based on the asynchronous location components) through a set of location components in the location module when the alarm reaches the user, so as to help the user perform fault location.

[0030] The solution of the present application will be further described in detail below with reference to the accompanying drawings.

[0031] Figure 1The figure shows a schematic flowchart of a method for anomaly location according to an embodiment of the present application. The method of this embodiment is implemented by a computer device; in some embodiments, the method of this embodiment is implemented by a network device equipped with a monitoring system. The method according to this embodiment includes step S11, step S12, and step S13. In step S11, if the current anomaly detection result indicates an anomaly alarm, the alarm module sends location request information to the location module, where the location module includes at least one synchronous location component and at least one asynchronous location component; in step S12, the location module receives the location request information and, according to the location request information, performs a synchronous location operation and an asynchronous location operation, where the synchronous location operation includes traversing and executing the at least one synchronous location component to obtain a synchronous location result and providing the synchronous location result to the alarm module, and the asynchronous location operation includes constructing a task sending queue based on the at least one asynchronous location component, consuming the task sending queue through an asynchronous thread to execute the at least one asynchronous location component to obtain an asynchronous location result, and writing the asynchronous location result into a database; in step S13, the alarm module generates anomaly location information according to the synchronous location result and sends the anomaly location information to the target user, where the anomaly location information includes the synchronous location result and query indication information corresponding to the anomaly detection result.

[0032] In step S11, if the current anomaly detection result indicates an anomaly alarm, the alarm module sends location request information to the location module, where the location module includes at least one synchronous location component and at least one asynchronous location component.

[0033] In some embodiments, the monitoring system includes an alarm module and a location module. The alarm module is used to issue an anomaly alarm when an anomaly is detected, and the location module is used to perform anomaly location for the anomaly alarm. In some embodiments, the location module includes multiple location components. Among them, the location component is the basic model for describing the location task in the monitoring system. One monitoring configuration can be associated with 1 or more location components, and different monitoring configurations can configure different location component input parameters. When an anomaly is monitored, it will trigger the execution of each location component in the location module. The location components in the present application include synchronous location components and asynchronous location components. In some embodiments, the synchronous location component is used to perform location tasks with relatively fast processing or not very large location data volume. In some embodiments, the asynchronous location component is used to perform location tasks with relatively slow processing or very large location data volume (not suitable for directly sending to users and needs to be presented separately). In some embodiments, each asynchronous location component or synchronous location component is a pluggable module, which can be quickly and freely plugged in according to different services, and the operator can arbitrarily combine each location component based on requirements.

[0034] In some embodiments, the synchronous positioning component includes an exception log information component, which is used to obtain the exception log information corresponding to the current exception alarm. As an example, when the exception log information component is executed, based on the time period when the exception alarm occurs, it tracks the log information of that time period and uses it as the exception log information. In some embodiments, the synchronous positioning component includes an alarm callback component, which is used to obtain alarm-related information (such as monitoring items) corresponding to the current exception alarm. After sending the alarm-related information to the target user, the target user can process it based on the corresponding positioning mechanism. In some embodiments, the synchronous positioning component includes the exception log information component and the alarm callback component. It should be noted that the above synchronous positioning component is only an example and not a limitation to this application. In actual applications, the number and functions of the synchronous positioning component can be set based on actual monitoring requirements.

[0035] In some embodiments, the asynchronous positioning component includes a dimension ratio component, which is used to analyze the ratio of different dimensions under alarm conditions. In some embodiments, after the monitoring system fails, the dimension ratio component is used to assist in positioning by analyzing the ratio of some dimensions (such as computer room, machine, IP, url, error code, status code, etc.) under alarm conditions. That is, the dimension ratio component can be used to determine which dimension is the root cause of the problem. As an example, the dimension ratio component is used to analyze the ratio of each dimension during the time period when the exception alarm occurs. If it is found that the number of alarms in a certain computer room is the largest, then the possibility of a failure in that computer room is the greatest.

[0036] In some embodiments, the asynchronous positioning component includes an alarm event correlation component, which is used to calculate the correlation relationship between alarm events by analyzing historical alarm data. In some embodiments, after a failure occurs at the bottom layer of the monitoring system, there are often a large number of alarms. The alarm event correlation component is used to analyze historical alarm data and calculate the correlation relationship between alarms through a clustering analysis algorithm. After failures occur in both the upstream system and the bottom layer system, it can indicate which part of the exception causes the alarm and provide the calculated correlation relationship to the user to assist in decision-making. That is, the alarm correlation analysis component can be used to infer the root cause of the exception. Since the causal relationship between alarms can be judged through the correlation relationship between alarm events, the result obtained by executing the alarm event correlation component can improve the positioning efficiency.

[0037] In some embodiments, the asynchronous positioning component includes an associated online operation component, which is used to determine whether the current abnormal alarm is caused by an online operation by analyzing the system operation records during the abnormal time period. In some embodiments, the abnormal time period is also the time period when the current abnormal alarm occurs. The associated online operation component is used to analyze whether there is an online event when the current abnormal alarm occurs according to the system operation records during the abnormal time period. If so, it is determined that the current abnormal alarm is caused by an online operation. Otherwise, the current abnormal alarm has nothing to do with the online operation, so as to determine whether the fault is caused by human operation (or human error).

[0038] In some embodiments, the asynchronous positioning component includes a multi-root cause analysis component, which is used to analyze the proportion of the change of each dimension value in different dimensions to the abnormality and the change difference of each dimension value. In some embodiments, the multi-root cause analysis component is used to measure the proportion of the change of the dimension value A in dimension A to the abnormality and the change difference of the dimension value A in dimension A through the Adtributor algorithm i to find out the value with a relatively large proportion in the total ratio and a relatively large change difference in A itself. This value may be the value that affects the final result, that is, the multi-root cause analysis component can analyze which dimension value changes in different dimensions lead to the change of the final result value. By executing the multi-root cause analysis component, the contribution degree of different dimensions to the result can be calculated, so as to recommend the most likely factors that may affect the result. i of the change difference, and find out the value of dimension value A i with a relatively large proportion in the total ratio and a relatively large change difference in A itself. This value may be the value that affects the final result, that is, the multi-root cause analysis component can analyze which dimension value changes in different dimensions lead to the change of the final result value. By executing the multi-root cause analysis component, the contribution degree of different dimensions to the result can be calculated, so as to recommend the most likely factors that may affect the result. i of the change difference, and find out the value of dimension value A with a relatively large proportion in the total ratio and a relatively large change difference in A itself. This value may be the value that affects the final result, that is, the multi-root cause analysis component can analyze which dimension value changes in different dimensions lead to the change of the final result value. By executing the multi-root cause analysis component, the contribution degree of different dimensions to the result can be calculated, so as to recommend the most likely factors that may affect the result.

[0039] In some embodiments, the asynchronous positioning component includes a trace information positioning component, which is used to trace the request call chain log and troubleshoot the abnormality according to the call chain log. In some embodiments, the asynchronous positioning component is used to obtain the request call chain according to the abnormal information, and troubleshoot the abnormality according to the call chain log to determine which code line has a fault.

[0040] In some embodiments, the asynchronous positioning component includes one or more of a dimension proportion component, an alarm event association component, an associated online operation component, a multi-root cause analysis component, and a trace information positioning component. It should be noted that the above asynchronous positioning components are only examples and not limitations to the present application. In actual applications, the number and functions of the asynchronous positioning components can be set based on actual monitoring requirements.

[0041] In some embodiments, the positioning request information is used to request positioning of an abnormal alarm; in some embodiments, the positioning request information includes, but is not limited to, the identification information of the abnormal alarm, the abnormal time, the alarm object, the product line, etc. In some embodiments, when the alarm module determines that there is an abnormal alarm based on the current abnormal detection result, it synchronously sends the positioning request information to the positioning module to request the positioning module to perform abnormal positioning for this abnormal alarm.

[0042] In step S12, the positioning module receives the positioning request information and, according to the positioning request information, performs synchronous positioning operations and asynchronous positioning operations. Among them, the synchronous positioning operation includes traversing and executing the at least one synchronous positioning component to obtain a synchronous positioning result, and providing the synchronous positioning result to the alarm module. The asynchronous positioning operation includes constructing a task sending queue based on the at least one asynchronous positioning component, consuming the task sending queue through an asynchronous thread to execute the at least one asynchronous positioning component to obtain an asynchronous positioning result, and writing the asynchronous positioning result into the database.

[0043] In some embodiments, after an abnormality occurs and before reaching the user, the positioning module will obtain the positioning results of all synchronous positioning components and provide them to the alarm module, and the alarm module will splice the synchronous positioning results into the reach copywriting in a unified format and send them to the user together. In some embodiments, the positioning module executes the at least one asynchronous positioning component based on the kafka message queue. After receiving the positioning request information, the positioning module will construct a kafka task sending queue according to the unique ID of this abnormal alarm. The asynchronous thread of the positioning module will consume kafka to execute the content of each asynchronous positioning component, calculate and process the positioning data to obtain an asynchronous positioning result, and then write the asynchronous positioning result into the database.

[0044] In some embodiments, the execution of the synchronous positioning operation and the asynchronous positioning operation is independent of each other, that is, they do not affect each other; when the positioning module provides the synchronous positioning result to the alarm module, the at least one asynchronous positioning component may have been executed, or may not have been executed yet. That is, the asynchronous positioning result may have been obtained and written into the database, may not have been obtained or written into the database yet, or there may be some asynchronous positioning results that have not been obtained or written into the database yet. In some embodiments, the synchronous positioning operation and the asynchronous positioning operation may start executing simultaneously, or may start executing in a predetermined order (such as starting the asynchronous positioning operation first and then starting the synchronous positioning operation).

[0045] In step S13, the alarm module generates abnormal positioning information according to the synchronous positioning result and sends the abnormal positioning information to the target user, where the abnormal positioning information includes the synchronous positioning result and query indication information corresponding to the abnormal detection result.

[0046] In some embodiments, the alarm module assembles the synchronous positioning result obtained by executing the at least one synchronous positioning component provided by the positioning module in a predetermined format into the reachable copy to generate abnormal positioning information. In some embodiments, the query indication information is used to query the asynchronous positioning result corresponding to the abnormal alarm; in some embodiments, the query indication information includes, but is not limited to: link information corresponding to the asynchronous positioning result, query code, alarm ID, etc.

[0047] In some embodiments, when the alarm module performs the sending operation, the positioning module may have obtained the asynchronous positioning result and written it into the database, or may not have obtained the asynchronous positioning result or not written it into the database yet.

[0048] In some embodiments, the method further includes: receiving an alarm details query request initiated by the target user based on the query indication information, reading the asynchronous positioning result from the database, and sending the asynchronous positioning result to the target user. In some embodiments, the alarm details query request includes identification information corresponding to the abnormal alarm (such as alarm ID, unique query code, etc.). In some embodiments, the positioning module (which may also be executed by other modules on the device) receives an alarm details query request initiated by the target user based on the query indication information in the abnormal positioning information received by the target user, reads the asynchronous positioning result corresponding to the identification information from the database according to the identification information in the alarm details query request, and sends the asynchronous positioning result to the target user. As an example, the abnormal positioning information includes link information corresponding to the asynchronous positioning result. The target user clicks on the link information on the user device, and the user device sends an alarm details query request to the network device (in which the system proposed in this application is arranged) in response to the click operation. The positioning module in the network device receives the alarm details query request, reads the asynchronous positioning result corresponding to the identification information from the database according to the identification information in the alarm details query request, and presents the asynchronous positioning result on the alarm details page of the user device. As another example, the abnormal positioning information includes a unique query code corresponding to the asynchronous positioning result. The target user enters the query code on a specific page of the user device and clicks the query button. The user device sends an alarm details query request to the network device in response to the click operation. The positioning module in the network device receives the alarm details query request, verifies the unique query code in the alarm details query request, reads the asynchronous positioning result corresponding to the identification information from the database after the verification is passed, and presents the asynchronous positioning result on the alarm details page of the user device. In some embodiments, if no corresponding content is found in the database after receiving the alarm details query request (that is, the asynchronous positioning result has not been written into the database), the user can be prompted to query again later.

[0049] Figure 2The following is a schematic flow diagram for anomaly location showing an example of the present application. The specific process is as follows: The Alarmer in the computer device processes the anomaly detection result. When it is determined that the anomaly detection result indicates an anomaly alarm, it synchronously requests the Locator. After receiving the request, the Locator traverses and executes all synchronous location components and returns the synchronous location result. Moreover, the Locator constructs a Kafka task sending queue for all asynchronous location components. The asynchronous thread of the Locator will consume Kafka to execute the content of the asynchronous location components, and then write the execution result (i.e., the asynchronous location result) into the database. The Alarmer generates anomaly location information based on the synchronous location result and sends it to the user (i.e., reaches the user). The anomaly location information includes a location details link (i.e., the link information corresponding to the asynchronous location result). The user can click on the location details link to query the detailed location information (i.e., initiate a query request for alarm details). The Locator can obtain the asynchronous location result from the database according to the alarm ID in the alarm details query request and send it to the alarm details page for display.

[0050] The present application also proposes a system for anomaly location. Among them, the system includes an Alarmer and a Locator. The Locator includes at least one synchronous location component and at least one asynchronous location component. Among them, the Alarmer is used for: if the current anomaly detection result indicates an anomaly alarm, sending a location request message to the Locator, and generating anomaly location information according to the synchronous location result provided by the Locator, and sending the anomaly location information to the target user. Among them, the anomaly location information includes the synchronous location result and query indication information corresponding to the anomaly detection result. Among them, the Locator is used for: receiving the location request message, and according to the location request message, performing synchronous location operations and asynchronous location operations. Among them, the synchronous location operation includes traversing and executing the at least one synchronous location component to obtain the synchronous location result, and providing the synchronous location result to the Alarmer. The asynchronous location operation includes constructing a task sending queue based on the at least one asynchronous location component, consuming the task sending queue through an asynchronous thread to execute the at least one asynchronous location component to obtain an asynchronous location result, and writing the asynchronous location result into the database. In some embodiments, the system is further used for: receiving a query request for alarm details initiated by the target user based on the query indication information, reading the asynchronous location result from the database, and sending the asynchronous location result to the target user. Figure 3 shows a schematic structural diagram of a system for anomaly location showing an example of the present application. The system of this example (such as Figure 3The "abnormality location system" shown includes an alarm module, a location module, and a database. The synchronization location component in the location module includes an exception log information component and the alarm callback component. The asynchronous location component in the location module includes a dimension ratio component, an alarm event association component, an associated online operation component, a multi-root cause analysis component, and a trace information location component. The functions of each module or component in the system proposed in this application have been described in detail in the foregoing embodiments and will not be elaborated herein.

[0051] According to the solution of this application, when an abnormal alarm is detected, it is possible to synchronously request the location module to perform synchronous location operations and asynchronous location operations. Furthermore, it is possible to reach the user with the abnormality location information generated based on the synchronous location results obtained by traversing all synchronous location components. And because the abnormality location information includes query indication information for querying asynchronous location results, the user can query the asynchronous location results based on the query indication information to further understand the alarm details, thus facilitating quick and accurate fault location. This application solves the problem of difficult fault location after system abnormalities in the monitoring field. Through a general system architecture for abnormality location, it is possible to improve the auxiliary location ability of the entire monitoring system.

[0052] This application also provides a computer device. Among them, the computer device includes: a memory for storing one or more programs; one or more processors connected to the memory. When the one or more programs are executed by the one or more processors, the one or more processors execute the method for abnormality location described in this application.

[0053] This application also provides a computer-readable storage medium, on which a computer program is stored. The computer program can be executed by a processor to execute the method for abnormality location described in this application.

[0054] This application also provides a computer program product. When the computer program product is executed by a device, the device executes the method for abnormality location described in this application.

[0055] Figure 4 An exemplary system that can be used to implement the various embodiments described in this application is shown.

[0056] In some embodiments, the system 1000 can serve as any one of the processing devices in the embodiments of this application. In some embodiments, the system 1000 may include one or more computer-readable media having instructions (such as a system memory or the NVM / storage device 1020) and one or more processors (such as (one or more) processors 1005) coupled to the one or more computer-readable media and configured to execute the instructions to implement modules and thereby perform the actions described in this application.

[0057] For one embodiment, the system control module 1010 may include any suitable interface controller to provide any suitable interface to at least one of the processor(s) 1005 and / or any suitable device or component communicating with the system control module 1010.

[0058] The system control module 1010 may include a memory controller module 1030 to provide an interface to the system memory 1015. The memory controller module 1030 may be a hardware module, a software module, and / or a firmware module.

[0059] The system memory 1015 may be used to load and store data and / or instructions for the system 1000, for example. For one embodiment, the system memory 1015 may include any suitable volatile memory, e.g., suitable DRAM. In some embodiments, the system memory 1015 may include double data rate type four synchronous dynamic random access memory (DDR4 SDRAM).

[0060] For one embodiment, the system control module 1010 may include one or more input / output (I / O) controllers to provide an interface to the NVM / storage device 1020 and the communication interface(s) 1025.

[0061] For example, the NVM / storage device 1020 may be used to store data and / or instructions. The NVM / storage device 1020 may include any suitable non-volatile memory (e.g., flash memory) and / or may include any suitable non-volatile storage device(s) (e.g., one or more hard disk drives (HDDs), one or more compact discs (CDs) drives, and / or one or more digital versatile discs (DVDs) drives).

[0062] The NVM / storage device 1020 may include storage resources that are physically part of the device on which the system 1000 is installed, or it may be accessible by the device without being part of the device. For example, the NVM / storage device 1020 may be accessed via the communication interface(s) 1025 over a network.

[0063] (The) communication interface(s) 1025 may provide an interface for the system 1000 to communicate through one or more networks and / or with any other suitable device. The system 1000 may communicate wirelessly with one or more components of a wireless network according to any of one or more wireless network standards and / or protocols.

[0064] For one embodiment, at least one of the (one or more) processors 1005 may be logically encapsulated with one or more controllers of the system control module 1010 (e.g., the memory controller module 1030). For one embodiment, at least one of the (one or more) processors 1005 may be logically encapsulated with one or more controllers of the system control module 1010 to form a system-in-package (SiP). For one embodiment, at least one of the (one or more) processors 1005 may be logically integrated with one or more controllers of the system control module 1010 on the same die. For one embodiment, at least one of the (one or more) processors 1005 may be logically integrated with one or more controllers of the system control module 1010 on the same die to form a system-on-chip (SoC).

[0065] In various embodiments, the system 1000 can be, but is not limited to: a server, a workstation, a desktop computing device, or a mobile computing device (e.g., a laptop computing device, a handheld computing device, a tablet computer, a netbook, etc.). In various embodiments, the system 1000 may have more or fewer components and / or a different architecture. For example, in some embodiments, the system 1000 includes one or more cameras, a keyboard, a liquid crystal display (LCD) screen (including a touchscreen display), a non-volatile memory port, multiple antennas, a graphics chip, an application specific integrated circuit (ASIC), and a speaker.

[0066] It will be apparent to those skilled in the art that the present application is not limited to the details of the above exemplary embodiments, and that the present application can be implemented in other specific forms without departing from the spirit or essential characteristics of the present application. Therefore, in any respect, the embodiments should be regarded as exemplary and non-limiting. The scope of the present application is defined by the appended claims rather than the above description. Accordingly, all changes that fall within the meaning and scope of the equivalent elements of the claims are intended to be embraced within the present application. Any reference signs in the claims should not be construed as limiting the claims involved. In addition, it is obvious that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. The multiple units or devices recited in the system claims can also be implemented by one unit or device through software or hardware. First, second, etc. are used to denote names and do not denote any particular order.

Claims

1. A method for anomaly localization, wherein, The method includes: If the current anomaly detection result indicates an anomaly alarm, the alarm module sends a positioning request message to the positioning module, where the positioning module includes at least one synchronous positioning component and at least one asynchronous positioning component; The positioning module receives the positioning request message and, according to the positioning request message, performs a synchronous positioning operation and an asynchronous positioning operation. The synchronous positioning operation includes traversing and executing the at least one synchronous positioning component to obtain a synchronous positioning result and providing the synchronous positioning result to the alarm module. The asynchronous positioning operation includes constructing a task sending queue based on the at least one asynchronous positioning component, consuming the task sending queue through an asynchronous thread to execute the at least one asynchronous positioning component to obtain an asynchronous positioning result, and writing the asynchronous positioning result into a database; The alarm module generates anomaly positioning information according to the synchronous positioning result and sends the anomaly positioning information to the target user, where the anomaly positioning information includes the synchronous positioning result and query indication information corresponding to the anomaly detection result; Among them, the synchronous positioning component includes an anomaly log information component, and the anomaly log information component is used to obtain anomaly log information corresponding to this anomaly alarm; the asynchronous positioning component includes a dimension ratio component, an alarm event association component, an associated online operation component, and a multi-root cause analysis component. The dimension ratio component is used to analyze the ratio of different dimensions under alarm conditions. The alarm event association component is used to calculate the association relationship between alarm events by analyzing historical alarm data. The associated online operation component is used to judge whether this anomaly alarm is caused by an online operation by analyzing the system operation records during the anomaly time period. The multi-root cause analysis component is used to analyze the ratio of the change of each dimension value under different dimensions to the anomaly and the change difference of each dimension value.

2. The method according to claim 1, wherein, The method further includes: Receiving an alarm detail query request initiated by the target user based on the query indication information, reading the asynchronous positioning result from the database, and sending the asynchronous positioning result to the target user.

3. The method according to claim 1 or 2, wherein The query indication information includes link information corresponding to the asynchronous positioning result.

4. The method according to claim 1, wherein, The synchronous positioning component includes an alarm callback component, and the alarm callback component is used to obtain alarm-related information corresponding to this anomaly alarm.

5. The method according to claim 1, wherein The asynchronous positioning component includes a trace information positioning component, and the trace information positioning component is used to trace the request call chain log and troubleshoot anomalies according to the call chain log.

6. The method according to claim 1, wherein Each asynchronous positioning component or synchronous positioning component is a pluggable module.

7. A system for anomaly localization, wherein, The system includes an alarm module and a positioning module, and the positioning module includes at least one synchronous positioning component and at least one asynchronous positioning component; The alarm module is configured to: if the current anomaly detection result indicates an anomaly alarm, send a positioning request message to the positioning module, and generate anomaly positioning information according to the synchronous positioning result provided by the positioning module, and send the anomaly positioning information to the target user, where the anomaly positioning information includes the synchronous positioning result and query indication information corresponding to the anomaly detection result; The positioning module is configured to: receive the positioning request message, and perform synchronous positioning operations and asynchronous positioning operations according to the positioning request message, where the synchronous positioning operation includes traversing and executing the at least one synchronous positioning component to obtain the synchronous positioning result, and providing the synchronous positioning result to the alarm module, and the asynchronous positioning operation includes constructing a task sending queue based on the at least one asynchronous positioning component, consuming the task sending queue through an asynchronous thread to execute the at least one asynchronous positioning component to obtain an asynchronous positioning result, and writing the asynchronous positioning result into the database; The synchronous positioning component includes an anomaly log information component, and the anomaly log information component is configured to obtain anomaly log information corresponding to the current anomaly alarm; the asynchronous positioning components include a dimension ratio component, an alarm event association component, an associated online operation component, and a multi-root cause analysis component. The dimension ratio component is configured to analyze the ratio of different dimensions under alarm conditions. The alarm event association component is configured to calculate the association relationship between alarm events by analyzing historical alarm data. The associated online operation component is configured to determine whether the current anomaly alarm is caused by an online operation by analyzing the system operation records during the anomaly time period. The multi-root cause analysis component is configured to analyze the ratio of the change of each dimension value in different dimensions to the anomaly and the change difference of each dimension value.

8. The system according to claim 7, wherein The system is further configured to: Receive an alarm details query request initiated by the target user based on the query indication information, read the asynchronous positioning result from the database, and send the asynchronous positioning result to the target user.

9. A computer device, wherein, The computer device includes: A memory for storing one or more programs; One or more processors connected to the memory, When the one or more programs are executed by the one or more processors, the one or more processors are caused to execute the method according to any one of claims 1 to 6.

10. A computer-readable storage medium, on which a computer program is stored, and the computer program can be executed by a processor to perform the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Abnormity positioning method and device for cloud platform monitoring system

    CN108259241A

  • Asynchronous interface detection method, asynchronous interface detection system and readable storage medium

    CN111274137A