Abnormality processing method and device, equipment, medium and program product
By obtaining log information from distributed cluster nodes and utilizing anomaly identification models and knowledge graphs, the difficulty of locating abnormal nodes in distributed architectures is solved, efficient anomaly identification and solution generation are achieved, and operation and maintenance efficiency is improved.
Patent Information
- Application Number
- CN202510851332.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-24
- Publication Date
- 2025-09-23
AI Technical Summary
In a distributed architecture, due to the technical problems of difficulty in locating abnormal nodes and low identification efficiency caused by multi-service coupling, existing technologies are unable to effectively identify complex abnormal data patterns, resulting in difficulty and low efficiency in locating abnormal nodes.
By obtaining the log information of distributed cluster nodes, using the preset anomaly recognition model to preliminarily judge the anomaly probability, combined with the random walk algorithm and knowledge graph, the abnormal nodes can be located and resolved.
It realizes closed-loop intelligent operation and maintenance from log anomaly perception to repair strategy solution generation, improves the accuracy of abnormal node positioning and identification efficiency, and avoids the low efficiency of manual troubleshooting and the pain points of long-term problem solving.
Smart Images

Figure CN120687291A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the fields of artificial intelligence and big data technology, and in particular to an exception handling method, apparatus, device, medium, and program product. Background Art
[0002] With the rapid development of internet technology, the rapid expansion of user numbers, and the continued expansion of business scale, traditional monolithic architectures are facing performance bottlenecks. Distributed architectures, through task decomposition and node communication, enable horizontal scalability to cope with complex business scenarios. However, each service is independently developed and deployed, and the complexity of cross-service call chains can weaken log correlation and make it difficult to locate the root cause of anomalies. Furthermore, the types of log anomalies are numerous and complex. Without experienced engineering and technical personnel, the entire process, from problem occurrence to problem location and resolution, can be time-consuming and have an immeasurable impact on critical production applications.
[0003] In the prior art, a keyword matching method is usually used to periodically match keywords to obtain error information. However, this method is difficult to identify complex data patterns, resulting in high cost and low efficiency in obtaining abnormal data in logs.
[0004] Therefore, in the exception handling process, the difficulty in locating abnormal nodes and the low identification efficiency caused by multi-service coupling are technical problems that need to be solved urgently. Summary of the Invention
[0005] In view of the above problems, the present disclosure provides an abnormality handling method, apparatus, device, medium and program product for improving the accuracy of abnormal node positioning and identification efficiency.
[0006] According to a first aspect of the present disclosure, a method for handling an exception is provided, the method comprising: obtaining first log information from a first node; performing exception recognition through a preset exception recognition model based on the first log information to obtain an recognition result; if the recognition result is abnormal, obtaining call chain data of the first node, the call chain data comprising log information of N nodes, where N is a positive integer; locating an abnormal node among the N nodes based on the log information of the N nodes; and determining an exception solution for the abnormal node.
[0007] According to an embodiment of the present disclosure, the abnormality recognition is performed based on the first log information through a preset abnormality recognition model to obtain an identification result, including: extracting preset parameter items in the first log information to obtain an event vector set; and outputting an abnormality probability value based on the event vector set through the preset abnormality recognition model, wherein, when the abnormality probability value is greater than a first preset threshold value, the identification result is determined to be abnormal.
[0008] According to an embodiment of the present disclosure, after performing abnormality recognition based on the first log information through a preset abnormality recognition model to obtain an identification result, it also includes: updating the preset abnormality recognition model based on the event vector set and the abnormality probability value.
[0009] According to an embodiment of the present disclosure, locating an abnormal node among the N nodes based on the log information of the N nodes includes: calculating anomaly scores of the N nodes based on a random walk algorithm to obtain N anomaly scores; and selecting a node whose median value of the N anomaly scores is greater than a second preset abnormal threshold as the abnormal node.
[0010] According to an embodiment of the present disclosure, determining an abnormal solution for the abnormal node includes: obtaining abnormal information of the abnormal node; and querying the abnormal solution through a preset knowledge graph based on the abnormal information.
[0011] According to an embodiment of the present disclosure, querying the exception solution through a preset knowledge graph based on the exception information includes: directly matching the exception solution in the preset knowledge graph based on the exception information.
[0012] According to an embodiment of the present disclosure, the querying of the exception solution through a preset knowledge graph based on the exception information also includes: in the case where the exception solution in the preset knowledge graph based on the exception information fails to be directly matched, the exception solution in the preset knowledge graph is matched based on the exception information by similarity.
[0013] According to an embodiment of the present disclosure, after determining the abnormal solution for the abnormal node, the method further includes: receiving feedback information for the abnormal solution; and maintaining the knowledge graph based on the feedback information.
[0014] A second aspect of the present disclosure provides an exception handling device, including: an acquisition module for acquiring first log information from a first node; an exception identification module for performing exception identification based on the first log information through a preset exception identification model to obtain an identification result; the acquisition module is also used to acquire call chain data of the first node when the identification result is abnormal, the call chain data including log information of N nodes, where N is a positive integer; an exception locating module for locating an abnormal node among the N nodes based on the log information of the N nodes; and a solution module for determining an exception solution for the abnormal node.
[0015] According to an embodiment of the present disclosure, the anomaly identification module is specifically used to extract preset parameter items in the first log information to obtain an event vector set; and output an anomaly probability value based on the event vector set through the preset anomaly identification model, wherein, when the anomaly probability value is greater than a first preset threshold, the identification result is determined to be abnormal.
[0016] According to an embodiment of the present disclosure, the apparatus further includes: a model updating module, configured to update the preset anomaly recognition model based on the event vector set and the anomaly probability value.
[0017] According to an embodiment of the present disclosure, the anomaly location module is specifically used to calculate the anomaly scores of the N nodes based on a random walk algorithm to obtain N anomaly scores; and select a node whose median of the N anomaly scores is greater than a second preset abnormal threshold as the anomaly node.
[0018] According to an embodiment of the present disclosure, the solution-solving module is specifically configured to obtain abnormal information of the abnormal node; and query the abnormal solution through a preset knowledge graph based on the abnormal information.
[0019] According to an embodiment of the present disclosure, the solution-solving module is specifically configured to directly match an exception solution in the designed knowledge graph based on the exception information.
[0020] According to an embodiment of the present disclosure, the solution-solving module is specifically used to match the exception solution in the set knowledge graph by similarity based on the exception information when the exception solution in the set knowledge graph directly matched based on the exception information fails.
[0021] According to an embodiment of the present disclosure, the device further includes a knowledge graph updating module for receiving feedback information on the exception solution; and maintaining the knowledge graph based on the feedback information.
[0022] A third aspect of the present disclosure provides an electronic device, comprising: one or more processors; and a memory for storing one or more computer programs, wherein the one or more processors execute the one or more computer programs to implement the steps of the above method.
[0023] The fourth aspect of the present disclosure further provides a computer-readable storage medium having a computer program or instructions stored thereon, which implements the steps of the above method when the computer program or instructions are executed by a processor.
[0024] The fifth aspect of the present disclosure further provides a computer program product, comprising a computer program or instructions, which implement the steps of the above method when executed by a processor.
[0025] In order to solve the technical problems of difficulty in locating abnormal nodes and low identification efficiency caused by multi-service coupling. In an embodiment of the present disclosure, by obtaining the log information of the nodes in the distributed cluster, the log information is identified, and the abnormal probability of the call chain in which the node is located is preliminarily judged. When the call chain of the node is in an abnormal state, the log information of other nodes in the call chain is obtained, and the abnormal node is finally determined. When the abnormal node is determined, the abnormal solution for the abnormal node is determined. The embodiment of the present disclosure can at least achieve the following beneficial effects: through multiple rounds of abnormality identification, closed-loop intelligent operation and maintenance from log anomaly perception to repair strategy solution generation is realized. Compared with traditional rule-based solutions, it has solved the core capabilities of log pattern drift in dynamic environments and difficulty in locating multi-service coupling faults. At the same time, it avoids the pain points of low efficiency of manual log checking and long time-consuming log anomaly resolution. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] The above contents and other objects, features and advantages of the present disclosure will become more apparent through the following description of the embodiments of the present disclosure with reference to the accompanying drawings, in which:
[0027] Figure 1 Schematically illustrates an application scenario diagram of the exception handling method, apparatus, device, medium, and program product according to an embodiment of the present disclosure;
[0028] Figure 2 The following schematically shows a flow chart of an exception handling method according to an embodiment of the present disclosure;
[0029] Figure 3 Schematically shows a structural block diagram of an exception handling device according to an embodiment of the present disclosure; and
[0030] Figure 4 A block diagram of an electronic device suitable for implementing the exception handling method according to an embodiment of the present disclosure is schematically shown. DETAILED DESCRIPTION
[0031] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present disclosure. In the detailed description below, for ease of explanation, many specific details are set forth to provide a comprehensive understanding of the embodiments of the present disclosure. However, it is apparent that one or more embodiments may also be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessary confusion of the concepts of the present disclosure.
[0032] The terms used herein are only for describing specific embodiments and are not intended to limit the present disclosure. The terms "comprise," "include," etc. used herein indicate the presence of the features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0033] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.
[0034] When expressions such as "at least one of A, B, and C, etc." are used, they should generally be interpreted in accordance with the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include but is not limited to a system having A alone, B alone, C alone, A and B, A and C, B and C, and / or A, B, C, etc.).
[0035] An embodiment of the present disclosure provides an exception handling method, which includes: obtaining first log information from a first node; performing exception recognition through a preset exception recognition model based on the first log information to obtain an recognition result; if the recognition result is abnormal, obtaining call chain data of the first node, the call chain data including log information of N nodes, where N is a positive integer; locating an abnormal node among the N nodes based on the log information of the N nodes; and determining an exception solution for the abnormal node.
[0036] In order to solve the technical problems of difficulty in locating abnormal nodes and low identification efficiency caused by multi-service coupling. In an embodiment of the present disclosure, by obtaining the log information of the nodes in the distributed cluster, the log information is identified, and the abnormal probability of the call chain in which the node is located is preliminarily judged. When the call chain of the node is in an abnormal state, the log information of other nodes in the call chain is obtained, and the abnormal node is finally determined. When the abnormal node is determined, the abnormal solution for the abnormal node is determined. The embodiment of the present disclosure can at least achieve the following beneficial effects: through multiple rounds of abnormality identification, closed-loop intelligent operation and maintenance from log anomaly perception to repair strategy solution generation is realized. Compared with traditional rule-based solutions, it has solved the core capabilities of log pattern drift in dynamic environments and difficulty in locating multi-service coupling faults. At the same time, it avoids the pain points of low efficiency of manual log checking and long time-consuming log anomaly resolution.
[0037] Figure 1 The following schematically illustrates an application scenario of the exception handling method according to an embodiment of the present disclosure.
[0038] like Figure 1 As shown, the application scenario 100 according to this embodiment may include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 is used as a medium for providing a communication link between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired or wireless communication links or optical fiber cables.
[0039] A user may use a first terminal device 101, a second terminal device 102, or a third terminal device 103 to interact with a server 105 via a network 104 to receive or send messages, etc. Various communication client applications may be installed on the first terminal device 101, the second terminal device 102, or the third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (for example only).
[0040] The first terminal device 101 , the second terminal device 102 , and the third terminal device 103 may be various electronic devices having display screens and supporting web browsing, including but not limited to smart phones, tablet computers, laptop computers, desktop computers, and the like.
[0041] The server 105 may be a server that provides various services, such as a background management server (for example only) that supports websites browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103. The background management server may analyze and process received data such as user requests, and feed back processing results (e.g., web pages, information, or data obtained or generated based on user requests) to the terminal devices.
[0042] It should be noted that the exception handling method provided in the embodiment of the present disclosure can generally be executed by the server 105. Accordingly, the exception handling device provided in the embodiment of the present disclosure can generally be set in the server 105. The exception handling method provided in the embodiment of the present disclosure can also be executed by a server or server cluster that is different from the server 105 and can communicate with the first terminal device 101, the second terminal device 102, the third terminal device 103 and / or the server 105. Accordingly, the exception handling device provided in the embodiment of the present disclosure can also be set in a server or server cluster that is different from the server 105 and can communicate with the first terminal device 101, the second terminal device 102, the third terminal device 103 and / or the server 105.
[0043] It should be understood that Figure 1The number of terminal devices, networks and servers in the embodiment is merely illustrative. Any number of terminal devices, networks and servers may be provided as required.
[0044] The following will be based on Figure 1 The scene described by Figure 2 The exception handling method of the disclosed embodiment is described in detail.
[0045] Figure 2 The flowchart of the exception handling method according to the embodiment of the present disclosure is schematically shown.
[0046] like Figure 2 As shown, the exception handling method of this embodiment includes operations S210 to S250, and the exception handling method can be executed by the server 105.
[0047] In operation S210 , first log information is obtained from a first node.
[0048] Specifically, the first node refers to a node in a distributed system, which can be in different call chains. For example, you can obtain log files from the log center for the node you want to test, filter out invalid data such as test environment logs and heartbeat detection noise, and unify the timestamp format.
[0049] In operation S220 , anomaly recognition is performed based on the first log information using a preset anomaly recognition model to obtain an identification result.
[0050] Specifically, the original first log information is converted into an event ID sequence with a timestamp, and parameters in the log (such as error code, response time, and keywords) are extracted to form a multidimensional vector as input data for the preset anomaly recognition model.
[0051] The preset anomaly recognition model can be a supervised learning model, trained using pre-set labeled features and various feature items in the logs. For example, a pre-trained model based on historical logs can be used to build a basic anomaly pattern library. Then, 10% of new logs are retained as a validation set, and the false alarm rate is monitored. If the validation loss continues to rise, training is paused and rolled back to the previous stable version using an adaptive learning rate.
[0052] According to an embodiment of the present disclosure, the abnormality recognition is performed based on the first log information through a preset abnormality recognition model to obtain an identification result, including: extracting preset parameter items in the first log information to obtain an event vector set; and outputting an abnormality probability value based on the event vector set through the preset abnormality recognition model, wherein, when the abnormality probability value is greater than a first preset threshold value, the identification result is determined to be abnormal.
[0053] Specifically, the preset anomaly recognition model can be trained to output an anomaly probability value for log information. By inputting preset parameters (such as error code, response time, and keywords), it outputs an anomaly probability value (such as an 80% anomaly probability value). It then determines whether the anomaly probability value is greater than a first preset threshold (such as 75%). If it is greater, it is determined that there is an anomaly in the call chain where the node is located; if it is less than this, it is determined that there is no anomaly at the node. A certain scope of investigation is preliminarily determined.
[0054] In a typical scenario, the pre-set anomaly recognition model can use multiple machine learning models in a serial or parallel prediction method to predict anomaly probabilities. For example, the anomaly recognition model includes parallel Model 1 and Model 2. Model 1 uses the event ID sequence in the input log information to predict the first probability distribution of the next event ID sequence. Model 2 uses the input parameter vector and the dynamic parameters of the log template (response time, error code, keyword) to output a second probability distribution. Finally, a weighted sum is taken from the first and second probability distributions to calculate a comprehensive anomaly probability value. The value is then compared with a threshold to determine whether it is less than the anomaly threshold.
[0055] According to an embodiment of the present disclosure, after performing abnormality recognition based on the first log information through a preset abnormality recognition model to obtain an identification result, it also includes: updating the preset abnormality recognition model based on the event vector set and the abnormality probability value.
[0056] Specifically, incremental learning can be achieved through incremental data. After each or every few times the anomaly recognition model is used, incremental learning is performed through the data extracted from the log and combined with the predicted anomaly probability value to update the anomaly recognition model.
[0057] For example, through online incremental learning, each time a log from a new time window is received, model fine-tuning is triggered (the learning rate is reduced to 10%-20% of the original value); feedback drives updates, and if the operation and maintenance marks a false positive or missed sample, it is added to the training set and re-optimized.
[0058] In operation S230, when the identification result is abnormal, the call chain data of the first node is obtained, and the call chain data includes log information of N nodes, where N is a positive integer.
[0059] Specifically, when the recognition result is determined to be abnormal, the log information of other nodes and other nodes that have a call chain relationship with the first node is obtained. It should be noted that the first node can be in the heavy weight of multiple different types of tasks. Therefore, the call chain data of the first node can include data of a single call chain or data of multiple call chains, which will not be elaborated here.
[0060] In operation S240 , an abnormal node among the N nodes is located based on the log information of the N nodes.
[0061] Specifically, after determining N nodes that have a call relationship with the first node, the log information is used to specifically locate whether these nodes have anomalies. For example, the error information appearing in the log information is matched with a preset error item to determine the abnormal node.
[0062] According to an embodiment of the present disclosure, locating an abnormal node among the N nodes based on the log information of the N nodes includes: calculating anomaly scores of the N nodes based on a random walk algorithm to obtain N anomaly scores; and selecting a node whose median value of the N anomaly scores is greater than a second preset abnormal threshold as the abnormal node.
[0063] Specifically, the log information can be used to output the task execution status of the node itself and the data on the correlation between nodes. By combining these data, a graph can be established, which contains nodes and edges. Then, the random walk algorithm is used to calculate the N nodes on the graph, calculate the abnormality score of each node, and then select one or more abnormal nodes with a score greater than the second preset threshold. It can be understood that
[0064] In a typical scenario, within a distributed microservices architecture, service call relationships may fluctuate with business traffic, and the launch of new features may result in the addition of new dependent services. Therefore, the knowledge graph weights must be constantly updated, assigning higher weights to edges with high dependency strengths. This accelerates root cause identification: in the graph traversal algorithm, searches are prioritized along high-weight edges. For example, if a payment service anomaly occurs, the system will prioritize checking its upstream order service over the low-dependency inventory service. Dynamic weights are adjusted based on the following rules: 1) Service dependency strength: Remote Procedure Call (RPC) call frequency calculated from call chain statistics. 2) Fault propagation probability: Number of historical faults divided by the total number of faults.
[0065] In operation S250 , an abnormality solution for the abnormal node is determined.
[0066] Specifically, the abnormality solution is determined through the log of the abnormal node. For example, the error-related information in the log of the abnormal node can be matched with some preset case libraries to achieve the output of the solution.
[0067] According to an embodiment of the present disclosure, determining an abnormal solution for the abnormal node includes: obtaining abnormal information of the abnormal node; and querying the abnormal solution through a preset knowledge graph based on the abnormal information.
[0068] Specifically, the exception information is used to query the exception solution in the preset knowledge graph, which is established by historical data.
[0069] In a typical scenario, building a knowledge graph begins with acquiring data sources, including: log metadata (service name, module, and error code); system topology (microservice dependencies and container deployment relationships); operations and maintenance knowledge base (historical troubleshooting and solution manuals); and monitoring metrics (CPU, memory, and network time series data). This data allows the creation of entity types and relationship definitions for the knowledge graph. Entity types include microservices, virtual machines, log event templates, error codes, code owners, historical event tickets, and monitoring metrics. Relationship definitions include service dependencies, deployment hosts, error code-triggered log events, code owners, and corresponding solutions.
[0070] According to an embodiment of the present disclosure, querying the exception solution through a preset knowledge graph based on the exception information includes:
[0071] Directly match the exception solution in the knowledge graph of the design based on the exception information.
[0072] Specifically, the knowledge graph can be matched directly by keywords without the need for fuzzy matching.
[0073] According to an embodiment of the present disclosure, querying the exception solution through a preset knowledge graph based on the exception information further includes:
[0074] In the event that direct matching of the exception solution in the assumed knowledge graph based on the exception information fails, the exception solution in the assumed knowledge graph is matched by similarity based on the exception information.
[0075] Specifically, when direct matching of anomaly solutions fails, fuzzy matching can also be performed.
[0076] In a typical scenario, if multiple recommended solutions are found based on the knowledge graph, they can be sorted according to the following strategy:
[0077] 1) Effectiveness score: The success rate of resolving the fault in historical events.
[0078] 2) Implementation complexity: the complexity of the required operations (number of operation steps, number of configuration environment modifications).
[0079] 3) Impact on production: expected downtime and scope of impact.
[0080] According to an embodiment of the present disclosure, after determining the abnormal solution for the abnormal node, the method further includes: receiving feedback information for the abnormal solution; and maintaining the knowledge graph based on the feedback information.
[0081] Feedback information is manually specified and includes information about whether a solution is effective or ineffective. If a solution is effective, its priority in subsequent recommendations is increased; if it is ineffective, its priority is decreased. Understandably, in a distributed architecture, changes in the system environment, the emergence of new faults, and the optimization of solutions all require the knowledge graph to reflect the latest information in a timely manner. Otherwise, root cause identification and recommended solutions will be inaccurate. Dynamic updates ensure that dependencies are correct, avoid missing critical paths, reduce false positives, and improve diagnostic accuracy.
[0082] In order to solve the technical problems of difficulty in locating abnormal nodes and low identification efficiency caused by multi-service coupling. In an embodiment of the present disclosure, by obtaining the log information of the nodes in the distributed cluster, the log information is identified, and the abnormal probability of the call chain in which the node is located is preliminarily judged. When the call chain of the node is in an abnormal state, the log information of other nodes in the call chain is obtained, and the abnormal node is finally determined. When the abnormal node is determined, the abnormal solution for the abnormal node is determined. The embodiment of the present disclosure can at least achieve the following beneficial effects: through multiple rounds of abnormality identification, closed-loop intelligent operation and maintenance from log anomaly perception to repair strategy solution generation is realized. Compared with traditional rule-based solutions, it has solved the core capabilities of log pattern drift in dynamic environments and difficulty in locating multi-service coupling faults. At the same time, it avoids the pain points of low efficiency of manual log checking and long time-consuming log anomaly resolution.
[0083] Based on the above exception handling method, the present disclosure also provides an exception handling device. Figure 3 The device is described in detail.
[0084] Figure 3 The following schematically shows a structural block diagram of an exception handling device according to an embodiment of the present disclosure.
[0085] like Figure 3 As shown, the exception handling device 300 of this embodiment includes an acquisition module 310 , an exception identification module 320 , an exception location module 330 and a solution module 340 .
[0086] The acquisition module 310 is used to acquire the first log information from the first node. In one embodiment, the acquisition module 310 can be used to perform the operation S210 described above, which will not be described in detail here.
[0087] The abnormality identification module 320 is used to perform abnormality identification based on the first log information through a preset abnormality identification model to obtain an identification result. In one embodiment, the abnormality identification module 320 can be used to perform the operation S220 described above, which will not be repeated here.
[0088] The acquisition module 310 is further configured to, when the identification result is abnormal, acquire the call chain data of the first node, wherein the call chain data includes log information of N nodes, where N is a positive integer. In one embodiment, the acquisition module 310 can be configured to execute the operation S230 described above, which will not be described in detail here.
[0089] The abnormality locating module 330 is used to locate abnormal nodes among the N nodes based on the log information of the N nodes. In one embodiment, the abnormality locating module 330 can be used to perform the operation S240 described above, which will not be repeated here.
[0090] The solution-solving module 340 is used to determine a solution to the abnormal node. In one embodiment, the solution-solving module 340 can be used to perform the operation S250 described above, which will not be described in detail here.
[0091] In order to solve the technical problems of difficulty in locating abnormal nodes and low identification efficiency caused by multi-service coupling. In an embodiment of the present disclosure, by obtaining the log information of the nodes in the distributed cluster, the log information is identified, and the abnormal probability of the call chain in which the node is located is preliminarily judged. When the call chain of the node is in an abnormal state, the log information of other nodes in the call chain is obtained, and the abnormal node is finally determined. When the abnormal node is determined, the abnormal solution for the abnormal node is determined. The embodiment of the present disclosure can at least achieve the following beneficial effects: through multiple rounds of abnormality identification, closed-loop intelligent operation and maintenance from log anomaly perception to repair strategy solution generation is realized. Compared with traditional rule-based solutions, it has solved the core capabilities of log pattern drift in dynamic environments and difficulty in locating multi-service coupling faults. At the same time, it avoids the pain points of low efficiency of manual log checking and long time-consuming log anomaly resolution.
[0092] According to an embodiment of the present disclosure, the anomaly identification module is specifically used to extract preset parameter items in the first log information to obtain an event vector set; and output an anomaly probability value based on the event vector set through the preset anomaly identification model, wherein, when the anomaly probability value is greater than a first preset threshold, the identification result is determined to be abnormal.
[0093] According to an embodiment of the present disclosure, the apparatus further includes: a model updating module, configured to update the preset anomaly recognition model based on the event vector set and the anomaly probability value.
[0094] According to an embodiment of the present disclosure, the anomaly location module is specifically used to calculate the anomaly scores of the N nodes based on a random walk algorithm to obtain N anomaly scores; and select a node whose median of the N anomaly scores is greater than a second preset abnormal threshold as the anomaly node.
[0095] According to an embodiment of the present disclosure, the solution-solving module is specifically configured to obtain abnormal information of the abnormal node; and query the abnormal solution through a preset knowledge graph based on the abnormal information.
[0096] According to an embodiment of the present disclosure, the solution-solving module is specifically configured to directly match an exception solution in the designed knowledge graph based on the exception information.
[0097] According to an embodiment of the present disclosure, the solution-solving module is specifically used to match the exception solution in the set knowledge graph by similarity based on the exception information when the exception solution in the set knowledge graph directly matched based on the exception information fails.
[0098] According to an embodiment of the present disclosure, the device further includes a knowledge graph updating module for receiving feedback information on the exception solution; and maintaining the knowledge graph based on the feedback information.
[0099] According to an embodiment of the present disclosure, any multiple modules among 310, the anomaly identification module 320, the anomaly location module 330, and the solution module 340 can be combined into a single module, or any one of these modules can be split into multiple modules. Alternatively, at least part of the functionality of one or more of these modules can be combined with at least part of the functionality of other modules and implemented in a single module. According to an embodiment of the present disclosure, at least one of 310, the anomaly identification module 320, the anomaly location module 330, and the solution module 340 can be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on a chip, a system on a substrate, a system on a package, an application-specific integrated circuit (ASIC), or can be implemented in hardware or firmware through any other reasonable means of circuit integration or packaging, or can be implemented in any one of the three implementation methods of software, hardware, and firmware, or in any appropriate combination of any of these. Alternatively, at least one of 310 , the anomaly identification module 320 , the anomaly location module 330 and the solution-solving module 340 may be at least partially implemented as a computer program module, which may perform corresponding functions when executed.
[0100] Figure 4 A block diagram of an electronic device suitable for implementing the exception handling method according to an embodiment of the present disclosure is schematically shown.
[0101] like Figure 4 As shown, the electronic device 900 according to an embodiment of the present disclosure includes a processor 901, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 902 or a program loaded from a storage unit 908 into a random access memory (RAM) 903. The processor 901 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or a related chipset and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 901 may also include onboard memory for caching purposes. The processor 901 may include a single processing unit or multiple processing units for performing different actions of the method flow according to the embodiment of the present disclosure.
[0102] Various programs and data required for the operation of the electronic device 900 are stored in the RAM 903. The processor 901, the ROM 902, and the RAM 903 are connected to each other via a bus 904. The processor 901 executes the various operations of the method flow according to the embodiment of the present disclosure by executing the programs in the ROM 902 and / or the RAM 903. It should be noted that the programs may also be stored in one or more memories other than the ROM 902 and the RAM 903. The processor 901 may also execute the various operations of the method flow according to the embodiment of the present disclosure by executing the programs stored in the one or more memories.
[0103] According to an embodiment of the present disclosure, electronic device 900 may further include an input / output (I / O) interface 905, which is also connected to bus 904. Electronic device 900 may also include one or more of the following components connected to I / O interface 905: an input section 906 including a keyboard, mouse, etc.; an output section 907 including devices such as a cathode ray tube (CRT), liquid crystal display (LCD), and speakers; a storage section 908 including a hard disk; and a communication section 909 including a network interface card such as a LAN card or modem. Communication section 909 performs communication processing via a network such as the Internet. A drive 910 is also connected to I / O interface 905 as needed. Removable media 911, such as a magnetic disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed in drive 910 as needed, so that computer programs read from the removable media can be installed into storage section 908 as needed.
[0104] The present disclosure also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments, or may exist independently and not be incorporated into the device / apparatus / system. The computer-readable storage medium carries one or more programs, and when executed, implements the method according to the embodiments of the present disclosure.
[0105] According to an embodiment of the present disclosure, a computer-readable storage medium may be a non-volatile computer-readable storage medium, and may include, for example, but is not limited to: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present disclosure, a computer-readable storage medium may include the ROM 902 and / or RAM 903 described above, and / or one or more memories other than ROM 902 and RAM 903.
[0106] The embodiments of the present disclosure also include a computer program product, which includes a computer program containing program code for executing the method shown in the flowchart. When the computer program product is run in a computer system, the program code is used to enable the computer system to implement the method provided by the embodiments of the present disclosure.
[0107] The computer program executes the above functions defined in the system / device of the embodiment of the present disclosure when the processor 901 executes the computer program. According to the embodiment of the present disclosure, the system, device, module, unit, etc. described above can be implemented by a computer program module.
[0108] In one embodiment, the computer program may be stored on a tangible storage medium such as an optical storage device or a magnetic storage device. In another embodiment, the computer program may be transmitted and distributed in the form of a signal on a network medium, downloaded and installed via the communication portion 909, and / or installed from a removable medium 911. The program code contained in the computer program may be transmitted using any appropriate network medium, including but not limited to wireless, wired, or any suitable combination thereof.
[0109] In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 909, and / or installed from a removable medium 911. When the computer program is executed by the processor 901, the above-described functions defined in the system of the embodiment of the present disclosure are performed. According to the embodiment of the present disclosure, the systems, devices, means, modules, units, etc. described above can be implemented by computer program modules.
[0110] According to an embodiment of the present disclosure, the program code for executing the computer program provided by the embodiment of the present disclosure can be written in any combination of one or more programming languages. Specifically, these computer programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include, but are not limited to, languages such as Java, C++, Python, "C" or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving a remote computing device, the remote computing device can be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (for example, using an Internet service provider to connect via the Internet).
[0111] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the above-mentioned module, program segment, or a part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, and the combination of boxes in the block diagram or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0112] Those skilled in the art will appreciate that the features described in the various embodiments of the present disclosure may be combined and / or coupled in various ways, even if such combinations or couplings are not explicitly described in the present disclosure. In particular, the features described in the various embodiments of the present disclosure may be combined and / or coupled in various ways without departing from the spirit and teachings of the present disclosure. All such combinations and / or couplings fall within the scope of the present disclosure.
[0113] The above describes the embodiments of the present disclosure. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present disclosure. Although each embodiment has been described separately above, this does not mean that the measures in each embodiment cannot be advantageously used in combination. Without departing from the scope of the present disclosure, those skilled in the art may make various substitutions and modifications, which should all fall within the scope of the present disclosure.
Claims
1. An exception handling method, characterized in that: The method comprises: Obtaining first log information from a first node; Performing anomaly recognition based on the first log information using a preset anomaly recognition model to obtain an identification result; If the identification result is abnormal, obtain call chain data of the first node, where the call chain data includes log information of N nodes, where N is a positive integer; Locating abnormal nodes among the N nodes based on the log information of the N nodes; and Determine an abnormality solution for the abnormal node.
2. The method according to claim 1, wherein performing anomaly identification based on the first log information using a preset anomaly identification model to obtain an identification result comprises: Extracting preset parameter items from the first log information to obtain an event vector set; as well as Based on the event vector set, the preset abnormality recognition model is used to output an abnormality probability value, wherein when the abnormality probability value is greater than a first preset threshold, the recognition result is determined to be abnormal.
3. The method according to claim 2, after performing abnormality identification based on the first log information using a preset abnormality identification model to obtain an identification result, further comprising: The preset anomaly recognition model is updated based on the event vector set and the anomaly probability value.
4. The method according to claim 1, wherein locating an abnormal node among the N nodes based on log information of the N nodes comprises: Calculating anomaly scores of the N nodes based on a random walk algorithm to obtain N anomaly scores; as well as Nodes whose median values of the N abnormality scores are greater than a second preset abnormality threshold are selected as the abnormal nodes.
5. The method according to claim 1, wherein determining an abnormality solution for the abnormal node comprises: Obtaining abnormal information of the abnormal node; as well as Based on the exception information, the exception solution is queried through a preset knowledge graph.
6. The method according to claim 5, wherein querying the exception solution through a preset knowledge graph based on the exception information comprises: Directly match the exception solution in the knowledge graph of the design based on the exception information.
7. The method according to claim 6, wherein querying the exception solution through a preset knowledge graph based on the exception information further comprises: In the event that direct matching of the exception solution in the assumed knowledge graph based on the exception information fails, the exception solution in the assumed knowledge graph is matched by similarity based on the exception information.
8. The method according to any one of claims 5 to 7, further comprising, after determining the abnormality solution for the abnormal node: receiving feedback information on the abnormality solution; as well as Maintain the knowledge graph based on the feedback information.
9. An exception handling device, characterized in that: The device comprises: An acquisition module, configured to acquire first log information from a first node; an anomaly identification module, configured to perform anomaly identification based on the first log information using a preset anomaly identification model to obtain an identification result; The acquisition module is further configured to, when the recognition result is abnormal, acquire call chain data of the first node, the call chain data including log information of N nodes, where N is a positive integer; an abnormality locating module, configured to locate an abnormal node among the N nodes based on log information of the N nodes; and A solution-solving module is used to determine an abnormal solution for the abnormal node.
10. An electronic device comprising: one or more processors; a memory for storing one or more computer programs, It is characterized in that the one or more processors execute the one or more computer programs to implement the steps of the method according to any one of claims 1 to 8.
11. A computer-readable storage medium having a computer program or instruction stored thereon, characterized in that: When the computer program or instruction is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.
12. A computer program product comprising a computer program or instructions, characterized in that When the computer program or instruction is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.