Root cause positioning method, device, equipment, storage medium and program product

CN122817062APending Publication Date: 2026-09-25INDUSTRIAL AND COMMERCIAL BANK OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611043176.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-14
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

但这种仅依赖单一类型数据的根因定位方法,缺乏对错误信息文本、代码调用栈和运行时环境参数进行综合分析的机制,导致定位结果缺乏全面性

Benefits of technology

[0013]根据本申请的第三方面提供了一种电子设备,包括:一个或多个处理器;存储器,用于存储一个或多个计算机程序,其中,上述一个或多个处理器执行上述一个或多个计算机程序以实现上述方法的步骤。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122817062A_ABST
    Figure CN122817062A_ABST
Patent Text Reader

Abstract

The application provides a root cause positioning method and device, equipment, storage medium and program product, which can be applied to the fields of artificial intelligence technology and big data technology. The method comprises the following steps: obtaining error information text of a target program, a code call stack of the target program and an environment parameter of the target program. Based on the error information text of the target program and the code call stack of the target program, root cause scores of n calling codes of the target program are generated. Based on the environment parameter of the target program, the n calling codes are subjected to environment parameter consistency verification. The root cause scores of m calling code lines that pass the environment parameter consistency verification are obtained, the root cause scores of the m calling code lines are sorted, and the calling code with the highest root cause score is determined as the target code of root cause positioning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of big data technology, specifically to the field of artificial intelligence technology, and in particular to a root cause localization method, apparatus, device, storage medium, and program product. Background Technology

[0002] With the development of computer technology, software program malfunctions or crashes are common and unavoidable problems for enterprises, including banks, during software program development, testing, and maintenance. When a program malfunctions or crashes, it is necessary to quickly and accurately locate the root cause of the error in order to fix the defect in a timely manner and ensure the stability and reliability of the system.

[0003] Current technologies typically rely on single-type log analysis, static code analysis, abstract syntax tree analysis, control flow graph analysis, or data flow analysis to locate the root cause of errors. However, this root cause localization method, which depends on only a single type of data, lacks a mechanism for comprehensive analysis of error message text, code call stack, and runtime environment parameters, resulting in incomplete localization results. Furthermore, the lack of structured verification methods for environmental factors leads to inaccurate localization results in environment-sensitive error scenarios. It also makes it impossible to reasonably filter the localization results, potentially listing code that is completely unreachable or lacks the conditions for execution as the root cause, easily resulting in poor accuracy and low reliability in root cause localization. Summary of the Invention

[0004] In view of the above problems, this application provides a root cause localization method, apparatus, device, storage medium, and program product.

[0005] According to a first aspect of this application, a root cause localization method is provided, the method comprising: obtaining error message text of a target program, the code call stack of the target program, and environment parameters of the target program; generating root cause scores for n call codes of the target program based on the error message text of the target program and the code call stack of the target program, wherein n is an integer and n is greater than 0; performing environment parameter consistency verification on the n call codes based on the environment parameters of the target program; and obtaining root cause scores for m call code lines that pass the environment parameter consistency verification, sorting the root cause scores of the m call code lines, and determining the call code with the highest root cause score as the target code for root cause localization, wherein m is an integer, m is greater than 0 and m is less than n.

[0006] According to an embodiment of this application, based on the error message text of the target program and the code call stack of the target program, a root cause score for n call codes of the target program is generated, including: extracting features from the error message text of the target program to generate a semantic vector of the error message text of the target program; extracting features from the code call stack of the target program to generate a temporal feature vector of the call stack of the target program; performing vector fusion on the semantic vector of the error message text of the target program and the temporal feature vector of the call stack of the target program to generate a fused feature vector; and normalizing the fused feature vector to generate a root cause score for the n call codes of the target program.

[0007] According to an embodiment of this application, feature extraction is performed on the error message text of the target program to generate a semantic vector of the error message text of the target program, including: performing data cleaning and data standardization on the error message text of the target program to generate standardized error message text; and inputting the standardized error message text into a semantic analysis model that is pre-trained and encoded at the embedding layer of the model, and outputting the semantic vector of the error message text.

[0008] According to an embodiment of this application, feature extraction is performed on the code call stack of the target program to generate a call stack temporal feature vector of the target program, including: performing data cleaning and data standardization on the code call stack of the target program to generate a standardized code call stack of the target program; inputting the standardized code call stack of the target program into a pre-trained bidirectional long short-term memory network model to output a forward temporal feature vector and a backward temporal feature vector of the code call stack of the target program; and concatenating the forward temporal feature vector and the backward temporal feature vector of the code call stack of the target program to generate a call stack temporal feature vector of the target program.

[0009] According to an embodiment of this application, vector fusion is performed on the error message text semantic vector of the target program and the call stack timing feature vector of the target program to generate a fused feature vector. This includes: inputting the error message text semantic vector of the target program into a pre-trained neural network model fine-tuned according to a text attention mechanism, and outputting a first weight of the error message text semantic vector of the target program; inputting the call stack timing feature vector of the target program into a pre-trained graph neural network model fine-tuned according to a text attention mechanism, and outputting a second weight of the call stack timing feature vector of the target program; and performing vector fusion on the error message text semantic vector of the target program and the call stack timing feature vector of the target program based on the first weight and the second weight of the call stack timing feature vector of the target program to generate the fused feature vector.

[0010] According to an embodiment of this application, based on the environment parameters of the target program, the consistency verification of the environment parameters of the n calling codes includes: performing data cleaning and data standardization processing on the environment parameters of the target program to generate standardized environment parameters of the target program; vectorizing the standardized environment parameters of the target program to generate a unified-dimensional environment parameter feature vector of the target program; and performing environment parameter consistency verification on the n calling codes based on the unified-dimensional environment parameter feature vector of the target program.

[0011] According to an embodiment of this application, the standardized environmental parameters of the target program are vectorized to generate a unified-dimensional environmental parameter feature vector of the target program, including: vectorizing the standardized environmental parameters of the target program using numerical encoding or categorical encoding to generate a unified-dimensional environmental parameter feature vector of the target program.

[0012] According to a second aspect of this application, a root cause localization apparatus is provided, comprising: a first acquisition module, configured to acquire error message text of a target program, the code call stack of the target program, and environment parameters of the target program; a first generation module, configured to generate root cause scores for n call codes of the target program based on the error message text of the target program and the code call stack of the target program, wherein n is an integer and n is greater than 0; a first verification module, configured to perform environment parameter consistency verification on the n call codes based on the environment parameters of the target program; and a first determination module, configured to acquire root cause scores for m call code lines that have passed the environment parameter consistency verification, sort the root cause scores of the m call code lines, and determine the call code with the highest root cause score as the target code for root cause localization, wherein m is an integer, m is greater than 0 and m is less than n.

[0013] According to a third aspect of this application, an electronic device is provided, comprising: one or more processors; and a memory for storing one or more computer programs, wherein the one or more processors execute the one or more computer programs to implement the steps of the method described above.

[0014] A fourth aspect of this application also provides a computer-readable storage medium having a computer program or instructions stored thereon, which, when executed by a processor, implement the steps of the above-described method.

[0015] The fifth aspect of this application also provides a computer program product, including a computer program or instructions that, when executed by a processor, implement the steps of the above-described method.

[0016] According to embodiments of this application, a multi-source information perception technology foundation is constructed by simultaneously acquiring data from three different dimensions of the target program: error message text, code call stack, and environment parameters, providing complete data support for subsequent comprehensive analysis. Root cause scores for each called code are jointly generated based on the error message text and code call stack, enabling root cause localization to consider both the semantic content of the error information and the temporal call relationships of the call stack, significantly improving the comprehensiveness and accuracy of localization compared to single-dimensional solutions. An environment parameter consistency verification step is introduced, using runtime environment parameters to filter the preliminary scoring results for rationality, eliminating invalid candidate code that does not match the current operating environment, thus ensuring the environmental reliability of the localization results. Overall, this achieves the technical effect of significantly improving the accuracy and reliability of root cause localization, effectively saving maintenance costs, and enhancing user experience. This invention addresses the shortcomings of existing root cause localization methods that rely solely on a single type of data. These methods lack a mechanism for comprehensive analysis of error message text, code call stack, and runtime environment parameters, resulting in incomplete localization results. Furthermore, the lack of structured verification methods for environmental factors leads to inaccurate localization results in environmentally sensitive error scenarios. Moreover, the inability to reasonably filter localization results may result in identifying code that is completely unreachable or lacks the conditions for execution as the root cause, leading to poor accuracy and low reliability in root cause localization. Attached Figure Description

[0017] The above-mentioned contents, other objects, features and advantages of this application will become clearer from the following description of embodiments with reference to the accompanying drawings, in which:

[0018] Figure 1 The illustrations depict application scenarios of root cause localization methods, apparatus, devices, media, and program products according to embodiments of this application.

[0019] Figure 2 A flowchart illustrating a root cause localization method according to an embodiment of this application is shown schematically.

[0020] Figure 3 This illustration schematically shows a flowchart of the root cause scoring of n calling codes of a target program in the root cause localization method according to an embodiment of the present application;

[0021] Figure 4 This illustration schematically shows a flowchart of generating error message text semantic vectors in the root cause localization method according to an embodiment of this application;

[0022] Figure 5 This illustration schematically shows a flowchart of generating the call stack timing feature vector of a target program in the root cause localization method according to an embodiment of this application;

[0023] Figure 6The flowchart illustrating the generation of fused feature vectors in the root cause localization method according to an embodiment of this application is shown in the illustration.

[0024] Figure 7 This illustration schematically shows a flowchart of verifying the consistency of environmental parameters for n calling codes in a root cause localization method according to an embodiment of this application;

[0025] Figure 8 A schematic diagram illustrating the structure of a root cause localization device according to an embodiment of this application is shown.

[0026] Figure 9 A block diagram schematically illustrates an electronic device suitable for implementing a root cause localization method according to an embodiment of this application. Detailed Implementation

[0027] The embodiments of this application will now be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of this application. In the following detailed description, numerous specific details are set forth to provide a thorough understanding of the embodiments of this application for ease of explanation. However, it will be apparent that one or more embodiments may be implemented without these specific details. Furthermore, descriptions of well-known structures and technologies are omitted in the following description to avoid unnecessarily obscuring the concepts of this application.

[0028] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of this application. The terms “comprising,” “including,” etc., as used herein indicate the presence of the stated features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.

[0029] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein are to be interpreted in a manner consistent with the context of this specification, and not in an idealized or overly rigid way.

[0030] When using expressions such as "at least one of A, B and C", they should generally be interpreted in accordance with the meaning that is commonly understood by those skilled in the art (e.g., "a system having at least one of A, B and C" should include, but is not limited to, a system having A alone, a system having B alone, a system having C alone, a system having A and B, a system having A and C, a system having B and C, and / or a system having A, B and C, etc.).

[0031] The accompanying drawings show some block diagrams and / or flowcharts. It should be understood that some blocks or combinations thereof in the block diagrams and / or flowcharts can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable control device, so that when executed by the processor, these instructions can create means for implementing the functions / operations described in these block diagrams and / or flowcharts.

[0032] First, the technical terms used in this article are explained as follows:

[0033] Multidimensional data: includes program error text, code call stack, and runtime environment parameters.

[0034] Attention mechanism: refers to dynamically adjusting the analysis focus based on the keyword weights in the context of program errors, prioritizing code segments that are highly relevant to the errors.

[0035] Root cause localization: Unlike error type classification, it refers to directly identifying the underlying code logic defects or environment configuration problems that cause program errors.

[0036] Softmax is the most commonly used activation function in the output layer of multi-class classification tasks. It is used to transform an entity vector into a probability distribution, where each value of the probability distribution is between 0 and 1, and all values ​​add up to 1.

[0037] Embodiments of this application provide a root cause localization method, which includes: obtaining error message text of a target program, the code call stack of the target program, and environment parameters of the target program; generating root cause scores for n call codes of the target program based on the error message text and the code call stack of the target program, where n is an integer and n is greater than 0; performing environment parameter consistency verification on the n call codes based on the environment parameters of the target program; and obtaining root cause scores for m call code lines that pass the environment parameter consistency verification, sorting the root cause scores of the m call code lines, and determining the call code with the highest root cause score as the target code for root cause localization, where m is an integer, m is greater than 0 and m is less than n.

[0038] According to embodiments of this application, a multi-source information perception technology foundation is constructed by simultaneously acquiring data from three different dimensions of the target program: error message text, code call stack, and environment parameters, providing complete data support for subsequent comprehensive analysis. Root cause scores for each called code are jointly generated based on the error message text and code call stack, enabling root cause localization to consider both the semantic content of the error information and the temporal call relationships of the call stack, significantly improving the comprehensiveness and accuracy of localization compared to single-dimensional solutions. An environment parameter consistency verification step is introduced, using runtime environment parameters to filter the preliminary scoring results for rationality, eliminating invalid candidate code that does not match the current operating environment, thus ensuring the environmental reliability of the localization results. Overall, this achieves the technical effect of significantly improving the accuracy and reliability of root cause localization, effectively saving maintenance costs, and enhancing user experience. This invention addresses the shortcomings of existing root cause localization methods that rely solely on a single type of data. These methods lack a mechanism for comprehensive analysis of error message text, code call stack, and runtime environment parameters, resulting in incomplete localization results. Furthermore, the lack of structured verification methods for environmental factors leads to inaccurate localization results in environmentally sensitive error scenarios. Moreover, the inability to reasonably filter localization results may result in identifying code that is completely unreachable or lacks the conditions for execution as the root cause, leading to poor accuracy and low reliability in root cause localization.

[0039] Figure 1 An exemplary distributed system architecture 100 that can be applied to the root cause localization method according to embodiments of this application is illustrated. It should be noted that... Figure 1 The examples shown are merely examples of the architecture that can be applied to the methods according to the embodiments of this application, in order to help those skilled in the art understand the technical content of this application, but do not imply any limitation on the application scenarios of the embodiments of this application.

[0040] Figure 1 The diagram illustrates an application scenario of the root cause localization method according to an embodiment of this application. For example... Figure 1 As shown, application scenario 100 according to an embodiment of this application may include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 serves as a medium for providing a communication link between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired or wireless communication links or fiber optic cables. For example, a user can use the first terminal device 101, the second terminal device 102, and the third terminal device 103 to interact with the server 105 through the network 104 to receive or send information, etc.

[0041] The first terminal device 101, the second terminal device 102, and the third terminal device 103 can be electronic devices such as smartphones, wearable devices, personal computers, intelligent voice interaction devices, smart home appliances, intelligent vehicles, in-vehicle terminals, aircraft, unmanned vending terminals, and extended reality devices. Extended reality devices can include virtual reality devices, augmented reality devices, and mixed reality devices. A client application for the target application can be installed and run on the terminal devices. This target application can include, but is not limited to, financial transaction applications, payment applications, shopping applications, web browser applications, search applications, instant messaging tools, email clients, and social media platform software (these are just examples). Furthermore, this application embodiment does not limit the form of the target application, and it can include, but is not limited to, applications, mini-programs, etc., installed on the terminal devices, and can also be in the form of web pages.

[0042] Server 105 can be a server providing various services, such as a backend management server supporting websites browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103 (this is just an example). The backend management server can analyze and process received user requests and other data, and feed back the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal devices. The server can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services such as cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks, and basic cloud computing services such as big data. The server can be the backend server of the aforementioned target application, used to provide backend services to the clients of the target application.

[0043] It should be noted that the root cause localization method provided in this application embodiment can generally be executed by server 105 and / or terminal devices 101-103. Accordingly, the root cause localization device provided in this application embodiment can generally be disposed in server 105 and / or terminal devices 101-103.

[0044] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.

[0045] The following will be based on Figure 1 The described scene, through Figures 2-7The root cause localization method according to the disclosed embodiments is described in detail. It should be noted that the above application scenarios are shown only to facilitate understanding of the spirit and principles of this application, and the implementation of this application is not limited in any way. Rather, the implementation of this application can be applied to any applicable scenario.

[0046] Figure 2 A flowchart illustrating a root cause localization method according to an embodiment of this application is shown schematically.

[0047] like Figure 2 As shown, the method 200 includes steps S210 to S240.

[0048] Step S210: Obtain the error message text of the target program, the code call stack of the target program, and the environment parameters of the target program.

[0049] For example, when a monitoring system detects a runtime exception in a target program, it acquires multi-dimensional data at the time of the exception. This multi-dimensional data typically includes three key parts: the error message text of the target program, the code call stack of the target program, and the environmental parameters of the target program. The error message text is the top information of the exception log stack generated by the service at the time of the failure, collected from the log center. The code call stack is the complete stack trace of the current request chain from the entry controller to the data access object layer, obtained from the application performance management system. Environmental parameters include the processor utilization, memory usage, memory size, garbage collection frequency, and request traffic rate of the system hosting the service at that moment.

[0050] Step S220: Based on the error message text of the target program and the code call stack of the target program, generate root cause scores for n call codes of the target program, where n is an integer and n is greater than 0.

[0051] Figure 3 The flowchart illustrating the root cause scoring of n calling codes of a target program in the root cause localization method according to an embodiment of the present application is shown in the illustration.

[0052] like Figure 3 As shown, the method 300 includes steps S310 to S340.

[0053] Step S310: Extract features from the error message text of the target program to generate a semantic vector of the error message text of the target program.

[0054] Figure 4 The flowchart illustrating the generation of error message text semantic vectors in the root cause localization method according to an embodiment of this application is shown in the illustration.

[0055] like Figure 4As shown, the method 400 includes steps S410 to S420.

[0056] Step S410: Perform data cleaning and data standardization on the error message text of the target program to generate standardized error message text.

[0057] For example, after obtaining the original error message text, regular expressions are used to remove dynamically changing content, such as specific IP addresses, port numbers, timestamps, and meaningless temporary variable values. Simultaneously, exception class names and method names are uniformly converted to lowercase, and special punctuation marks are removed, thus obtaining standardized error message text.

[0058] Step S420: Input the standardized error message text into a pre-trained semantic analysis model that is encoded at the embedding layer of the model, and output the semantic vector of the error message text.

[0059] For example, standard error message text can be input into a pre-trained semantic analysis model, and position encoding can be superimposed on the semantic analysis model to retain the positional order information of words in the error message text. Then, the output data of the semantic analysis model can be pooled to generate a semantic vector of the error message text.

[0060] By overlaying positional encoding into the model's embedding layer, the model can retain the positional order information of words in the error message text while extracting semantic features, making the generated semantic vector of the error message text more accurate and reliable.

[0061] Return to reference Figure 3 In step S320, feature extraction is performed on the code call stack of the target program to generate the call stack timing feature vector of the target program.

[0062] Figure 5 The flowchart illustrating the generation of the call stack timing feature vector of a target program in the root cause localization method according to an embodiment of this application is shown.

[0063] like Figure 5 As shown, the method 500 includes steps S510 to S530.

[0064] Step S510: Perform data cleaning and data standardization on the code call stack of the target program to generate a standardized code call stack of the target program.

[0065] For example, data cleaning of the target program's code call stack can include: denoising the target program's code call stack to generate a cleaned code call stack. Specifically, this can include: obtaining the original code call stack, which is usually presented in text form, and removing underlying framework call frames that are irrelevant to business logic, as these frames exist in most faults and contribute significant noise to root cause localization.

[0066] Standardizing the target program's code call stack can include: standardizing and parsing the cleaned code call stack, extracting the function identifiers and location information of each call level according to the call order from top to bottom of the stack, generating a standard call stack sequence as the target program's code call stack; arranging the remaining business-related stack frames after data cleansing in order from top to bottom, and performing unified naming and encoding to generate a standardized code call stack for the target program.

[0067] Step S520: Input the standardized code call stack of the target program into the pre-trained bidirectional long short-term memory network model, and output the forward temporal feature vector and the backward temporal feature vector of the code call stack of the target program.

[0068] For example, a standard call stack sequence can be input into a pre-trained bidirectional long short-term memory network. The bidirectional long short-term memory network performs temporal modeling on the standard call stack sequence in both the forward and backward directions, and outputs a forward hidden state sequence and a backward hidden state sequence as the forward and backward temporal feature vectors of the target program's code call stack.

[0069] Step S530: The forward timing feature vector and the backward timing feature vector of the code call stack of the target program are concatenated to generate the call stack timing feature vector of the target program.

[0070] For example, the last hidden state of the forward hidden state sequence can be concatenated or summed with the last hidden state of the reverse hidden state sequence to obtain a call stack timing feature vector with doubled dimensions, which can be used as the call stack timing feature vector of the target program.

[0071] By modeling the call stack using a pre-trained bidirectional long short-term memory network, features are extracted along the forward and backward time sequences of the call stack using two neural networks, one forward and one backward. This simultaneously captures historical accumulation information from the bottom to the top of the stack and anomaly propagation backtracking information from the top to the bottom. The forward and backward time sequence feature vectors are then concatenated, resulting in a call stack time sequence feature vector that comprehensively integrates information from both directions. This more fully represents the temporal structure and contextual dependencies of the call stack, effectively improving the quality and reliability of the generated call stack time sequence feature vector for the target program, and providing higher-quality call stack feature input for subsequent fusion steps.

[0072] Return to reference Figure 3 In step S330, the error message text semantic vector of the target program and the call stack timing feature vector of the target program are fused to generate a fused feature vector.

[0073] Figure 6 The flowchart illustrating the generation of fused feature vectors in the root cause localization method according to an embodiment of this application is shown in the illustration.

[0074] like Figure 6 As shown, the method 600 includes steps S610 to S630.

[0075] Step S610: By inputting the error message text semantic vector of the target program into a pre-trained neural network model that has been fine-tuned according to a text attention mechanism, the first weight of the error message text semantic vector of the target program is output.

[0076] For example, the semantic vector of the error message text from the target program is input into a pre-trained neural network model. This model is fine-tuned based on a basic multilayer perceptron using a text attention mechanism. This attention mechanism allows the model to dynamically focus on the most critical semantic dimensions of the error root cause, such as those representing "null pointer" or "timeout." The model outputs a scalar value, denoted as the first weight, which ranges from 0 to 1 and represents the importance of the error text in the root cause determination.

[0077] Step S620: By inputting the call stack timing feature vector of the target program into a pre-trained graph neural network model that has been fine-tuned according to a text attention mechanism, the second weight of the call stack timing feature vector of the target program is output.

[0078] For example, the call stack temporal feature vector is input into another pre-trained graph neural network model. This model treats each method call in the call stack as a node in the graph, the call relationships as edges, and is fine-tuned using a text attention mechanism to focus on the most anomalous subgraph structure in the call chain. The model outputs another scalar value, denoted as the second weight, which ranges from 0 to 1 and represents the importance of the call stack temporal features in root cause determination.

[0079] Step S630: Based on the first weight of the error message text semantic vector of the target program and the second weight of the call stack timing feature vector of the target program, perform vector fusion on the error message text semantic vector of the target program and the call stack timing feature vector of the target program to generate the fused feature vector.

[0080] By utilizing a pre-trained neural network model fine-tuned with a text attention mechanism to process the semantic vector of error message text, the model can adaptively filter out the key semantic segments most relevant to root cause identification from the error message text and assign them higher first weights, effectively suppressing the interference of redundant text information on localization. A pre-trained graph neural network model fine-tuned with an attention mechanism is used to process the temporal feature vector of the call stack. The graph neural network can construct a graph structure using the call relationships between function call frames in the call stack, and perform attention calculations on the graph, enabling the model to focus on the key call frames most relevant to the root cause in the call chain, suppressing the noise influence of irrelevant call frames. Adaptive vector fusion based on first and second weights achieves dynamic weight allocation between the text modality and the call stack modality, significantly improving the fused features' ability to represent and discriminate root causes. The two attention mechanisms complement each other, simultaneously filtering root cause clues from two dimensions: "semantic keywords" and "key nodes in the call relationship," improving the model's ability to focus on root causes and its localization accuracy.

[0081] Return to reference Figure 3 In step S340, the fused feature vector is normalized to generate root cause scores for the n calling codes of the target program.

[0082] For example, the fused feature vector is input into a fully connected layer with a Softmax activation function for normalization. The number of output nodes of this fully connected layer is the same as the number of lines of code n in the call stack, and its output value is the root cause score for each line of code. The sum of all scores is 1, representing the probability distribution that each line of code is the root cause. By extracting features from the error message text to generate semantic vectors, unstructured natural language text is transformed into a structured numerical vector representation, enabling the computer to quantitatively understand and process the semantic content of error information, effectively saving computer resources. Furthermore, fusing feature vectors from two different modalities to obtain a fused feature vector achieves complementarity and integration of textual semantic information and call stack temporal information, effectively improving the accuracy and reliability of generating root cause scores for the n call codes of the target program.

[0083] Return to reference Figure 2 In step S230, the environment parameter consistency of the n calling codes is verified based on the environment parameters of the target program.

[0084] Figure 7 The flowchart illustrating the process of verifying the consistency of environmental parameters for n calling codes in the root cause localization method according to an embodiment of this application is shown.

[0085] like Figure 7 As shown, the method 700 includes steps S710 to S730.

[0086] Step S710: Perform data cleaning and data standardization on the environmental parameters of the target program to generate standardized environmental parameters of the target program.

[0087] For example, data cleaning and standardization of the environmental parameters of the target program may specifically include: aligning all time series data into a 5-minute window before and after the time of the failure, and filling missing values ​​with linear interpolation; converting text-type parameters into standard codes to generate a standardized environmental parameter dataset.

[0088] Step S720: The standardized environmental parameters of the target program are vectorized to generate a uniform-dimensional environmental parameter feature vector for the target program.

[0089] For example, the standardized environmental parameters of the target program can be vectorized using numerical encoding or categorical encoding to generate a uniform-dimensional environmental parameter feature vector for the target program.

[0090] For example, categorical parameters can be processed using one-hot encoding or label encoding, and all processed parameters can be concatenated into an environmental parameter feature vector. For numerical environmental parameters, numerical encoding is used. Specifically, these values ​​are first standardized using standard scores. Then, all standardized numerical parameters are arranged in a fixed lexicographical order of their parameter names and concatenated into a one-dimensional floating-point vector. For categorical or enumerated environmental parameters, categorical encoding is used. In this embodiment, one-hot encoding or embedding encoding is preferred. For parameters with fewer value types, one-hot encoding is used; for parameters with more value types, a vocabulary is constructed, and then a pre-trained word embedding model is used to map it into a low-dimensional dense vector.

[0091] By processing categorical environmental parameters through categorical encoding, discrete category labels are transformed into numerical vector representations, avoiding the incorrect assignment of order to unordered categories. After processing different types of environmental parameters with appropriate encoding strategies and then concatenating them uniformly, a unified-dimensional environmental parameter feature vector is generated. This ensures that semantic information of different types of parameters is not lost or distorted during the encoding process, and also makes the final generated feature vector have a standardized fixed dimension, which can serve as a stable input for the consistency verification step.

[0092] Step S730: Based on the unified dimension of the environment parameter feature vector of the target program, perform environment parameter consistency verification on the n calling codes.

[0093] For example, the environmental parameter feature vector can be compared with the fused feature vectors of the n calling codes generated in the previous steps using a dot product or cosine similarity calculation. A dynamic threshold is set; if the similarity between the feature vector of a calling code and the environmental parameter feature vector exceeds the threshold, the execution logic of that code line is determined to be highly relevant to the current environmental condition, i.e., it has passed the consistency verification; otherwise, the code line is considered to be environment-independent general logic and is removed from the candidate root cause list.

[0094] According to embodiments of this application, a configuration environment and code compatibility rule base can be constructed based on the environment parameter feature vector of the target program in a unified dimension. Based on the matching judgment between the configuration environment and code compatibility rule base and the environment parameter feature vector, environment parameter consistency verification is performed on the n calling code lines. The root cause scores of the g calling code lines that failed the environment parameter consistency verification are obtained and set to invalid values ​​or deleted from the score list, where g is an integer, g is greater than 0 and g is less than n.

[0095] By using the generated environment parameter feature vector to perform consistency verification on n calling codes, and by matching and calculating the preconditions such as compilation conditions and runtime constraints that each calling code depends on with the environment parameter feature vector, candidate code that has execution reachability and compatibility in the current running environment is selected, and calling code that is unreachable, incompatible or invalid in the current environment is eliminated. This improves the reliability of the final location result, effectively avoids locating code that cannot be executed in the current environment, and improves operation and maintenance efficiency.

[0096] Return to reference Figure 2 In step S240, the root cause scores of m lines of code that have passed the environmental parameter consistency verification are obtained, the root cause scores of the m lines of code are sorted, and the call code with the highest root cause score is determined as the target code for root cause localization, where m is an integer, m is greater than 0 and m is less than n. Figure 8 A schematic block diagram of a root cause localization device according to an embodiment of this application is shown.

[0097] like Figure 8 As shown, the device 800 includes: a first acquisition module 810, a first generation module 820, a first verification module 830, and a first determination module 840.

[0098] The first acquisition module 810 is used to acquire the error message text of the target program, the code call stack of the target program, and the environment parameters of the target program. In one embodiment, the first acquisition module 810 can be used to execute step S210 described above, which will not be repeated here.

[0099] The first generation module 820 is configured to generate root cause scores for n call codes of the target program based on the error message text of the target program and the code call stack of the target program, where n is an integer and n is greater than 0. In one embodiment, the first generation module 820 can be used to execute step S220 described above.

[0100] The first generation module 820 includes: a second generation module, a third generation module, a fourth generation module, and a fifth generation module.

[0101] The second generation module is used to extract features from the error message text of the target program and generate a semantic vector of the error message text of the target program. In one embodiment, the second generation module can be used to execute step S310 described above.

[0102] The second generation module includes: the sixth generation module and the seventh generation module.

[0103] The sixth generation module is used to perform data cleaning and data standardization on the error message text of the target program to generate standardized error message text. In one embodiment, the sixth generation module can be used to execute step S410 described above, which will not be repeated here.

[0104] The seventh generation module is used to pre-train and generate a semantic analysis model that encodes the standardized error message text input at a position superimposed on the embedding layer of the model, and outputs the semantic vector of the error message text. In one embodiment, the seventh generation module can be used to perform step S420 described above, which will not be repeated here.

[0105] The third generation module is used to extract features from the code call stack of the target program and generate a call stack timing feature vector of the target program. In one embodiment, the third generation module can be used to execute step S320 described above.

[0106] The third generation module includes: the eighth generation module, the ninth generation module, and the tenth generation module.

[0107] The eighth generation module is used to perform data cleaning and standardization on the code call stack of the target program to generate a standardized code call stack of the target program. In one embodiment, the eighth generation module can be used to execute step S510 described above, which will not be repeated here.

[0108] The ninth generation module is used to input the standardized code call stack of the target program into a pre-trained bidirectional long short-term memory network model, and output the forward temporal feature vector and the backward temporal feature vector of the code call stack of the target program. In one embodiment, the ninth generation module can be used to execute step S520 described above, which will not be repeated here.

[0109] The tenth generation module is used to concatenate the forward and backward timing feature vectors of the target program's code call stack to generate the call stack timing feature vector of the target program. In one embodiment, the tenth generation module can be used to execute step S530 described above, which will not be repeated here.

[0110] The fourth generation module is used to perform vector fusion on the error message text semantic vector of the target program and the call stack timing feature vector of the target program to generate a fused feature vector. In one embodiment, the fourth generation module can be used to execute step S330 described above.

[0111] The fourth generation module includes: the eleventh generation module, the twelfth generation module, and the thirteenth generation module.

[0112] The eleventh generation module is used to input the error message text semantic vector of the target program into a pre-trained neural network model that is fine-tuned according to a text attention mechanism, and output the first weight of the error message text semantic vector of the target program. In one embodiment, the eleventh generation module can be used to execute step S610 described above, which will not be repeated here.

[0113] The twelfth generation module is used to input the call stack timing feature vector of the target program into a pre-trained graph neural network model that has been fine-tuned according to a text attention mechanism, and output the second weight of the call stack timing feature vector of the target program. In one embodiment, the twelfth generation module can be used to execute step S620 described above, which will not be repeated here.

[0114] The thirteenth generation module is used to perform vector fusion on the error message text semantic vector and the call stack timing feature vector of the target program based on the first weight of the error message text semantic vector and the second weight of the call stack timing feature vector of the target program, to generate the fused feature vector. In one embodiment, the thirteenth generation module can be used to execute step S630 described above, which will not be repeated here.

[0115] The fifth generation module is used to normalize the fused feature vector to generate root cause scores for the n calling codes of the target program. In one embodiment, the fifth generation module can be used to perform step S340 described above, which will not be repeated here.

[0116] The first verification module 830 is used to perform environment parameter consistency verification on the n calling codes based on the environment parameters of the target program. In one embodiment, the first verification module 830 can be used to execute step S230 described above.

[0117] The first verification module 830 includes: the fourteenth generation module, the fifteenth generation module, and the second verification module.

[0118] The fourteenth generation module is used to perform data cleaning and data standardization on the environmental parameters of the target program to generate standardized environmental parameters for the target program. In one embodiment, the fourteenth generation module can be used to execute step S710 described above, which will not be repeated here.

[0119] The fifteenth generation module is used to vectorize the standardized environmental parameters of the target program to generate a uniform-dimensional environmental parameter feature vector for the target program. In one embodiment, the fifteenth generation module can be used to perform step S720 described above.

[0120] The fifteenth generation module includes: the sixteenth generation module, which is used to vectorize the standardized environmental parameters of the target program using numerical encoding or category encoding to generate a uniform-dimensional environmental parameter feature vector of the target program.

[0121] The second verification module is used to perform environment parameter consistency verification on the n calling codes based on the unified-dimensional environment parameter feature vector of the target program. In one embodiment, the second verification module can be used to execute step S730 described above, which will not be repeated here.

[0122] The first determining module 840 is used to obtain the root cause scores of m lines of code that have passed the environmental parameter consistency verification, sort the root cause scores of the m lines of code, and determine the calling code with the highest root cause score as the target code for root cause localization, where m is an integer, m is greater than 0 and m is less than n. In one embodiment, the first determining module 840 can be used to execute step S240 described above, which will not be repeated here.

[0123] According to embodiments of this application, any multiple modules among the first acquisition module 810, first generation module 820, first verification module 830, and first determination module 840 can be combined into one module, or any one of these modules can be split into multiple modules. Alternatively, at least some of the functions of one or more of these modules can be combined with at least some of the functions of other modules and implemented in one module. According to embodiments of this application, at least one of the first acquisition module 810, first generation module 820, first verification module 830, and first determination module 840 can be at least partially implemented as hardware circuits, such as field-programmable gate arrays, programmable logic arrays, systems-on-a-chip, systems-on-a-substrate, systems-on-package, application-specific integrated circuits, or implemented by any other reasonable means of integrating or packaging circuits, or implemented by any one of software, hardware, and firmware implementations, or by a suitable combination of any of these. Alternatively, at least one of the first acquisition module 810, the first generation module 820, the first verification module 830, and the first determination module 840 may be implemented at least partially as a computer program module, which can perform corresponding functions when the computer program module is run.

[0124] Figure 9 A block diagram schematically illustrates an electronic device suitable for implementing a root cause localization method according to an embodiment of this application.

[0125] like Figure 9As shown, an electronic device 900 according to an embodiment of this application includes a processor 901, which can perform various appropriate actions and processes according to a program stored in a read-only memory 902 or a program loaded from a storage portion 908 into a random access memory 903. The processor 901 may include, for example, a general-purpose microprocessor, an instruction set processor and / or an associated chipset and / or a dedicated microprocessor. The processor 901 may also include onboard memory for caching purposes. The processor 901 may include a single processing unit or multiple processing units for executing different steps of the method flow according to an embodiment of this application.

[0126] Random access memory 903 stores various programs and data required for the operation of electronic device 900. Processor 901, read-only memory 902, and random access memory 903 are interconnected via bus 904. Processor 901 executes various steps of the method flow according to embodiments of this application by executing programs stored in read-only memory 902 and / or random access memory 903. It should be noted that the programs may also be stored in one or more memories other than read-only memory 902 and random access memory 903. Processor 901 may also execute various steps of the method flow according to embodiments of this application by executing programs stored in said one or more memories.

[0127] According to embodiments of this application, the electronic device 900 may further include an input / output interface 905, which is also connected to a bus 904. The electronic device 900 may also include one or more of the following components connected to the input / output interface 905: an input section 906 including a keyboard, mouse, etc.; an output section 907 including a cathode ray tube, liquid crystal display, etc., and a speaker, etc.; a storage section 908 including a hard disk, etc.; and a communication section 909 including a network interface card, such as a local area network card, modem, etc. The communication section 909 performs communication processing via a network such as the Internet. A drive 910 is also connected to the input / output interface 905 as needed. A removable medium 911, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on the drive 910 as needed so that computer programs read from it can be installed into the storage section 908 as needed.

[0128] Embodiments of this application also provide a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or it may exist independently and not assembled into the device / apparatus / system. The computer-readable storage medium carries one or more programs, which, when executed, implement the method according to the embodiments of this application.

[0129] According to embodiments of this application, the computer-readable storage medium can be a non-volatile computer-readable storage medium, such as including but not limited to: portable computer disks, hard disks, random access memory, read-only memory, erasable programmable read-only memory, portable compact disk read-only memory, optical storage devices, magnetic storage devices, or any suitable combination thereof. In embodiments of this application, the computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. For example, according to embodiments of this application, the computer-readable storage medium may include the read-only memory 902 described above, and / or random access memory 903, and / or one or more memories other than read-only memory 902 and random access memory 903.

[0130] Embodiments of this application also include a computer program product comprising a computer program containing program code for performing the methods shown in the flowchart. When the computer program product is run on a computer system, the program code is used to cause the computer system to implement the methods provided in the embodiments of this application.

[0131] In one embodiment, the computer program may rely on a tangible storage medium such as an optical storage device or a magnetic storage device. In another embodiment, the computer program may also be transmitted and distributed in the form of signals over a network medium, and downloaded and installed via the communication section 909, and / or installed from a removable medium 911. The program code contained in the computer program can be transmitted using any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination thereof.

[0132] In embodiments of this application, the computer program can be downloaded and installed from a network via communication section 909, and / or installed from removable medium 911. When the computer program is executed by processor 901, it performs the functions defined in the system of embodiments of this application. According to embodiments of this application, the systems, devices, apparatuses, modules, units, etc., described above can be implemented by computer program modules.

[0133] According to embodiments of this application, program code for executing the computer programs provided in the embodiments of this application can be written in any combination of one or more programming languages. Specifically, these computational programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. The program code can be executed entirely on the user's computing device, partially on the user's device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).

[0134] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0135] Those skilled in the art will understand that the features described in the various embodiments of this application can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in this application. In particular, the features described in the various embodiments of this application can be combined and / or combined in various ways without departing from the spirit and teachings of this application. All such combinations and / or combinations fall within the scope of this application.

Claims

1. A root cause localization method, characterized in that, The method includes: Obtain the error message text of the target program, the code call stack of the target program, and the environment parameters of the target program; Based on the error message text of the target program and the code call stack of the target program, the root cause scores of n call codes of the target program are generated, where n is an integer and n is greater than 0; Based on the environment parameters of the target program, the consistency of the environment parameters is verified for the n calling codes; and Obtain the root cause scores of m lines of code that have passed the environmental parameter consistency verification, sort the root cause scores of the m lines of code, and determine the call code with the highest root cause score as the target code for root cause localization, where m is an integer, m is greater than 0 and m is less than n.

2. The method according to claim 1, characterized in that, Based on the error message text of the target program and the code call stack of the target program, root cause scores are generated for n call codes of the target program, including: Feature extraction is performed on the error message text of the target program to generate a semantic vector of the error message text of the target program; Feature extraction is performed on the code call stack of the target program to generate the call stack timing feature vector of the target program; The error message text semantic vector of the target program and the call stack timing feature vector of the target program are fused to generate a fused feature vector; and The fused feature vector is normalized to generate root cause scores for the n calling codes of the target program.

3. The method according to claim 2, characterized in that, Feature extraction is performed on the error message text of the target program to generate a semantic vector of the error message text of the target program, including: The error message text of the target program is cleaned and standardized to generate standardized error message text; and The standardized error message text is input into a pre-trained semantic analysis model that is encoded at the embedding layer of the model, and the output is the semantic vector of the error message text.

4. The method according to claim 2, characterized in that, Feature extraction is performed on the code call stack of the target program to generate a call stack timing feature vector of the target program, including: The code call stack of the target program is cleaned and standardized to generate a standardized code call stack of the target program. The standardized code call stack of the target program is input into a pre-trained bidirectional long short-term memory network model, which outputs the forward and backward temporal feature vectors of the target program's code call stack; and The forward and backward timing feature vectors of the target program's code call stack are concatenated to generate the call stack timing feature vector of the target program.

5. The method according to claim 2, characterized in that, The error message text semantic vector of the target program and the call stack timing feature vector of the target program are fused to generate a fused feature vector, including: By inputting the error message text semantic vector of the target program into a pre-trained neural network model and fine-tuning the model according to a text attention mechanism, the first weight of the error message text semantic vector of the target program is output. By inputting the call stack timing feature vector of the target program into a pre-trained graph neural network model that is fine-tuned based on a text attention mechanism, the model outputs the second weight of the call stack timing feature vector of the target program; and Based on the first weight of the error message text semantic vector of the target program and the second weight of the call stack timing feature vector of the target program, the error message text semantic vector and the call stack timing feature vector of the target program are fused to generate the fused feature vector.

6. The method according to any one of claims 1 to 5, characterized in that, Based on the environment parameters of the target program, the consistency of the environment parameters of the n calling codes is verified, including: The environmental parameters of the target program are cleaned and standardized to generate standardized environmental parameters for the target program. The standardized environmental parameters of the target program are vectorized to generate a uniform-dimensional environmental parameter feature vector for the target program; and Based on the unified-dimensional environmental parameter feature vector of the target program, the consistency of environmental parameters is verified for the n calling codes.

7. The method according to claim 6, characterized in that, The standardized environmental parameters of the target program are vectorized to generate a uniform-dimensional environmental parameter feature vector for the target program, including: The standardized environmental parameters of the target program are vectorized using numerical encoding or categorical encoding to generate a uniform-dimensional environmental parameter feature vector for the target program.

8. A root cause localization device, characterized in that, The device includes: The first acquisition module is used to acquire the error message text of the target program, the code call stack of the target program, and the environment parameters of the target program; The first generation module is used to generate root cause scores for n call codes of the target program based on the error message text of the target program and the code call stack of the target program, where n is an integer and n is greater than 0; The first verification module is used to perform environment parameter consistency verification on the n calling codes based on the environment parameters of the target program; and The first determining module is used to obtain the root cause scores of m lines of code that have passed the environmental parameter consistency verification, sort the root cause scores of the m lines of code, and determine the calling code with the highest root cause score as the target code for root cause localization, where m is an integer, m is greater than 0 and m is less than n.

9. An electronic device, comprising: One or more processors; Memory, used to store one or more computer programs. The characteristic feature is that the one or more processors execute the one or more computer programs to implement the steps of the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program or instructions are executed by a processor, they implement the steps of the method according to any one of claims 1 to 7.

11. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.