An abnormal event fault locating method, device, equipment, medium and product

CN122816951APending Publication Date: 2026-09-25INDUSTRIAL AND COMMERCIAL BANK OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610624740.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-08
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0003]现有联机接口调度的异常处理方式中,多依赖人工排查日志或单一维度调用分析,缺乏对接口调度全链路关联特征的覆盖,容易出现故障根因定位滞后、误判等情况,导致异常事件处理效率低、根因定位精准度不足,无法为金融领域下的联机接口调度的稳定运行提供高效可靠的故障防控

Benefits of technology

[0008]根据本发明的另一方面,提供了一种计算机可读存储介质,计算机可读存储介质存储有计算机指令,计算机指令用于使处理器执行时实现本发明任一实施例的一种异常事件故障定位方法。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122816951A_ABST
    Figure CN122816951A_ABST
Patent Text Reader

Abstract

The application discloses an abnormal event fault positioning method, device, equipment, medium and product. It can be applied to the field of financial technology. The method comprises the following steps: when an abnormal event of business processing in an online interface scheduling process is monitored, link tracking identification and abnormal event context information of the abnormal event of business processing are acquired, and online call full logs associated with the abnormal event of business processing are acquired; a thread-level call relationship graph, a container-level call relationship graph and a cross-container-level call relationship graph are generated; feature coding is performed on the thread-level call relationship graph to generate a thread-level coding vector, feature coding is performed on the container-level call relationship graph to generate a container-level coding vector, and feature coding is performed on the cross-container-level call relationship graph to generate a cross-container-level coding vector; a fusion coding vector is generated according to the thread-level coding vector, the container-level coding vector and the cross-container-level coding vector; and the fusion coding vector is input into a fault prediction model to output a target fault root cause positioning result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of financial technology, and in particular to a method, apparatus, equipment, medium and product for locating faults in abnormal events. Background Technology

[0002] In online interface scheduling scenarios in the fintech field, rapid root cause localization of abnormal business processing events is a core element in ensuring stable system operation and improving service continuity, and is directly related to the processing efficiency of financial business, customer experience, and operational security.

[0003] Existing methods for handling anomalies in online interface scheduling often rely on manual log checks or single-dimensional call analysis, lacking coverage of the full-link correlation characteristics of interface scheduling. This can easily lead to delayed root cause location and misjudgment, resulting in low efficiency in handling anomalies and insufficient accuracy in root cause location. Consequently, these methods cannot provide efficient and reliable fault prevention for the stable operation of online interface scheduling in the financial sector. Summary of the Invention

[0004] This invention provides a method, apparatus, device, medium, and product for locating abnormal events, so as to improve the accuracy of determining the root cause of abnormal events in financial business scenarios.

[0005] According to one aspect of the present invention, a method for locating faults in abnormal events is provided, the method comprising: When an abnormal event in the online interface scheduling process is detected, the link tracing identifier and abnormal event context information of the abnormal event are obtained, and based on the link tracing identifier and abnormal event context information, the full log of the online call associated with the abnormal event is obtained. Based on the full log of online calls, generate thread-level call relationship diagrams, container-level call relationship diagrams, and cross-container-level call relationship diagrams. Feature encoding is performed on the thread-level call relationship graph to generate thread-level encoding vectors, feature encoding is performed on the container-level call relationship graph to generate container-level encoding vectors, and feature encoding is performed on the cross-container-level call relationship graph to generate cross-container-level encoding vectors. Generate a fused encoding vector based on thread-level encoding vectors, container-level encoding vectors, and cross-container-level encoding vectors; The fused encoding vector is input into the pre-trained fault prediction model, which performs fault root cause analysis and outputs the target fault root cause location result of the abnormal business processing event.

[0006] According to another aspect of the present invention, an abnormal event fault location device is provided, the device comprising: The full log call module is used to obtain the link tracing identifier and exception event context information of the business processing exception event when an exception event is detected in the online interface scheduling process, and to obtain the full log of the online call associated with the business processing exception event based on the link tracing identifier and exception event context information. The call relationship graph generation module is used to generate thread-level call relationship graphs, container-level call relationship graphs, and cross-container-level call relationship graphs based on the full log of online calls. The encoding vector generation module is used to perform feature encoding on the thread-level call relationship graph to generate thread-level encoding vectors, to perform feature encoding on the container-level call relationship graph to generate container-level encoding vectors, and to perform feature encoding on the cross-container-level call relationship graph to generate cross-container-level encoding vectors. The fusion encoding vector generation module is used to generate fusion encoding vectors based on thread-level encoding vectors, container-level encoding vectors, and cross-container-level encoding vectors. The fault root cause localization module is used to input the fused encoded vector into the pre-trained fault prediction model, which then performs fault root cause analysis and outputs the target fault root cause localization result for abnormal business processing events.

[0007] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising: At least one processor; and A memory that is communicatively connected to at least one processor; wherein, The memory stores a computer program that can be executed by at least one processor, such that the at least one processor is able to execute an abnormal event fault location method according to any embodiment of the present invention.

[0008] According to another aspect of the present invention, a computer-readable storage medium is provided, which stores computer instructions for causing a processor to execute an abnormal event fault location method according to any embodiment of the present invention.

[0009] According to another aspect of the present invention, a computer program product is provided, comprising a computer program that, when executed by a processor, implements an abnormal event fault location method according to any embodiment of the present invention.

[0010] When an abnormal service processing event is detected during the online interface scheduling process, the technical solution of this invention obtains the link tracing identifier and abnormal event context information of the abnormal service processing event, and obtains the full log of online calls associated with the abnormal service processing event based on the link tracing identifier and abnormal event context information; generates thread-level call relationship graphs, container-level call relationship graphs, and cross-container-level call relationship graphs based on the full log of online calls; performs feature encoding on the thread-level call relationship graphs to generate thread-level encoding vectors, performs feature encoding on the container-level call relationship graphs to generate container-level encoding vectors, and performs feature encoding on the cross-container-level call relationship graphs to generate cross-container-level encoding vectors; generates a fused encoding vector based on the thread-level encoding vector, container-level encoding vector, and cross-container-level encoding vector; inputs the fused encoding vector into a pre-trained fault prediction model, which performs fault root cause analysis and outputs the target fault root cause location result of the abnormal service processing event. It can obtain the full online call log based on link tracing identifiers and abnormal context information, generate thread-level, container-level, and cross-container-level call relationship graphs, and complete feature encoding and vector fusion. Then, through a pre-trained fault prediction model, it can achieve root cause analysis and accurate location of abnormal events, effectively solving the problems of incomplete information coverage and low location efficiency in traditional single-dimensional investigation methods. It improves the processing efficiency and root cause location accuracy of online interface scheduling abnormal events in the financial field, and provides reliable technical support for the stable operation of the system.

[0011] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0012] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0013] Figure 1 This is a flowchart of an abnormal event fault location method provided in Embodiment 1 of the present invention; Figure 2 This is a flowchart of an abnormal event fault location method provided in Embodiment 2 of the present invention; Figure 3 This is a flowchart of an abnormal event fault location method provided in Embodiment 3 of the present invention; Figure 4 This is a schematic diagram of the structure of an abnormal event fault location device provided in Embodiment 4 of the present invention; Figure 5 This is a schematic diagram of the structure of an electronic device that implements an abnormal event fault location method according to an embodiment of the present invention. Detailed Implementation

[0014] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0015] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0016] Example 1 Figure 1 This is a flowchart of an abnormal event fault location method provided in Embodiment 1 of the present invention. This embodiment is applicable to the situation of accurately locating abnormal events occurring during interface scheduling in the financial field. This method can be executed by an abnormal event fault location device, which can be implemented in hardware and / or software and can be configured in an electronic device. Figure 1 As shown, the method includes: S101. When an abnormal event in the online interface scheduling process is detected, the link tracing identifier and abnormal event context information of the abnormal event are obtained, and the full log of the online call associated with the abnormal event is obtained based on the link tracing identifier and abnormal event context information.

[0017] S102. Based on the full log of online calls, generate thread-level call relationship diagrams, container-level call relationship diagrams, and cross-container-level call relationship diagrams.

[0018] S103. Perform feature encoding on the thread-level call relationship graph to generate a thread-level encoding vector, perform feature encoding on the container-level call relationship graph to generate a container-level encoding vector, and perform feature encoding on the cross-container-level call relationship graph to generate a cross-container-level encoding vector.

[0019] S104. Generate a fused encoding vector based on the thread-level encoding vector, container-level encoding vector, and cross-container-level encoding vector.

[0020] S105. Input the fused encoding vector into the pre-trained fault prediction model, and the fault prediction model performs fault root cause analysis and outputs the target fault root cause location result of the abnormal business processing event.

[0021] Among them, business processing exceptions can be service execution failures, response timeouts, abnormal return results, or error messages caused by the mixing of multiple containers, multiple threads, and synchronous and asynchronous calls during online calls in a microservice architecture. The trace identifier can be a TraceID, a unique string generated when the online call is initiated, used to associate the logs of the same call in different containers and threads. The exception event context information can be a set of basic information directly related to the business processing exception, specifically including the execution thread identifier, service identifier, container identifier, call timestamp, exception occurrence time, and call parameters. The full online call log can be the complete log of all microservice containers and execution threads involved in the same online call covered by the trace identifier.

[0022] For example, after detecting an abnormal event in business processing, the tracing identifier and abnormal event context information of the interface scheduling process are first extracted. Then, using the tracing identifier as an index, all logs corresponding to that tracing identifier are collected and aggregated to form a complete set of full call logs. For instance, if a transfer interface call of a financial microservice returns an execution failure, and the corresponding tracing identifier is 123, then the abnormal event context information of the tracing identifier 123 is extracted, such as service identifier A, container identifier 003, etc. All relevant logs containing the tracing identifier 123 are collected and integrated to obtain the full online call logs.

[0023] The thread-level call graph can be a graphical structure with thread-level execution units as nodes and inter-thread synchronization methods as edges, where each node contains attributes such as service identifier and execution time. The container-level call graph can be constructed by aggregating multiple thread-level execution units within the same container into nodes and using intra-container thread call relationships as node edges. The cross-container-level call graph can be a global link graph structure with each container as a node and cross-container thread call edges.

[0024] Specifically, thread-level encoding vectors can be obtained by vector encoding the thread-level call relationship graph using a graph neural network. Container-level encoding vectors can be obtained by vector encoding the container-level call relationship graph using a graph neural network. Cross-container-level encoding vectors can be obtained by vector encoding the cross-container-level call relationship graph using a graph neural network.

[0025] For example, for thread-level call relationship graphs, graph neural networks are used to mine the call relationship structure between threads and generate vectors of structural features. For container-level call relationship graphs, graph neural networks are used to extract structural features of the overall execution state and related logic of the aggregated threads in the container and encode them to obtain container-level encoding vectors. For cross-container-level call relationship graphs, the topology of the call chain and the environmental relationship features between different containers are captured to complete the structural encoding and obtain cross-container-level encoding vectors.

[0026] Furthermore, to accurately describe the structural and temporal characteristics of the thread-level call relationship graph and improve the accuracy of subsequent analysis based on the thread-level call relationship graph, in an optional embodiment, feature encoding is performed on the thread-level call relationship graph to generate a thread-level encoding vector, including: Step a1: Perform structural encoding on the thread-level call relationship graph to generate a thread-level structure-aware vector.

[0027] Step a2: Perform timing encoding on the thread-level call relationship graph to generate a thread-level timing relationship vector.

[0028] Step a3: Perform vector fusion on the thread-level structure awareness vector and the thread-level temporal relationship vector to generate the thread-level encoding vector.

[0029] Among them, the thread-level structure-aware vector can be a vector obtained by encoding the topology of the thread-level call relationship graph through a graph neural network.

[0030] For example, for the thread call relationship graph of the order thread unit and the payment thread unit, if the structure encoding generates a 5-dimensional thread-level structure awareness vector [0.1, 0.98, 1.0, 1.0, 1], each dimension can correspond to the abnormal state of the order thread unit, the average memory usage of the payment thread unit, the synchronous call type identifier, the call confidence, and the number of connections of the payment thread unit, respectively.

[0031] Among them, the thread-level time sequence relationship vector can be obtained by encoding the time sequence features of the thread-level call relationship graph.

[0032] For example, for the thread call relationship diagram of the order thread unit and the payment thread unit, if the timing encoding generates a 5-dimensional timing relationship vector [0, 77, 5000, 5000, 0.99], where each dimension corresponds to the timing identifier of the order thread unit, the start time difference between the two, the execution time of the order thread unit, the execution time of the payment thread unit, and the timing correlation weight, respectively.

[0033] For example, the thread-level structure-aware vector and the thread-level temporal relationship vector can be added together to obtain the thread-level encoding vector. For instance, the structure-aware vector [0.1, 0.98, 1.0, 1.0, 1] and the temporal relationship vector [0, 77, 5000, 5000, 0.99] can be fused with equal weights to generate the final 5-dimensional thread-level encoding vector [0.1, 77.98, 5001.0, 5001.0, 1.99].

[0034] The above technical solution accurately captures the global structure of thread calls, effectively analyzes key temporal correlations and structural features, and then integrates thread-level structure perception vectors and thread-level temporal relationship vectors through vector fusion to eliminate multi-threaded concurrency interference, fully preserve the structural and temporal information of thread calls, and improve the completeness and accuracy of thread-level features.

[0035] The fused encoding vector can be generated by concatenating or weighting thread-level encoding vectors, container-level encoding vectors, and cross-container-level encoding vectors.

[0036] For example, a vector concatenation method is used to connect the thread-level encoding vector, container-level encoding vector, and cross-container-level encoding vector end to end in sequence to form a fused encoding vector with a dimension equal to the sum of the three. For instance, if the thread-level encoding vector is [0.36, 0.31, 0.34], the container-level encoding vector is [0.35, 0.41], and the cross-container-level encoding vector is [0.42, 0.36], concatenating them in sequence will result in a fused encoding vector of [0.36, 0.31, 0.34, 0.35, 0.41, 0.42, 0.36].

[0037] Among them, the fault prediction model can be used to locate the root cause of abnormal events in business processing. The target fault root cause location result can be the fault root cause location result of the abnormal event in business processing output by the fault prediction model.

[0038] Furthermore, to address the limitations of a single analytical dimension in adapting to complex faults and to improve the comprehensiveness and accuracy of root cause analysis, in one optional embodiment, the fault prediction model includes at least one expert prediction module and a routing decision module. The fused encoded vector is input into the pre-trained fault prediction model, which performs root cause analysis and outputs the target fault root cause localization result for the abnormal business processing event, including: Step b1: Input the fused coding vector into at least one expert prediction module in the fault prediction model. The expert prediction module performs fault root cause analysis based on the fused coding vector to obtain the fault root cause prediction results output by each expert prediction module.

[0039] Step b2: Input the root cause prediction results output by each expert prediction module into the routing decision module in the fault prediction model. The routing decision module then fuses the root cause prediction results of each expert prediction module according to the current weight parameters of each module to obtain the target fault root cause location result corresponding to the abnormal business processing event.

[0040] Different expert prediction modules can perform root cause analysis and localization from different dimensions of root cause analysis, such as a container-level fault responsibility probability analysis expert module, a container-internal thread relationship analysis expert module, and a cross-container overall call chain analysis expert module. The root cause prediction result can be the root cause result output by the expert prediction module after analyzing the fused coding vector.

[0041] For example, different expert prediction modules can extract relevant features from the fused encoding vector, output the root cause prediction result of the fault, and give the probability. For instance, the container-level expert module can extract container-related features, and the output root cause prediction result is a solver failure probability of 0.91. The thread relationship expert module can extract thread-related features, and the output root cause prediction result is a thread parameter error probability of 0.96.

[0042] The routing decision module can dynamically adjust the weights of each expert module based on the characteristics of abnormal scenarios, and integrate the root cause prediction results output by each expert module. The current weight parameters can be the weight values ​​assigned to each expert module by the routing decision module based on scenario characteristics.

[0043] For example, the routing decision module performs scene identification on the fused encoding vector, then assigns weights to each expert module according to the scene type. The prediction results of each expert module for different candidate fault nodes are weighted and summed, and the candidate node with the highest comprehensive score is selected as the target fault root cause localization result. For instance, if the scene identification is an e-commerce order payment timeout anomaly, the routing decision module assigns weights of 0.4, 0.3, and 0.3 to the container resource expert module, the cross-container structure expert module, and the thread-level expert module, respectively. Each expert module outputs root cause confidence scores for the two candidate fault nodes: payment service and accounting service. The container-level expert module predicts payment service anomalies with a confidence score of 0.95 and accounting service anomalies with a confidence score of 0.3; the cross-container-level expert module predicts payment service anomalies with a confidence score of 0.9 and accounting service anomalies with a confidence score of 0.6; and the thread-level expert module predicts payment service anomalies with a confidence score of 0.85 and accounting service anomalies with a confidence score of 0.55. The routing decision module calculates the comprehensive score by weighted summation: the confidence level for payment service anomalies is 0.95×0.4 + 0.9×0.3 + 0.85×0.3 = 0.905, and the confidence level for billing service anomalies is 0.3×0.4 + 0.6×0.3 + 0.55×0.3 = 0.465. Since the comprehensive score for payment service is higher than that for billing service, payment service anomalies are selected as the target root cause localization result.

[0044] The above technical solution can set up multiple expert modules to ensure coverage of multiple dimensions such as topology, resource status, thread execution, log characteristics, and anomaly propagation, avoiding the one-sidedness of a single dimension. At the same time, a routing decision module is set up to dynamically adjust the current weight parameters based on scenario characteristics, so that the analysis results are more in line with the actual fault scenario, effectively solving the problem of insufficient adaptability of a single model in complex environments, and significantly improving the accuracy and reliability of root cause localization.

[0045] Furthermore, to clarify the model training process of the fault prediction model and improve the accuracy of the fault root cause prediction results, in an optional embodiment, the model training method of the fault prediction model is as follows: Step c1: Obtain the historical fusion coding vector and its corresponding historical fault root cause location under the historical time period, and use the historical fault root cause location as the sample label of the historical fusion coding vector.

[0046] Step c2: Input the historical fusion encoding vector with historical fault root cause localization sample labels into the fault prediction model to obtain the fault prediction results output by each expert prediction module, and input the fault prediction results output by each expert prediction module into the routing decision module to obtain the predicted fault root cause localization output by the routing decision module.

[0047] Step c3: Based on the historical fault root cause localization and the predicted fault root cause localization, update the initial weight parameters corresponding to each expert prediction module in the routing decision module of the fault prediction model until the preset model training termination condition is met, and obtain the fault prediction model that has completed training.

[0048] Here, the historical time period can be a preset continuous time period for collecting historical training data, such as the past 6 months. The historical fusion encoding vector can be the fusion encoding vector within the historical time period. The historical fault root cause localization can be the actual fault root cause corresponding to the historical fusion encoding vector. The sample label can be the structured label of the historical fault root cause localization result corresponding to the historical fusion encoding vector. For example, resource anomalies are labeled as 0, network latency anomalies as 1, parameter error anomalies as 2, etc.

[0049] For example, the historical time period is set to the past 6 months. All historical online call logs within this period are collected. A historical fusion encoding vector is generated for each event. Relevant technical personnel label the event. For example, a parameter error scenario with incorrect thread parameter format is labeled as 2, an accounting service timeout exception propagation scenario is labeled as 3, and a normal call without exception is labeled as 4. The sample labels are associated with the corresponding historical fusion encoding vectors to build a model training dataset.

[0050] Among them, the predicted root cause localization of faults can be the predicted root cause of faults output by the routing decision module after analyzing the historical fused coding vectors.

[0051] The initial weight parameters can be the default weight values ​​set for each expert prediction module before model training. The model training termination condition can be a training termination standard pre-set by relevant technical personnel, such as the fault prediction model reaching the maximum number of training iterations on the validation set, or the change in the model loss value being less than a set threshold in 100 consecutive iterations.

[0052] For example, at the start of model training, the training dataset is divided into a training set and a validation set in a 7:3 ratio. The training set is input into the fault prediction model, and the error between the predicted fault root cause location and the historical fault root cause location is calculated using the cross-entropy loss function. Based on the error, the initial weight parameters of each expert module are adjusted using the gradient descent method. When the training is iterated to 1800 times, the encoding matching accuracy of the fault prediction model on the validation set reaches 96%, which meets the threshold set by relevant technical personnel. Training is then stopped, and the fault prediction model that has been trained is obtained.

[0053] The above technical solution constructs training samples by collecting sufficient historical data from multiple scenarios and optimizes model weights based on real labels, enabling the fault prediction model to fully learn the fault characteristics and patterns, improve prediction accuracy and scenario adaptability, and ensure that the model outputs accurate predictions of the root cause of the fault in practical applications.

[0054] When an abnormal service processing event is detected during the online interface scheduling process, the technical solution of this invention obtains the link tracing identifier and abnormal event context information of the abnormal service processing event, and obtains the full log of online calls associated with the abnormal service processing event based on the link tracing identifier and abnormal event context information; generates thread-level call relationship graphs, container-level call relationship graphs, and cross-container-level call relationship graphs based on the full log of online calls; performs feature encoding on the thread-level call relationship graphs to generate thread-level encoding vectors, performs feature encoding on the container-level call relationship graphs to generate container-level encoding vectors, and performs feature encoding on the cross-container-level call relationship graphs to generate cross-container-level encoding vectors; generates a fused encoding vector based on the thread-level encoding vector, container-level encoding vector, and cross-container-level encoding vector; inputs the fused encoding vector into a pre-trained fault prediction model, which performs fault root cause analysis and outputs the target fault root cause location result of the abnormal service processing event. It can obtain the full online call log based on link tracing identifiers and abnormal context information, generate thread-level, container-level, and cross-container-level call relationship graphs, and complete feature encoding and vector fusion. Then, through a pre-trained fault prediction model, it can achieve root cause analysis and accurate location of abnormal events, effectively solving the problems of incomplete information coverage and low location efficiency in traditional single-dimensional investigation methods. It improves the processing efficiency and root cause location accuracy of online interface scheduling abnormal events in the financial field, and provides reliable technical support for the stable operation of the system.

[0055] Example 2 Figure 2This is a flowchart of an abnormal event fault location method provided in Embodiment 2 of the present invention. This embodiment optimizes and improves upon the above technical solutions. The step "generating a thread-level call relationship graph, a container-level call relationship graph, and a cross-container-level call relationship graph based on the full online call logs" is refined to: "Based on a preset first key field, extracting thread-level associated log content from the full online call logs to obtain thread-level associated log data, and constructing a thread-level call relationship graph based on the thread-level associated log data; based on a preset second key field, extracting container-level associated log content from the full online call logs to obtain container-level associated log data, and constructing a container-level call relationship graph based on the container-level associated log data; based on a preset third key field, extracting cross-container-level associated log content from the full online call logs to obtain cross-container-level associated log data, and constructing a cross-container-level call relationship graph based on the cross-container-level associated log data." This improves the generation method of the thread-level call relationship graph, container-level call relationship graph, and cross-container-level call relationship graph.

[0056] It should be noted that for parts not described in detail in the embodiments of the present invention, please refer to the descriptions in other embodiments. For example... Figure 2 As shown, the method includes the following specific steps: S201. When an abnormal event in the online interface scheduling process is detected, obtain the link tracing identifier and abnormal event context information of the abnormal event, and obtain the full log of the online call associated with the abnormal event based on the link tracing identifier and abnormal event context information.

[0057] S202. Based on the preset first key field, extract the thread-level associated log content from the full log of online calls to obtain thread-level associated log data, and construct a thread-level call relationship graph based on the thread-level associated log data.

[0058] S203. Based on the preset second key field, extract the container-level associated log content from the full log of online calls to obtain container-level associated log data, and construct a container-level call relationship graph based on the container-level associated log data.

[0059] S204. Based on the preset third key field, extract the cross-container-level related log content from the full log of online calls to obtain cross-container-level related log data, and construct a cross-container-level call relationship graph based on the cross-container-level related log data.

[0060] S205. Perform feature encoding on the thread-level call relationship graph to generate a thread-level encoding vector, perform feature encoding on the container-level call relationship graph to generate a container-level encoding vector, and perform feature encoding on the cross-container-level call relationship graph to generate a cross-container-level encoding vector.

[0061] S206. Generate a fused encoding vector based on the thread-level encoding vector, container-level encoding vector, and cross-container-level encoding vector.

[0062] S207. Input the fused encoding vector into the pre-trained fault prediction model, and let the fault prediction model perform fault root cause analysis and output the target fault root cause location result of the abnormal business processing event.

[0063] The preset first key field can be a set of core fields pre-defined by relevant technical personnel for extracting thread-level related logs from the full log. For example, the preset first key field could be exception type, service name, etc. The thread-level related log content can be all logs in the full log that contain the preset first key field. The thread-level related log data can be a structured data set obtained after structuring the thread-level related log content.

[0064] For example, if the preset first key field is an exception type of solver error, then the entire online call log can be traversed to filter out logs containing all the preset first key fields. The filtered logs can then undergo field extraction and format standardization to form thread-level associated log data. For instance, after filtering out logs containing the preset first key field, all log information with thread parameter format errors can be extracted to form thread-level associated log data. Then, each log in the thread-level associated log data can be used as a node, and the edges between nodes can represent the relationships between logs, thus constructing a thread-level call relationship graph.

[0065] Furthermore, to accurately describe the calling logic in multi-threaded concurrent scenarios and improve the accuracy of root cause analysis, in one optional embodiment, the thread-level correlation log data includes thread-level basic correlation data, thread-level call sequence data, thread-level call relationship data, and thread-level auxiliary correlation data. Correspondingly, a thread-level call relationship graph is constructed based on the thread-level correlation log data, including: Step d1: Parse the thread-level basic association data to obtain at least one execution thread identifier, and use the execution thread unit corresponding to the execution thread identifier as a node in the initial call relationship graph.

[0066] Step d2: Parse the thread-level call timing data and thread-level call relationship data to determine the thread call relationship between execution thread units.

[0067] Step d3: Based on the thread call relationship, construct the node edges of the initial call relationship graph.

[0068] Step d4: Determine the node attribute information of each node and the edge attribute information of each node in the initial call relationship graph based on the thread-level auxiliary association data to obtain the thread-level call relationship graph.

[0069] Specifically, the thread-level basic association data can be a set of core association fields extracted from logs aggregated by link tracing identifiers after structured preprocessing. The execution thread identifier can be an identifier composed of core fields from the thread-level basic association data. The execution thread unit can be an independent analysis unit formed by integrating information such as the call start and end time, execution duration, and exception status of the corresponding thread, using the execution thread identifier as a unique key. The initial call relationship graph can be a basic graph structure containing only execution thread units as nodes.

[0070] For example, if the thread-level basic association data specifically consists of two sets of core association fields, the first set of association fields corresponds to the order service, the order container, and the thread-related basic information under that service container; the second set of association fields corresponds to the payment service, the payment container, and the thread-related basic information under that service container. From these two sets of thread-level basic association data, three types of core information—service name, container identifier, and thread identifier—are extracted respectively. Two execution thread identifiers are generated through field association combinations. Each execution thread identifier corresponds to an independent execution thread unit, namely the order thread unit and the payment thread unit. These two execution thread units are then used as independent nodes to construct a node graph containing only node information and an initial call relationship diagram.

[0071] The thread-level call timing data can include the start and end times of each execution thread unit. The thread-level call relationship data can include the request task identifier and call type identifier of each execution thread unit. The thread call relationship can be a synchronous or asynchronous call relationship between execution thread units.

[0072] For example, if the thread-level call timing data shows that the order thread unit starts at 10:08 and ends at 10:20, while the payment thread unit starts at 10:10, and the thread-level call relationship data shows that both requests contain the same order number, then the thread call relationship between the order thread unit and the payment thread unit can be considered a synchronous call relationship.

[0073] In the initial call relationship graph, the node edges can be used to connect two execution thread unit nodes that have a call relationship. The type and direction of the edges can correspond to the thread call relationship.

[0074] For example, if the thread call relationship between the order thread unit and the payment thread unit is a synchronous call relationship, then a synchronous call edge can be drawn from the order thread unit to the payment thread unit.

[0075] Among them, thread-level auxiliary correlation data can include resource utilization, exception type, parameter matching similarity, and call confidence calculation results for each execution thread unit. Node attribute information can include core feature data of the execution thread unit, including execution time, average resource utilization, and exception status. Edge attribute information of node edges can provide supplementary data for call relationships, such as call type and parameter matching similarity.

[0076] For example, in thread-level auxiliary association data, if the attributes of the payment thread unit are: average memory usage: 0.98, abnormal status: no abnormality, add attributes to the node edge, such as the edge attribute from the order thread unit to the payment thread unit having a parameter matching similarity of 0.9, and finally obtain the complete thread-level call relationship graph.

[0077] The above embodiment achieves structured construction of thread-level call relationship graphs by extracting nodes from the initial call relationship graph, determining thread call relationships, constructing node edges in the call relationship graph, and finally supplementing node attribute information and edge attribute information. It generates unique execution thread identifiers by combining thread-level basic association data, eliminating multi-threaded concurrency interference, accurately locating each independent execution unit, enriching node and edge attribute information, transforming scattered log data into structured graph data, and reducing the difficulty of root cause localization in complex call scenarios.

[0078] The preset second key field can be a set of core fields extracted from structured logs to determine container-level related logs, such as container identifier, service name, and thread execution duration. The container-level related log content can be all log entries in the full structured log containing the preset second key field. The container-level related log data can be a set of structured data obtained after structuring the container-level related log content.

[0079] For example, if the preset second key field is container A, then all logs containing the preset second key field are filtered out, and all logs with container A are aggregated to form container-level associated log data. A container-level call relationship graph is constructed based on the container-level associated log data: nodes can be different containers, and node edges can represent container call relationships, ultimately forming container-level associated log data and the corresponding container-level call relationship graph.

[0080] The preset third key field can be a set of core fields extracted from structured logs for cross-container-level related logs, such as call type and call confidence. The cross-container-level related log content can be all log entries in the full structured log containing the preset third key field. The cross-container-level related log data can be a set of structured data obtained after structuring the cross-container-level related log content.

[0081] For example, the pre-processed structured logs are traversed according to the third key field pre-defined by relevant technical personnel, and log entries containing all the preset third fields are filtered out. Then, the call relationships between different containers are matched, and a cross-container level call relationship graph is constructed based on the processed data: nodes can be various services and corresponding containers, and the edges between nodes are cross-container call relationships, forming cross-container level associated log data and the corresponding cross-container level call relationship graph.

[0082] This embodiment's technical solution uses three preset key fields to accurately filter related log data at each level, constructs a targeted call relationship graph, and refines the construction process step by step to ensure that each level of graph accurately represents the corresponding characteristics, comprehensively covers key information in complex call scenarios, effectively solves the shortcomings of traditional single-dimensional anomaly investigation, and improves the accuracy and efficiency of root cause localization of abnormal events.

[0083] Example 3 Figure 3 This is a flowchart of an abnormal event fault location method provided in Embodiment 3 of the present invention. Based on the above embodiments, this embodiment provides a preferred example.

[0084] S301. When an abnormal event in the online interface scheduling process is detected, the link tracing identifier and abnormal event context information of the abnormal event are obtained, and the full log of the online call associated with the abnormal event is obtained based on the link tracing identifier and abnormal event context information. S302. Based on the preset first key field, extract the thread-level associated log content from the full log of online calls to obtain thread-level associated log data, and construct a thread-level call relationship graph based on the thread-level associated log data.

[0085] The thread-level association log data includes thread-level basic association data, thread-level call sequence data, thread-level call relationship data, and thread-level auxiliary association data. Correspondingly, a thread-level call relationship graph is constructed based on the thread-level association log data, including: parsing the thread-level basic association data to obtain at least one execution thread identifier, and using the execution thread unit corresponding to the execution thread identifier as a node in the initial call relationship graph; parsing the thread-level call sequence data and thread-level call relationship data to determine the thread call relationships between execution thread units; constructing the node edges of the initial call relationship graph based on the thread call relationships; and determining the node attribute information of each node and the edge attribute information of each node edge in the initial call relationship graph based on the thread-level auxiliary association data to obtain the thread-level call relationship graph. S303. Based on the preset second key field, extract the container-level associated log content from the full log of online calls to obtain container-level associated log data, and construct a container-level call relationship graph based on the container-level associated log data.

[0086] S304. Based on the preset third key field, extract the cross-container-level related log content from the full log of online calls to obtain cross-container-level related log data, and construct a cross-container-level call relationship graph based on the cross-container-level related log data.

[0087] S305. Perform feature encoding on the thread-level call relationship graph to generate a thread-level encoding vector, perform feature encoding on the container-level call relationship graph to generate a container-level encoding vector, and perform feature encoding on the cross-container-level call relationship graph to generate a cross-container-level encoding vector.

[0088] The process of feature encoding the thread-level call relationship graph to generate a thread-level encoding vector includes: structural encoding of the thread-level call relationship graph to generate a thread-level structure-aware vector; temporal encoding of the thread-level call relationship graph to generate a thread-level temporal relationship vector; and vector fusion of the thread-level structure-aware vector and the thread-level temporal relationship vector to generate a thread-level encoding vector.

[0089] S306. Generate a fused encoding vector based on the thread-level encoding vector, container-level encoding vector, and cross-container-level encoding vector.

[0090] S307. Input the fused coding vector into at least one expert prediction module in the fault prediction model, and have the expert prediction module perform fault root cause analysis based on the fused coding vector to obtain the fault root cause prediction results output by each expert prediction module.

[0091] S308. Input the fault root cause prediction results output by each expert prediction module into the routing decision module in the fault prediction model. The routing decision module then performs result fusion on each fault root cause prediction result according to the current weight parameters corresponding to each expert prediction module to obtain the target fault root cause location result corresponding to the abnormal business processing event.

[0092] The fault prediction model includes at least one expert prediction module and a routing decision module.

[0093] The training method for the fault prediction model is as follows: Historical fusion encoding vectors and their corresponding historical fault root cause locations are obtained for historical time periods, and these historical fault root cause locations are used as sample labels for the historical fusion encoding vectors; the historical fusion encoding vectors with historical fault root cause location sample labels are input into the fault prediction model to obtain the fault prediction results output by each expert prediction module, and the fault prediction results output by each expert prediction module are input into the routing decision module to obtain the predicted fault root cause location output by the routing decision module; based on the historical fault root cause location and the predicted fault root cause location, the initial weight parameters corresponding to each expert prediction module in the routing decision module of the fault prediction model are updated until the preset model training termination condition is met, resulting in a completed fault prediction model.

[0094] The information collected in the above embodiments of the present invention is all information and data authorized by the user or fully authorized by all parties. The collection, storage, use, processing, transmission, provision, disclosure and application of the relevant data all comply with the relevant laws, regulations and standards of the relevant countries and regions, necessary confidentiality measures have been taken, and they do not violate public order and good morals. Corresponding operation entry points are provided for users to choose to authorize or refuse.

[0095] Example 4 Figure 4 This is a schematic diagram of the structure of an abnormal event fault location device provided in Embodiment 4 of the present invention. The abnormal event fault location device provided in this embodiment of the present invention is applicable to abnormal event fault location scenarios in the financial technology field. This abnormal event fault location device can be implemented in hardware and / or software, and can be applied to an abnormal event fault location method. Specifically, it can be configured in a controller, such as... Figure 4 As shown, the device includes: a full log retrieval module 401, a retrieval relationship graph generation module 402, an encoding vector generation module 403, a fused encoding vector generation module 404, and a fault root cause localization module 405. Wherein: The full log call module 401 is used to obtain the link tracing identifier and exception event context information of the business processing exception event when an exception event is detected in the online interface scheduling process, and to obtain the full log of the online call associated with the exception event based on the link tracing identifier and exception event context information.

[0096] The call relationship graph generation module 402 is used to generate thread-level call relationship graphs, container-level call relationship graphs, and cross-container-level call relationship graphs based on the full log of online calls.

[0097] The encoding vector generation module 403 is used to perform feature encoding on the thread-level call relationship graph to generate a thread-level encoding vector, to perform feature encoding on the container-level call relationship graph to generate a container-level encoding vector, and to perform feature encoding on the cross-container-level call relationship graph to generate a cross-container-level encoding vector.

[0098] The fused encoding vector generation module 404 is used to generate a fused encoding vector based on thread-level encoding vectors, container-level encoding vectors, and cross-container-level encoding vectors.

[0099] The fault root cause localization module 405 is used to input the fused encoded vector into the pre-trained fault prediction model, and the fault prediction model performs fault root cause analysis and outputs the target fault root cause localization result of the abnormal event in business processing.

[0100] When an abnormal service processing event is detected during the online interface scheduling process, the technical solution of this invention obtains the link tracing identifier and abnormal event context information of the abnormal service processing event, and obtains the full log of online calls associated with the abnormal service processing event based on the link tracing identifier and abnormal event context information; generates thread-level call relationship graphs, container-level call relationship graphs, and cross-container-level call relationship graphs based on the full log of online calls; performs feature encoding on the thread-level call relationship graphs to generate thread-level encoding vectors, performs feature encoding on the container-level call relationship graphs to generate container-level encoding vectors, and performs feature encoding on the cross-container-level call relationship graphs to generate cross-container-level encoding vectors; generates a fused encoding vector based on the thread-level encoding vector, container-level encoding vector, and cross-container-level encoding vector; inputs the fused encoding vector into a pre-trained fault prediction model, which performs fault root cause analysis and outputs the target fault root cause location result of the abnormal service processing event. It can obtain the full online call log based on link tracing identifiers and abnormal context information, generate thread-level, container-level, and cross-container-level call relationship graphs, and complete feature encoding and vector fusion. Then, through a pre-trained fault prediction model, it can achieve root cause analysis and accurate location of abnormal events, effectively solving the problems of incomplete information coverage and low location efficiency in traditional single-dimensional investigation methods. It improves the processing efficiency and root cause location accuracy of online interface scheduling abnormal events in the financial field, and provides reliable technical support for the stable operation of the system.

[0101] Optionally, the fault prediction model includes at least one expert prediction module and a routing decision module. Correspondingly, the fault root cause localization module 405 is specifically used for: The fused coding vector is input into at least one expert prediction module in the fault prediction model. The expert prediction module performs fault root cause analysis based on the fused coding vector, and the fault root cause prediction results output by each expert prediction module are obtained.

[0102] The root cause prediction results output by each expert prediction module are input into the routing decision module in the fault prediction model. The routing decision module then fuses the root cause prediction results of each expert prediction module according to the current weight parameters of each module to obtain the target fault root cause location result corresponding to the abnormal business processing event.

[0103] In one optional embodiment, the fault prediction model is trained as follows: Obtain the historical fusion coding vectors and their corresponding historical fault root cause locations under historical time periods, and use the historical fault root cause locations as sample labels for the historical fusion coding vectors.

[0104] The historical fusion encoding vector with historical fault root cause localization sample labels is input into the fault prediction model to obtain the fault prediction results output by each expert prediction module. The fault prediction results output by each expert prediction module are then input into the routing decision module to obtain the predicted fault root cause localization output by the routing decision module.

[0105] Based on the historical root cause localization and the predicted root cause localization, the initial weight parameters corresponding to each expert prediction module in the routing decision module of the fault prediction model are updated until the preset model training termination condition is met, and the trained fault prediction model is obtained.

[0106] Optionally, the relationship diagram generation module 402 is invoked, including: The thread-level call relationship graph construction unit extracts thread-level related log content from the full online call log based on the preset first key field, obtains thread-level related log data, and constructs a thread-level call relationship graph based on the thread-level related log data.

[0107] The container-level call relationship graph construction unit extracts container-level associated log content from the full log of online calls based on the preset second key field, obtains container-level associated log data, and constructs a container-level call relationship graph based on the container-level associated log data.

[0108] The cross-container-level call relationship graph construction unit extracts cross-container-level related log content from the full log of online calls based on the preset third key field, obtains cross-container-level related log data, and constructs a cross-container-level call relationship graph based on the cross-container-level related log data.

[0109] Optionally, thread-level correlation log data includes thread-level basic correlation data, thread-level call sequence data, thread-level call relationship data, and thread-level auxiliary correlation data. Correspondingly, the thread-level call relationship graph construction unit is specifically used for: The thread-level basic association data is parsed to obtain at least one execution thread identifier, and the execution thread unit corresponding to the execution thread identifier is used as a node in the initial call relationship graph.

[0110] The thread-level call timing data and thread-level call relationship data are parsed to determine the thread call relationship between execution thread units.

[0111] Based on the thread call relationships, construct the node edges of the initial call relationship graph.

[0112] Based on the thread-level auxiliary association data, the node attribute information of each node and the edge attribute information of each node in the initial call relationship graph are determined, and the thread-level call relationship graph is obtained.

[0113] Optionally, the encoding vector generation module 403 is specifically used for: The thread-level call relationship graph is structurally encoded to generate a thread-level structure-aware vector.

[0114] Perform time-series encoding on the thread-level call relationship graph to generate a thread-level time-series relationship vector.

[0115] The thread-level structure awareness vector and the thread-level temporal relationship vector are fused to generate the thread-level encoding vector.

[0116] The abnormal event fault location device provided in the embodiments of the present invention can execute an abnormal event fault location method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.

[0117] Example 5 Figure 5 A schematic diagram of an electronic device 50 that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0118] like Figure 5As shown, the electronic device 50 includes at least one processor 51 and a memory, such as a read-only memory (ROM) 52 and a random access memory (RAM) 53, communicatively connected to the at least one processor 51. The memory stores computer programs executable by the at least one processor. The processor 51 can perform various appropriate actions and processes based on the computer program stored in the ROM 52 or loaded from storage unit 58 into the RAM 53. The RAM 53 can also store various programs and data required for the operation of the electronic device 50. The processor 51, ROM 52, and RAM 53 are interconnected via a bus 54. An input / output (I / O) interface 55 is also connected to the bus 54.

[0119] Multiple components in electronic device 50 are connected to I / O interface 55, including: input unit 56, such as keyboard, mouse, etc.; output unit 57, such as various types of monitors, speakers, etc.; storage unit 58, such as disk, optical disk, etc.; and communication unit 59, such as network card, modem, wireless transceiver, etc. Communication unit 59 allows electronic device 50 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0120] Processor 51 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 51 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 51 performs the various methods and processes described above, such as an abnormal event fault location method.

[0121] In some embodiments, an abnormal event fault location method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 58. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 50 via ROM 52 and / or communication unit 59. When the computer program is loaded into RAM 53 and executed by processor 51, one or more steps of the abnormal event fault location method described above may be performed. Alternatively, in other embodiments, processor 51 may be configured as an abnormal event fault location method by any other suitable means (e.g., by means of firmware).

[0122] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0123] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0124] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0125] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0126] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0127] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.

[0128] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0129] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A method for locating faults in abnormal events, characterized in that, include: When an abnormal event in the online interface scheduling process is detected, the link tracing identifier and abnormal event context information of the abnormal event are obtained, and based on the link tracing identifier and the abnormal event context information, the full log of the online call associated with the abnormal event is obtained. Based on the full log of online calls, generate thread-level call relationship graphs, container-level call relationship graphs, and cross-container-level call relationship graphs. The thread-level call relationship graph is feature-encoded to generate a thread-level encoding vector; the container-level call relationship graph is feature-encoded to generate a container-level encoding vector; and the cross-container-level call relationship graph is feature-encoded to generate a cross-container-level encoding vector. Based on the thread-level encoding vector, container-level encoding vector, and cross-container-level encoding vector, a fused encoding vector is generated; The fused encoding vector is input into a pre-trained fault prediction model, which performs fault root cause analysis and outputs the target fault root cause localization result of the abnormal business processing event.

2. The method according to claim 1, characterized in that, The fault prediction model includes at least one expert prediction module and a routing decision module; correspondingly, the step of inputting the fused encoding vector into the pre-trained fault prediction model, and having the fault prediction model perform fault root cause analysis to output the target fault root cause localization result of the abnormal business processing event includes: The fused coding vector is input into at least one expert prediction module in the fault prediction model, and the expert prediction module performs fault root cause analysis based on the fused coding vector to obtain the fault root cause prediction results output by each expert prediction module. The root cause prediction results output by each of the expert prediction modules are input into the routing decision module in the fault prediction model. The routing decision module then fuses the root cause prediction results according to the current weight parameters corresponding to each of the expert prediction modules to obtain the target fault root cause location result corresponding to the abnormal business processing event.

3. The method according to claim 2, characterized in that, The training method for the fault prediction model is as follows: Obtain the historical fusion coding vector and its corresponding historical fault root cause location under the historical time period, and use the historical fault root cause location as the sample label of the historical fusion coding vector; The historical fusion encoding vector with historical fault root cause localization sample labels is input into the fault prediction model to obtain the fault prediction results output by each of the expert prediction modules. The fault prediction results output by each of the expert prediction modules are then input into the routing decision module to obtain the predicted fault root cause localization output by the routing decision module. Based on the historical fault root cause localization and the predicted fault root cause localization, the initial weight parameters corresponding to each expert prediction module in the routing decision module of the fault prediction model are updated until the preset model training termination condition is met, and the trained fault prediction model is obtained.

4. The method according to claim 1, characterized in that, The step of generating thread-level call relationship graphs, container-level call relationship graphs, and cross-container-level call relationship graphs based on the full log of online calls includes: Based on a preset first key field, thread-level related log content is extracted from the full online call log to obtain thread-level related log data, and a thread-level call relationship graph is constructed based on the thread-level related log data; and, Based on the preset second key field, container-level associated log content is extracted from the full online call log to obtain container-level associated log data, and a container-level call relationship graph is constructed based on the container-level associated log data; and, Based on the preset third key field, the cross-container-level associated log content is extracted from the full log of the online call to obtain cross-container-level associated log data, and a cross-container-level call relationship graph is constructed based on the cross-container-level associated log data.

5. The method according to claim 4, characterized in that, The thread-level association log data includes thread-level basic association data, thread-level call sequence data, thread-level call relationship data, and thread-level auxiliary association data; correspondingly, the step of constructing a thread-level call relationship graph based on the thread-level association log data includes: The thread-level basic association data is parsed to obtain at least one execution thread identifier, and the execution thread unit corresponding to the execution thread identifier is used as a node in the initial call relationship graph; The thread-level call timing data and thread-level call relationship data are parsed to determine the thread call relationships between the execution thread units; Based on the thread call relationships, construct the node edges of the initial call relationship graph; Based on the thread-level auxiliary association data, the node attribute information of each node and the edge attribute information of each node in the initial call relationship graph are determined to obtain the thread-level call relationship graph.

6. The method according to claim 1, characterized in that, The step of performing feature encoding on the thread-level call relationship graph to generate a thread-level encoding vector includes: The thread-level call relationship graph is structurally encoded to generate a thread-level structure-aware vector; and, The thread-level call relationship graph is time-encoded to generate a thread-level time relationship vector; The thread-level structure perception vector and the thread-level temporal relationship vector are fused to generate a thread-level encoding vector.

7. An abnormal event fault location device, characterized in that, include: The full log call module is used to obtain the link tracing identifier and exception event context information of the business processing exception event when an exception event is detected in the online interface scheduling process, and to obtain the full log of the online call associated with the exception event based on the link tracing identifier and the exception event context information. The call relationship graph generation module is used to generate thread-level call relationship graphs, container-level call relationship graphs, and cross-container-level call relationship graphs based on the full log of online calls. The encoding vector generation module is used to perform feature encoding on the thread-level call relationship graph to generate a thread-level encoding vector, to perform feature encoding on the container-level call relationship graph to generate a container-level encoding vector, and to perform feature encoding on the cross-container-level call relationship graph to generate a cross-container-level encoding vector. The fusion encoding vector generation module is used to generate a fusion encoding vector based on the thread-level encoding vector, container-level encoding vector, and cross-container-level encoding vector. The fault root cause localization module is used to input the fused encoding vector into the pre-trained fault prediction model, and the fault prediction model performs fault root cause analysis to output the target fault root cause localization result of the abnormal business processing event.

8. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform an abnormal event fault location method according to any one of claims 1-6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that are used to cause a processor to execute the fault location method for any one of claims 1-6.

10. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements an abnormal event fault location method according to any one of claims 1-6.