A system error processing method and device, electronic equipment and storage medium

By acquiring monitoring data and call relationship information of system errors, and using information analysis models to automatically analyze the causes of errors, the problem of low efficiency in system error handling in existing technologies is solved, and efficient error handling is achieved.

CN114398200BActive Publication Date: 2026-02-27AGRICULTURAL BANK OF CHINA
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202210064233.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-20
Publication Date
2026-02-27
Estimated Expiration
2042-01-20

AI Technical Summary

Technical Problem

Existing technologies have low error handling efficiency, requiring manual troubleshooting of error causes, which is inefficient.

Method used

By acquiring monitoring data from multiple systems, including transaction identifiers and call relationship information, the faulty system is identified based on the call relationship information. The error information is then input into a pre-trained information analysis model, which utilizes multiple analysis sub-models and fusion sub-models to perform automated analysis of the error causes.

Benefits of technology

It enables accurate acquisition and location of error messages, automated analysis of error causes, and improved error handling efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114398200B_ABST
    Figure CN114398200B_ABST
Patent Text Reader

Abstract

The application discloses a system error processing method and device, electronic equipment and a storage medium. The method comprises the following steps: in the case of detecting a system error, acquiring monitoring data of multiple systems, wherein the monitoring data comprises transaction identification and call relationship information; determining a fault system based on the call relationship information, and acquiring error information corresponding to the transaction identification in the fault system; inputting the error information into a pre-trained information analysis model to obtain error analysis information, wherein the information analysis model comprises multiple analysis sub-models and a fusion sub-model, the analysis sub-models are used for predicting error cause information of the fault system, and the fusion sub-model is used for fusing the output error cause information of the multiple analysis sub-models to obtain the error analysis information. By using the above technical scheme, the automatic analysis of error causes is realized through the information analysis model, and the error processing efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of troubleshooting, and in particular to a system error processing method and device, electronic equipment and storage medium. BACKGROUND

[0002] In today's rapid development of digital transformation in various industries, the number of internal systems in enterprises is gradually increasing, the division of systems is gradually refined, and the calling relationship between systems is becoming more and more complex.

[0003] In the prior art, the whole life cycle process of a business involves as many as ten systems. These mutually called systems form a system group. In the system group, each system is generally developed by different departments within the enterprise, each system has its own data format and is isolated from each other, forming a data island.

[0004] At present, once a customer initiates a business on the system terminal of the system group and an error occurs, the relevant personnel of each system need to manually troubleshoot the error reason, which is low in efficiency. SUMMARY

[0005] The present application provides a system error processing method and device, electronic equipment and storage medium to solve the problem of system error processing efficiency and improve error processing efficiency.

[0006] According to an aspect of the present application, a system error processing method is provided, comprising:

[0007] In the case of detecting a system error, monitoring data of multiple systems is obtained, wherein the monitoring data includes transaction identification and calling relationship information;

[0008] Based on the calling relationship information, a fault system is determined, and in the fault system, error information corresponding to the transaction identification is obtained;

[0009] The error information is input into a pre-trained information analysis model to obtain error analysis information, wherein the information analysis model includes multiple analysis sub-models and a fusion sub-model, the analysis sub-model is used to predict error reason information of the fault system, and the fusion sub-model is used to fuse the output error reason information of the multiple analysis sub-models to obtain the error analysis information.

[0010] According to another aspect of the present application, a system error processing device is provided, comprising:

[0011] A monitoring data acquisition module is configured to execute, in the case of detecting a system error, acquire monitoring data of multiple systems, wherein the monitoring data includes transaction identification and calling relationship information;

[0012] An error information obtaining module is configured to determine a fault system based on the call relationship information, and obtain error information corresponding to the transaction identifier in the fault system.

[0013] An analysis model processing module is configured to input the error information into a pre-trained information analysis model to obtain error analysis information, wherein the information analysis model comprises a plurality of analysis sub-models and a fusion sub-model, the analysis sub-models are configured to predict error cause information of the fault system, and the fusion sub-model is configured to fuse the error cause information output by the plurality of analysis sub-models to obtain the error analysis information.

[0014] According to another aspect of the present application, an electronic device is provided, which comprises:

[0015] at least one processor; and

[0016] a memory connected to the at least one processor in communication; wherein

[0017] The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to perform the system error processing method according to any one of the embodiments of the present application.

[0018] According to another aspect of the present application, a computer readable storage medium is provided, which stores computer instructions for enabling a processor to perform the system error processing method according to any one of the embodiments of the present application when executed.

[0019] The technical solution of the embodiments of the present application achieves accurate acquisition of monitoring data by acquiring monitoring data of multiple systems in the case of detecting system error, wherein the monitoring data comprises a transaction identifier and call relationship information; achieves positioning and acquisition of error information by determining a fault system based on the call relationship information and acquiring error information corresponding to the transaction identifier in the fault system; and achieves automatic analysis of error causes by inputting the error information into a pre-trained information analysis model to obtain error analysis information, wherein the information analysis model comprises a plurality of analysis sub-models and a fusion sub-model, the analysis sub-models are configured to predict error cause information of the fault system, and the fusion sub-model is configured to fuse the error cause information output by the plurality of analysis sub-models to obtain the error analysis information, thereby solving the problem of low error processing efficiency and improving the error processing efficiency.

[0020] It should be understood that the content described in this part is not intended to identify key or important features of the embodiments of the present application, nor to limit the scope of the present application. Other features of the present application will become apparent from the following description. Attached Figure Description

[0021] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0022] Figure 1 This is a flowchart of a system error handling method provided in Embodiment 1 of the present invention;

[0023] Figure 2 This is a flowchart of a system error handling method provided in Embodiment 2 of the present invention;

[0024] Figure 3 This is a flowchart of a system error handling method provided in Embodiment 3 of the present invention;

[0025] Figure 4 This is a flowchart of a system error handling method provided in Embodiment 4 of the present invention;

[0026] Figure 5 This is a schematic diagram of an intelligent error analysis model architecture provided in Embodiment 4 of the present invention;

[0027] Figure 6 This is a schematic diagram of the structure of a system error processing device according to Embodiment 5 of the present invention;

[0028] Figure 7 This is a schematic diagram of the structure of an electronic device that implements the system error handling method of the present invention. Detailed Implementation

[0029] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0030] It should be noted that the terms "first", "second", and the like in the description and claims of the application and the above drawings are used to distinguish similar objects, and do not necessarily indicate a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the application described herein can be implemented in other than the order illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to include only those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0031] Embodiment one

[0032] Figure 1 A flowchart of a system error processing method is provided for the first embodiment of the application. The present embodiment can be applied to the automatic troubleshooting of system groups. The method can be executed by a system error processing device, which can be implemented in the form of hardware and / or software. The system error processing device can be configured in a terminal that monitors multiple systems. As shown in FIG. 1, the method comprises the following steps. Figure 1

[0033] S110, in the case of detecting a system error, acquiring monitoring data of multiple systems, wherein the monitoring data comprises transaction identification and call relationship information.

[0034] The multiple systems refer to a system group comprising multiple systems. The systems can call each other. The call relationship information of the systems and between the systems collectively constitutes a system group. Typically, the multiple systems can be internal systems of an enterprise. The monitoring data refers to data acquired during the monitoring of the multiple systems. The monitoring data can include but is not limited to transaction identification and call relationship information. The transaction identification refers to a unique identification generated when each transaction is initiated. The transaction identification can be passed to other related systems along with the transaction data. By setting the transaction identification, the cross-system transaction identification and the data island problem between systems can be effectively solved. The call relationship information refers to the call relationship chain between the systems, which can be a chain structure. By tracking the call relationship information, the fault system can be quickly located. Optionally, the call relationship information can be the gateway relationship of the systems.

[0035] By way of example, the terminal that monitors multiple systems can monitor the systems in real time. When an error occurs in the systems, the terminal acquires the transaction identification and call relationship information of the multiple systems to perform troubleshooting analysis based on the transaction identification and call relationship information, and obtains the error cause to achieve automatic error processing.

[0036] ​On the basis of the above-mentioned embodiments, the monitoring data of the multiple systems is acquired, including: acquiring the monitoring data of the multiple systems based on a preset time interval; or acquiring the monitoring data of the multiple systems based on a preset trigger event.

[0037] Specifically, in some embodiments, the polling mode can be adopted to acquire the monitoring data of the multiple systems, that is, the abnormal log table of the terminal system is polled every certain time interval to query whether there is an abnormal record. In some embodiments, the event trigger mode is adopted to acquire the monitoring data of the multiple systems, that is, the abnormal monitoring module is embedded in the terminal system, and if an abnormal event occurs, the monitoring data of the multiple systems is actively triggered to be acquired, which can effectively improve the response speed of error reporting.

[0038] S120, determining a fault system based on the call relationship information, in which the error information corresponding to the transaction identifier is acquired.

[0039] The fault system refers to a system in which an error occurs in the multiple systems, and the fault system can be determined through the call relationship information, that is, the fault system is determined through the call relationship chain between the systems. The error information refers to the information recorded by the fault system and containing the error content, which can be matched in the fault system through the transaction identifier to acquire the error information, for example, the error information can be recorded in the fault system log.

[0040] On the basis of the above-mentioned embodiments, the fault system is determined based on the call relationship information, including: determining the system at the tail of the call chain structure in the call relationship information; and determining the system at the tail of the call chain structure as the fault system.

[0041] The call relationship information includes but is not limited to the calling system number and the system number of the called system, and the calling system number and the system number of the called system can be used to establish the call chain structure of each system. By tracking the call chain structure, the system at the tail of the call chain structure can be determined, and the system can be determined as the fault system.

[0042] On the basis of the above-mentioned embodiments, the fault system includes an error log; correspondingly, the error information corresponding to the transaction identifier is acquired, including: acquiring the error log of the fault system; and extracting the error information corresponding to the transaction identifier from the error log.

[0043] Specifically, the error information corresponding to the transaction identifier can be extracted from the error log of the fault system, wherein the error log can be a log table recording system errors. The error information can include two types of information, one of which can be information identified in advance by the user and actively reported, and the other of which can be error stack information thrown by the system. The two types of information can be recorded in the error log for the user to retrieve.

[0044] S130, input the error information into the information analysis model pre-trained to obtain error analysis information, wherein the information analysis model includes a plurality of analysis sub-models and a fusion sub-model, the analysis sub-models are used to predict error cause information of the fault system, and the fusion sub-model is used to fuse the output error cause information of the plurality of analysis sub-models to obtain the error analysis information.

[0045] The information analysis model can be a machine learning model for predicting and analyzing error causes. The information analysis model can include a plurality of analysis sub-models and a fusion sub-model. The number of analysis sub-models can be two or more. The analysis sub-models can be used to predict error cause information of the fault system. The fusion sub-model is used to fuse the output error cause information of the plurality of analysis sub-models to obtain error analysis information. The error analysis information is the fused error cause, which is unique.

[0046] It can be understood that the information analysis model can generate a plurality of error cause information, and fuse each error cause information through the fusion sub-model to generate unique error analysis information. This can avoid the situation that a single analysis model has poor prediction accuracy of error causes due to performance differences, and effectively improve the prediction accuracy of error causes.

[0047] The embodiment of the application provides a system error processing method. By detecting system errors, the monitoring data of multiple systems is obtained, wherein the monitoring data includes transaction identifiers and call relationship information, accurate monitoring data is obtained; the fault system is determined based on the call relationship information, the error information corresponding to the transaction identifier is obtained in the fault system, and the error information is positioned and obtained; the error information is input into the information analysis model pre-trained to obtain error analysis information, wherein the information analysis model includes a plurality of analysis sub-models and a fusion sub-model, the analysis sub-models are used to predict error cause information of the fault system, and the fusion sub-model is used to fuse the output error cause information of the plurality of analysis sub-models to obtain error analysis information, automatic analysis of error causes is realized, thereby solving the problem of low error processing efficiency, and improving the error processing efficiency.

[0048] Embodiment two

[0049] Figure 2A flowchart of a system error processing method provided for the second embodiment of the present application, which can be combined with the various optional solutions in the above embodiments. In the second embodiment of the present application, optionally, the inputting of the error information into the pre-trained information analysis model to obtain error analysis information includes: performing word segmentation processing on the error information to obtain error segmented information; performing cleaning on the error segmented information to obtain error cleaned information; performing feature extraction on the error cleaned information, and inputting the extracted feature information into each analysis sub-model respectively to obtain error cause information corresponding to each analysis sub-model; and performing fusion of the error cause information output by each analysis sub-model through the fusion sub-model to obtain error analysis information.

[0050] As shown in Figure 2 , the method includes:

[0051] S210, in the case of detecting a system error, obtaining monitoring data of multiple systems, wherein the monitoring data includes transaction identification and call relationship information.

[0052] S220, determining a fault system based on the call relationship information, and obtaining error information corresponding to the transaction identification in the fault system.

[0053] S230, performing word segmentation processing on the error information to obtain error segmented information.

[0054] The word segmentation processing refers to segmenting the error information in units of words to obtain multiple error segmented information. The specific method of word segmentation processing can include a dictionary-based word segmentation algorithm and a statistical-based word segmentation algorithm. In the second embodiment of the present application, the method of word segmentation processing is not limited, for example, the jieba word segmentation library can be called for word segmentation processing.

[0055] S240, performing cleaning on the error segmented information to obtain error cleaned information.

[0056] Specifically, the cleaning of the error segmented information can remove useless words such as virtual words, and only keep the key words associated with the error, that is, the error cleaned information can be multiple error key words cleaned. The cleaning method can specifically be to establish a key word list, compare the error segmented information with the key word list, and remove the same words in the error segmented information as the key word list, so that the obtained error cleaned information is more reliable.

[0057] S250, performing feature extraction on the error cleaned information, and inputting the extracted feature information into each analysis sub-model respectively to obtain error cause information corresponding to each analysis sub-model.

[0058] It should be noted that the number of feature extraction methods is the same as the number of analysis sub-models, and is also multiple, each feature extraction method can be randomly used with each analysis sub-model, or can be fixedly used, and the embodiment does not limit this. In the embodiment of the application, the extracted feature information is input into each analysis sub-model, and the error reason information corresponding to each analysis sub-model can be obtained, that is, multiple error reason information can be obtained, which can avoid the situation that the prediction error reason accuracy is poor due to the performance difference of a single analysis model, and effectively improves the prediction error reason accuracy.

[0059] On the basis of the above embodiment, the analysis sub-model includes at least two types of models; correspondingly, the feature extraction of the error cleaning information and the input of the extracted feature information into each analysis sub-model to obtain the error reason information corresponding to each analysis sub-model includes: based on the frequency information of each keyword in the error cleaning information, the weight value information corresponding to each keyword is obtained by weight value calculation, and the first feature vector is constructed according to the weight value information corresponding to each keyword, and the first feature vector is input into the first analysis sub-model to obtain the error reason information corresponding to the first analysis sub-model; based on the context information of each keyword in the error cleaning information, the second feature vector is obtained by vector conversion, and the second feature vector is input into the second analysis sub-model to obtain the error reason information corresponding to the second analysis sub-model; each keyword in the error cleaning information is encoded to obtain error encoding information, and the third feature vector is obtained by feature extraction of the error encoding information based on the bag-of-words model, and the third feature vector is input into the third analysis sub-model to obtain the error reason information corresponding to the third analysis sub-model.

[0060] Among them, the analysis sub-model can include multiple types of models, and different types of models correspond to different prediction strategies, so that the information analysis model can obtain multiple error reason information for fusion, and improve the reliability and accuracy of the prediction result.

[0061] In the embodiment of the application, typically, the number of analysis sub-models can include three, which are a first analysis sub-model, a second analysis sub-model and a third analysis sub-model, and the types of each analysis sub-model are different. The input of the first analysis sub-model is the first feature vector, which can be obtained by calculating the weight value of the frequency information of each keyword in the error cleaning information, and then constructing according to the weight value corresponding to each keyword. The input of the second analysis sub-model is the second feature vector, which can be obtained by vector conversion according to the context information of each keyword in the error cleaning information. The input of the third analysis sub-model is the third feature vector, which can be obtained by encoding each keyword in the error cleaning information, and then extracting features of the error encoding information according to the bag-of-words model.

[0062] S260. The error cause information output by each analysis sub-model is fused through the fusion sub-model to obtain error analysis information.

[0063] For example, error analysis information can be the result of fusing multiple error cause information. The error cause information can include a first error cause, a second error cause, and a third error cause. By inputting the three error causes into the fusion sub-model and performing operations such as weighting or voting on each error cause, a unique error analysis information can be obtained, which can effectively improve the accuracy of error analysis information.

[0064] This invention provides a system error handling method. By segmenting error information into words to obtain error segmentation information, and then cleaning this segmentation information to obtain cleaned error information, the method ensures the cleaned error information is cleaner and more reliable. Furthermore, feature extraction is performed on the cleaned error information, and the extracted feature information is input into each analysis sub-model to obtain the error cause information corresponding to each analysis sub-model. A fusion sub-model is then used to fuse the error cause information output from each analysis sub-model, enabling comparative analysis of multiple error cause information to obtain more accurate error analysis information. Moreover, the analysis sub-models and the fusion sub-model achieve automated error cause analysis, thereby solving the problem of low error handling efficiency and improving overall error handling efficiency.

[0065] Example 3

[0066] Figure 3 This is a flowchart of a system error handling method provided in Embodiment 3 of the present invention. The embodiments of the present invention can be combined with the various optional solutions in the above embodiments. Optionally, in this embodiment, after inputting the error information into a pre-trained information analysis model to obtain error analysis information, the method further includes: sending the error analysis information to a target terminal, wherein the target terminal is the device used by the person in charge of the faulty system.

[0067] like Figure 3 As shown, the method includes:

[0068] S310. In the event of a system error, acquire monitoring data from multiple systems, wherein the monitoring data includes transaction identifiers and call relationship information.

[0069] S320. Based on the call relationship information, determine the faulty system, and in the faulty system, obtain the error information corresponding to the transaction identifier.

[0070] S330. Input the error message into the pre-trained information analysis model to obtain error analysis information.

[0071] S340, send the error analysis information to a target terminal, wherein the target terminal is a device used by a responsible person corresponding to the fault system.

[0072] In the embodiment of the application, the target terminal can be a device used by a responsible person corresponding to the fault system, and can include but is not limited to a mobile phone, a computer, and the like.

[0073] For example, the error analysis information can be sent to the target terminal by means of WeChat, short message, or email, so that the response speed of system error can be improved, the economic loss caused by system error can be reduced, and user experience can be improved.

[0074] The embodiment of the application provides a system error processing method, which can improve the response speed of system error, reduce the economic loss caused by system error, and improve user experience by inputting error information into a pre-trained information analysis model to obtain error analysis information, automatically analyzing error causes, and sending the error analysis information obtained by analysis to a device used by a responsible person corresponding to a fault system.

[0075] Embodiment Four

[0076] Figure 4 A flowchart of a system error processing method provided in the fourth embodiment of the application is provided, and the embodiment is a preferred example of the above-mentioned embodiment. It should be noted that the system group unique identifier in the embodiment is the transaction identifier in the above-mentioned embodiment, the calling chain, the information header, and the information body are the calling relationship information in the above-mentioned embodiment, the error system is the fault system in the above-mentioned embodiment, the error cause is the error analysis information in the above-mentioned embodiment, and the multi-model fusion machine learning algorithm is the information analysis model in the above-mentioned embodiment.

[0077] As shown in Figure 4 , the method comprises the following steps:

[0078] Step 1: terminal system error.

[0079] The terminal system refers to a terminal system of a monitoring system group. The terminal system and the customer can directly interact, and the terminal system is an entry of a transaction initiation. If an error occurs, the terminal system can be used as a starting point for error troubleshooting.

[0080] Step 2: trigger an automatic processing program.

[0081] In this embodiment, the monitoring of the terminal system adopts real-time monitoring and active triggering mode. The active triggering mode includes but is not limited to two schemes of intrusive and non-intrusive. For the non-intrusive scheme, a polling mode is adopted, and the abnormal log table is polled every certain time interval to query whether there is an abnormal record. For the intrusive scheme, an event triggering mode is adopted, and an abnormal monitoring module is embedded in the terminal system, and once an abnormality occurs, the error handling is triggered actively.

[0082] Step 4: Obtain the transaction identifier.

[0083] Step 5: Locate the single transaction error business log.

[0084] Step 6: Analyze the log header information.

[0085] Step 7: Determine whether it is an error of the system.

[0086] In the system group, the mutual call between systems forms a complex call network. The terminal system error can be caused by any related system. The system error is divided into two categories: one is the error caused by the called system when calling between systems. The other is the error caused by the system itself. In order to facilitate the error system to be checked, the embodiment specifies the format of the error information when calling between systems. The error information is composed of information header and information body. The information header includes three parts of the system number of the calling system, the system number of the called system and the unique identifier of the system group; the message body contains the content of the error information. According to the calling relationship, the system error forms a chain structure. Tracking the call chain can quickly locate the error system (the system at the tail of the chain is the error system).

[0087] Step 8: Obtain the transaction identifier.

[0088] In this embodiment, after locating the error system, the single transaction causing the error is located through the transaction identifier.

[0089] Specifically, different systems in an enterprise are responsible for different departments, each system is isolated from each other, and each system has its own log recording mode, and the data of each system forms a data island. In order to break through the data island, the present application proposes a unique identifier of the system group. A unique identifier is generated when each transaction is initiated in the terminal, and is transferred to the related system as the unique identifier of the transaction in the system group, so as to realize the identification of cross-system transaction. The unique identifier of the system group can break through the data island between systems without changing the data recording mode of each system. Through the unique identifier of the system group, the error transaction in the error system is uniquely located.

[0090] Step 9: Obtain the log body content.

[0091] Specifically, after locating the error reporting system and the corresponding error reporting transaction, the detailed information of the error reporting is extracted from the error log table of the system. The error reporting information is divided into two categories: one is the information identified in advance by the developer and actively reported, and the other is the error stack information thrown by the system.

[0092] Step 10: Intelligent analysis of error content.

[0093] Specifically, as shown in Figure 5 , first, the error reporting information is subjected to word segmentation processing, and useless words such as virtual words are removed, and only key words are retained. Then, TF-IDF, word2vec, and oneHot encoding are used to extract features, and XGboost model, LightGBM model, and CatBoost model are trained respectively, and finally, stacking integration algorithm is used to integrate the results of the three models. Compared with single model algorithm, the algorithm based on multi-model fusion can effectively improve the accuracy and generalization performance of the algorithm.

[0094] Step 11: Obtain error cause.

[0095] Step 12: Inform system administrator.

[0096] Specifically, based on the message sending module, the relevant system administrator and the relevant contact person are informed in time through the mobile phone, the email and the like.

[0097] The system error processing method provided by the embodiment of the application, compared with the original manual error elimination, realizes the feedback of the error by the customer to the active discovery of the error by the system, the manual investigation of the error by the cross-department system to the chain automatic positioning of the system, the analysis of the error cause by the professional personnel to the intelligent cause analysis, enhances the system error response capability, achieves the active monitoring, timely response, automatic positioning, intelligent analysis, timely notification, and improves the user experience.

[0098] Embodiment five

[0099] Figure 6 A structural schematic diagram of a system error processing device provided for the embodiment five of the application.

[0100] As shown in Figure 6 , the device comprises:

[0101] The monitoring data acquisition module 510 is configured to acquire monitoring data of multiple systems when a system error is detected, wherein the monitoring data comprises transaction identification and call relationship information.

[0102] The error information acquisition module 520 is configured to determine a fault system based on the call relationship information, and acquire error information corresponding to the transaction identification in the fault system.

[0103] The analysis model processing module 530 is configured to input the error information into a pre-trained information analysis model to obtain error analysis information, wherein the information analysis model comprises a plurality of analysis sub-models and a fusion sub-model, the analysis sub-models are configured to predict error cause information of the fault system, and the fusion sub-model is configured to fuse the error cause information output by the plurality of analysis sub-models to obtain the error analysis information.

[0104] The embodiment of the present application provides a system error processing device, which realizes accurate acquisition of monitoring data by acquiring monitoring data of multiple systems in the case that a system error is detected, wherein the monitoring data comprises transaction identification and call relationship information; determines a fault system based on the call relationship information, acquires error information corresponding to the transaction identification in the fault system, and realizes positioning and acquisition of the error information; inputs the error information into a pre-trained information analysis model to obtain error analysis information, wherein the information analysis model comprises a plurality of analysis sub-models and a fusion sub-model, the analysis sub-models are configured to predict error cause information of the fault system, and the fusion sub-model is configured to fuse the error cause information output by the plurality of analysis sub-models to obtain the error analysis information, thereby realizing automatic analysis of error causes, and solving the problem of low error processing efficiency, and improving the error processing efficiency.

[0105] Optionally, the monitoring data acquisition module 510 can be further configured to:

[0106] acquire the monitoring data of the multiple systems based on a preset time interval; or

[0107] acquire the monitoring data of the multiple systems based on a preset trigger event.

[0108] Optionally, the error information acquisition module 520 can be further configured to:

[0109] determine a system at a tail of a call chain structure in the call relationship information as a fault system.

[0110] Optionally, the fault system comprises an error log; and correspondingly, the error information acquisition module 520 can be further configured to:

[0111] acquire the error log of the fault system;

[0112] extract error information corresponding to the transaction identification from the error log.

[0113] In any optional technical solution in the embodiment of the application, optionally, the analysis model processing module 530 comprises:

[0114] The word segmentation processing unit is configured to perform word segmentation processing on the error reporting information to obtain error reporting word segmentation information.

[0115] The information cleaning unit is configured to perform cleaning on the error reporting word segmentation information to obtain error reporting cleaning information.

[0116] The error analysis unit is configured to perform feature extraction on the error reporting cleaning information, and input the extracted feature information into each analysis sub-model respectively to obtain error cause information corresponding to each analysis sub-model.

[0117] The information fusion unit is configured to perform fusion on the error cause information output by each analysis sub-model through the fusion sub-model to obtain error analysis information.

[0118] In any optional technical solution in the embodiment of the application, optionally, the error analysis unit can also be configured to:

[0119] perform weight value calculation based on frequency information of each keyword in the error reporting cleaning information to obtain weight value information corresponding to each keyword, construct a first feature vector according to the weight value information corresponding to each keyword, and input the first feature vector into a first analysis sub-model to obtain error cause information corresponding to the first analysis sub-model;

[0120] perform vector conversion based on context information of each keyword in the error reporting cleaning information to obtain a second feature vector, and input the second feature vector into a second analysis sub-model to obtain error cause information corresponding to the second analysis sub-model;

[0121] perform encoding on each keyword in the error reporting cleaning information to obtain error encoding information, perform feature extraction on the error encoding information based on a bag-of-words model to obtain a third feature vector, and input the third feature vector into a third analysis sub-model to obtain error cause information corresponding to the third analysis sub-model.

[0122] In any optional technical solution in the embodiment of the application, optionally, the device further comprises:

[0123] The information sending module is configured to send the error analysis information to a target terminal, wherein the target terminal is a device used by a responsible person corresponding to the fault system.

[0124] The system error processing device provided in the embodiment of the application can perform the system error processing method provided in any embodiment of the application, and has the corresponding functional modules and beneficial effects of the performing method.

[0125] Example Six

[0126] Figure 7 A structural diagram of an electronic device 10 that can be used to implement embodiments of the present application is shown. The electronic device is intended to represent various forms of digital computers, such as laptops, desktops, tablets, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular telephones, smart phones, wearable devices (e.g., headsets, glasses, watches, etc.), and other similar computing devices. The components shown here, their connections and relationships, and their functions, are meant to be examples only, and are not meant to limit implementations of the applications described and / or claimed in this document.

[0127] As shown in Figure 7 The electronic device 10 includes at least one processor 11, and a memory, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., connected to the at least one processor 11 in communication, where the memory stores computer programs executable by the at least one processor 11, and the processor 11 can perform various appropriate actions and processes according to the computer programs stored in the read-only memory (ROM) 12 or loaded into the random access memory (RAM) 13 from the storage unit 18. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0128] A plurality of components in the electronic device 10 are connected to the I / O interface 15, including an input unit 16, such as a keyboard, a mouse, etc., an output unit 17, such as various types of displays, speakers, etc., a storage unit 18, such as a magnetic disk, an optical disk, etc., and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices through a computer network, such as the Internet, and / or various telecommunications networks.

[0129] The processor 11 can be various general and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 performs various methods and processes described above, such as the system error reporting processing method.

[0130] In some embodiments, the system error handling method can be implemented as a computer program tangibly embodied in a computer readable storage medium, e.g., storage unit 18. In some embodiments, parts or all of the computer program can be loaded and / or installed onto electronic device 10 via, e.g., ROM 12 and / or communication unit 19. When the computer program is loaded onto RAM 13 and executed by processor 11, one or more steps of the above-described method XXX can be performed. Alternatively, in other embodiments, processor 11 can be configured to perform the system error handling method by other means, e.g., with the aid of firmware.

[0131] Various implementations of the systems and techniques described above can be realized in digital electronic circuitry, integrated circuitry, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on a chip (SOC), a programmable logic device (PLD), a computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.

[0132] Computer programs used to implement the methods of the present application can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the computer program, when executed by the processor of the machine, implements the functions / acts specified in the flowcharts and / or block diagrams. The computer program can be executed entirely on a machine, partially on a machine, partially on a machine and partially on a remote machine or entirely on a remote machine or server.

[0133] In the context of the present application, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. A computer-readable storage medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium can be a machine-readable signal medium. More specific examples of a machine-readable storage medium will include one or more lines of a program of instructions in a transitory signal, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0134] To provide for interaction with a user, the systems and techniques described here can be implemented on an electronic device having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the electronic device. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.

[0135] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), blockchain network, and the Internet.

[0136] The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system, and solves the defects of large management difficulty and weak business scalability in traditional physical hosts and VPS services.

[0137] It should be understood that the various forms of flow shown above can be reordered, additional steps added, or steps deleted. For example, the steps described in the present application can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions of the present application can be achieved, and the present application does not limit this.

[0138] Embodiment seven

[0139] Embodiment seven of the present application also provides a computer readable storage medium comprising computer executable instructions for executing a system error processing method when executed by a computer processor, the method comprising:

[0140] In the case of detecting a system error, monitoring data of multiple systems is obtained, wherein the monitoring data includes transaction identification and call relationship information;

[0141] Based on the call relationship information, a fault system is determined, and in the fault system, error information corresponding to the transaction identification is obtained;

[0142] The error information is input into a pre-trained information analysis model to obtain error analysis information, wherein the information analysis model includes multiple analysis sub-models and a fusion sub-model, the analysis sub-models are used to predict error cause information of the fault system, and the fusion sub-model is used to fuse the output error cause information of the multiple analysis sub-models to obtain the error analysis information.

[0143] The computer readable storage medium of the embodiments of the present application can adopt any combination of one or more computer readable media. The computer readable medium can be a computer readable signal medium or a computer readable storage medium. The computer readable storage medium may, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or apparatus, or any combination thereof. More specific examples (non-exhaustive list) of the computer readable storage medium include an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this document, the computer readable storage medium can be any tangible medium that contains or stores a program that can be used by or in connection with an instruction execution system, apparatus or device.

[0144] The computer readable signal medium can include a data signal propagated in baseband or propagated as a carrier wave, in which computer readable program code is embodied. Such propagated data signals can take a wide variety of forms, including but not limited to electro-magnetic signals, optical signals, or any suitable combination thereof. The computer readable signal medium can also be any computer readable medium that is not a storage medium, that is capable of storing the program for use by or in connection with the instruction execution system, apparatus or device.

[0145] The program code embodied on the computer readable medium can be transmitted using any appropriate medium, including but not limited to wireless, wire line, optical fiber cable, RF, etc., or any suitable combination of the above.

[0146] The computer program code for carrying out operations of the embodiments of the present application can be written in one or more programming languages or combinations of languages including an object oriented programming language such as Java, Smalltalk, C++ or the like and conventional procedural programming languages such as "C" or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).

[0147] The above detailed description does not limit the scope of the application. Various modifications, combinations, sub-combinations and alternatives can be made to the detailed description. Any modification, equivalent replacement and improvement etc. made within the spirit and principle of the application shall be included in the scope of the application.

Claims

1. A system error handling method, characterized in that, include: In the event of a system error, monitoring data from multiple systems is acquired, including transaction identifiers and call relationship information. Based on the call relationship information, the faulty system is identified, and in the faulty system, the error information corresponding to the transaction identifier is obtained; The error information is input into a pre-trained information analysis model to obtain error analysis information. The information analysis model includes multiple analysis sub-models and a fusion sub-model. The analysis sub-models are used to predict the error cause information of the faulty system, and the fusion sub-model is used to fuse the error cause information output by the multiple analysis sub-models to obtain the error analysis information. The step of inputting the error information into a pre-trained information analysis model to obtain error analysis information includes: The error message is segmented into words to obtain the error segmentation information; The error segmentation information is cleaned to obtain error cleanup information; The error cleanup information is subjected to feature extraction, and the extracted feature information is input into each analysis sub-model to obtain the error reason information corresponding to each analysis sub-model; The error reason information output by each analysis sub-model is fused through the fusion sub-model to obtain error analysis information; The analysis sub-model includes at least two types of models; Accordingly, the step of extracting features from the error cleanup information and inputting the extracted feature information into each analysis sub-model to obtain the error reason information corresponding to each analysis sub-model includes: Weights are calculated based on the frequency information of each keyword in the error cleanup information to obtain the weight information corresponding to each keyword. A first feature vector is constructed based on the weight information corresponding to each keyword, and the first feature vector is input into the first analysis sub-model to obtain the error reason information corresponding to the first analysis sub-model. Based on the context information of each keyword in the error cleanup information, a vector transformation is performed to obtain a second feature vector, and the second feature vector is input into the second analysis sub-model to obtain the error reason information corresponding to the second analysis sub-model; Each keyword in the error cleanup information is encoded to obtain error encoding information. Features are extracted from the error encoding information based on the bag-of-words model to obtain a third feature vector. The third feature vector is then input into a third analysis sub-model to obtain the error reason information corresponding to the third analysis sub-model. The first analysis sub-model is the XGboost model; the second analysis sub-model is the LightGBM model; and the third analysis sub-model is the CatBoost model.

2. The method according to claim 1, characterized in that, The acquisition of monitoring data from multiple systems includes: Acquire monitoring data from multiple systems based on preset time intervals; or... Monitoring data from multiple systems is acquired based on preset trigger events.

3. The method according to claim 1, characterized in that, The step of determining the faulty system based on the call relationship information includes: Identify the system at the end of the call chain structure in the call relationship information; The system at the end of the call chain structure is identified as a faulty system.

4. The method according to claim 1, characterized in that, The faulty system includes an error log; Accordingly, obtaining the error information corresponding to the transaction identifier includes: Obtain the error log of the faulty system; Extract the error information corresponding to the transaction identifier from the error log.

5. The method according to claim 1, characterized in that, After inputting the error message into a pre-trained information analysis model to obtain error analysis information, the method further includes: The error analysis information is sent to the target terminal, wherein the target terminal is the device used by the person in charge of the faulty system.

6. A system error processing device, characterized in that, include: The monitoring data acquisition module is used to acquire monitoring data from multiple systems when a system error is detected, wherein the monitoring data includes transaction identifiers and call relationship information; The error information acquisition module is used to execute a fault determination system based on the call relationship information, and to acquire the error information corresponding to the transaction identifier in the fault system; An analysis model processing module is used to input the error information into a pre-trained information analysis model to obtain error analysis information. The information analysis model includes multiple analysis sub-models and a fusion sub-model. The analysis sub-models are used to predict the error cause information of the faulty system, and the fusion sub-model is used to fuse the error cause information output by the multiple analysis sub-models to obtain the error analysis information. The analysis model processing module includes: The word segmentation processing unit is used to perform word segmentation processing on the error information to obtain error segmentation information; The information cleaning unit is used to clean the error segmentation information to obtain error cleanup information; The error analysis unit is used to extract features from the error cleanup information and input the extracted feature information into each analysis sub-model to obtain the error reason information corresponding to each analysis sub-model. The information fusion unit is used to fuse the error cause information output by each analysis sub-model through the fusion sub-model to obtain error analysis information; The error analysis unit can also be used for: Weights are calculated based on the frequency information of each keyword in the error cleanup information to obtain the weight information corresponding to each keyword. A first feature vector is constructed based on the weight information corresponding to each keyword, and the first feature vector is input into the first analysis sub-model to obtain the error reason information corresponding to the first analysis sub-model. Based on the context information of each keyword in the error cleanup information, a vector transformation is performed to obtain a second feature vector, and the second feature vector is input into the second analysis sub-model to obtain the error reason information corresponding to the second analysis sub-model; Each keyword in the error cleanup information is encoded to obtain error encoding information. Based on the bag-of-words model, the error encoding information is used to extract features to obtain a third feature vector. The third feature vector is then input into the third analysis sub-model to obtain the error reason information corresponding to the third analysis sub-model. The first analysis sub-model is the XGboost model; the second analysis sub-model is the LightGBM model; and the third analysis sub-model is the CatBoost model.

7. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the system error handling method according to any one of claims 1-5.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that are used to cause a processor to execute the system error handling method according to any one of claims 1-5.

Citation Information

Patent Citations

  • Method and device for processing abnormal logs of operating system based on deep learning

    CN111104242A

  • Transaction fault detection method and device, computing equipment and medium

    CN111611100A

  • Service fault positioning method and device based on machine learning and text classification

    CN113094198A