Abnormal data processing method and device, electronic equipment and computer readable medium

By performing abnormal detection and screening of log data, combined with large language model analysis, the problem of inefficiency of traditional root cause positioning methods in complex systems is solved, and fast and accurate fault positioning and abnormal data analysis are achieved.

CN119961027APending Publication Date: 2025-05-09BEIJING WODONG TIANJUN INFORMATION TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311474382.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-07
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

Traditional artificial or rules-based root cause positioning methods are difficult to effectively handle abnormal data in complex systems, especially in intelligent operation and maintenance scenarios where data is growing and system differences are large.

Method used

By performing exception detection on the acquired log data, exception keywords are obtained, and the target call stack is filtered out from the log data based on these keywords, the root cause analysis results are determined, and the cause of the exception is finally obtained. This method can combine pre-trained large language models for in-depth analysis of the causes of abnormalities.

Benefits of technology

Effectively locate the root cause of abnormality, assist in fast and accurate positioning of faults, improve the efficiency of abnormal data analysis, and is suitable for intelligent operation and maintenance scenarios in complex systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119961027A_ABST
    Figure CN119961027A_ABST
Patent Text Reader

Abstract

The invention provides an abnormal data processing method and device. A specific embodiment of the method comprises the steps of performing anomaly detection on acquired log data to obtain an anomaly keyword; based on the abnormal keyword, screening out a target call stack from the log data; determining a root cause analysis result based on the target call stack; and obtaining an abnormal cause based on the target call stack and the root cause analysis result. According to the embodiment, the accuracy of root cause positioning can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technology, and in particular to an abnormal data processing method and device, an electronic device, and a computer-readable medium. Background Art

[0002] In intelligent operation and maintenance scenarios, artificial intelligence is generally applied to the operation and maintenance field. Based on existing operation and maintenance data, machine learning is used to further solve problems that cannot be solved by automated operation and maintenance.

[0003] Due to the constraints of manpower, experience and other conditions, as well as the differences between different systems and the complexity of locating the root cause of anomalies, traditional manual or rule-based root cause location methods are no longer applicable to the current complex systems with growing data. Summary of the invention

[0004] The embodiments of the present disclosure provide an abnormal data processing method and device, an electronic device, and a computer-readable medium.

[0005] In a first aspect, an embodiment of the present disclosure provides a method for processing abnormal data, the method comprising: performing anomaly detection on acquired log data to obtain abnormal keywords; based on the abnormal keywords, filtering out a target call stack from the log data; based on the target call stack, determining a root cause analysis result; based on the target call stack and the root cause analysis result, obtaining the cause of the abnormality.

[0006] In some embodiments, obtaining the cause of the exception based on the target call stack and the root cause analysis results includes: obtaining a pre-trained large language model; inputting the target call stack and the root cause analysis results into the large language model to obtain the cause of the exception output by the large language model.

[0007] In some embodiments, the above-mentioned log data includes streaming real-time data, and the above-mentioned anomaly detection on the acquired log data to obtain anomaly keywords includes: in response to the data volume of the streaming real-time data being greater than a data threshold, performing data preprocessing on the streaming real-time data to obtain preprocessed log data; based on the preprocessed log data, obtaining filtering keywords; based on the filtering keywords, obtaining anomaly keywords.

[0008] In some embodiments, the above-mentioned filtering out the target call stack from the log data based on the abnormal keyword includes: determining at least one original call stack corresponding to the abnormal keyword in the log data; determining the frequency of occurrence of each original call stack, and taking the original call stack with the highest frequency of occurrence in the log data as the target call stack.

[0009] In some embodiments, determining the root cause analysis result based on the target call stack includes: sorting the occurrence frequencies of the IPs corresponding to the target call stack, and taking the IP with the highest occurrence frequency as the root cause analysis result.

[0010] In some embodiments, determining the root cause analysis result based on the target call stack includes: responding to a key indicator corresponding to the target call stack in a database; obtaining the number of occurrences of the target call stack; and determining the root cause analysis result based on the number of occurrences, the target call stack and the semantic similarity of the key indicator.

[0011] In some embodiments, the above-mentioned log data includes alarm data in a preset time period after the key indicator alarm, and the acquired log data is subjected to anomaly detection to obtain abnormal keywords, including: performing data cleaning and data filtering on the alarm data to obtain processed data; obtaining filtered keywords based on the processed data; and obtaining abnormal keywords based on the filtered keywords.

[0012] In some embodiments, the above-mentioned filtering out the target call stack from the log data based on the abnormal keyword includes: determining at least one original call stack corresponding to the abnormal keyword in the log data; determining the frequency of occurrence of each original call stack; and determining the target call stack based on the frequency of occurrence, the semantic similarity of the root cause method and the key indicator in each original call stack.

[0013] In some embodiments, determining the root cause analysis result based on the target call stack includes: taking the root cause method with the greatest semantic similarity to the key indicator in the target call stack as the root cause analysis result.

[0014] In a second aspect, an embodiment of the present disclosure provides an abnormal data processing device, which includes: a word obtaining unit, configured to perform abnormality detection on the acquired log data to obtain abnormal keywords; a screening unit, configured to filter out a target call stack from the log data based on the abnormal keywords; a determination unit, configured to determine a root cause analysis result based on the target call stack; and a cause obtaining unit, configured to obtain the cause of the abnormality based on the target call stack and the root cause analysis result.

[0015] In some embodiments, the cause obtaining unit is configured to: obtain a pre-trained large language model; input the target call stack and the root cause analysis result into the large language model to obtain the abnormal cause output by the large language model.

[0016] In some embodiments, the above-mentioned log data includes streaming real-time data, and the above-mentioned word obtaining unit is configured to: in response to the data volume of the streaming real-time data being greater than a data threshold, perform data preprocessing on the streaming real-time data to obtain preprocessed log data; obtain filtering keywords based on the preprocessed log data; and obtain abnormal keywords based on the filtering keywords.

[0017] In some embodiments, the screening unit is configured to: determine at least one original call stack corresponding to the abnormal keyword in the log data; determine the frequency of occurrence of each original call stack, and use the original call stack with the highest frequency in the log data as the target call stack.

[0018] In some embodiments, the determination unit is configured to: sort the occurrence frequencies of the IPs corresponding to the target call stack, and use the IP with the highest occurrence frequency as the root cause analysis result.

[0019] In some embodiments, the determination unit is configured to: respond to a database having a key indicator corresponding to a target call stack; obtain the number of occurrences of the target call stack; and determine a root cause analysis result based on the number of occurrences, the target call stack, and the semantic similarity of the key indicator.

[0020] In some embodiments, the above-mentioned log data includes alarm data in a preset time period after the key indicator alarm, and the above-mentioned word obtaining unit is configured to: perform data cleaning and data filtering on the alarm data to obtain processed data; obtain filtering keywords based on the processed data; and obtain abnormal keywords based on the filtering keywords.

[0021] In some embodiments, the above-mentioned screening unit is configured to: determine at least one original call stack corresponding to the abnormal keyword in the log data; determine the frequency of occurrence of each original call stack; and determine the target call stack based on the frequency of occurrence, the semantic similarity of the root cause method and the key indicators in each original call stack.

[0022] In some embodiments, the determination unit is configured to: take the root cause method in the target call stack that has the greatest semantic similarity with the key indicator as the root cause analysis result.

[0023] In a third aspect, an embodiment of the present disclosure provides an electronic device, comprising: one or more processors; a storage device on which one or more programs are stored; when the one or more programs are executed by one or more processors, the one or more processors implement the method described in any embodiment of the first aspect.

[0024] In a fourth aspect, an embodiment of the present disclosure provides a computer-readable medium having a computer program stored thereon, which, when executed by a processor, implements the method described in any embodiment of the first aspect.

[0025] The abnormal data processing method and device provided by the embodiments of the present disclosure perform abnormal detection on the acquired log data to obtain abnormal keywords; based on the abnormal keywords, filter out the target call stack from the log data; based on the target call stack, determine the root cause analysis result; based on the target call stack and the root cause analysis result, obtain the cause of the abnormality. Thus, the target call stack is obtained based on the abnormal keywords, and the cause of the abnormality is obtained based on the target call stack and the root cause analysis result, which effectively locates the root cause of the abnormality, can assist in quickly and accurately locating the fault, and improves the efficiency of abnormal data analysis. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Other features, objects and advantages of the present disclosure will become more apparent from the detailed description of non-limiting embodiments made with reference to the following drawings:

[0027] Figure 1 is an exemplary system architecture diagram in which an embodiment of the present disclosure may be applied;

[0028] Figure 2 is a flow chart of an embodiment of an abnormal data processing method according to the present disclosure;

[0029] Figure 3 It is a structural schematic diagram obtained by the abnormal cause in the abnormal data processing method disclosed in the present invention;

[0030] Figure 4 is a structural schematic diagram of an embodiment of an abnormal data processing device according to the present disclosure;

[0031] Figure 5 It is a schematic diagram of the structure of an electronic device suitable for implementing the embodiments of the present disclosure. DETAILED DESCRIPTION

[0032] The present disclosure is further described in detail below in conjunction with the accompanying drawings and embodiments. It is understood that the specific embodiments described herein are only used to explain the relevant invention, rather than to limit the invention. It is also necessary to explain that, for ease of description, only the parts related to the relevant invention are shown in the accompanying drawings.

[0033] It should be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present disclosure may be combined with each other. The present disclosure will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0034] Figure 1 An exemplary system architecture 100 is shown to which the abnormal data processing method of the present disclosure can be applied.

[0035] like Figure 1As shown, the system architecture 100 may include terminal devices 101, 102, 103, a network 104 and a server 105. The network 104 is used to provide a medium for communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, and may generally include wireless communication links and the like.

[0036] The terminal devices 101, 102, 103 interact with the server 105 via the network 104 to receive or send messages, etc. Various communication client applications, such as instant messaging tools, email clients, etc., may be installed on the terminal devices 101, 102, 103.

[0037] The terminal devices 101, 102, 103 can be hardware or software; when the terminal devices 101, 102, 103 are hardware, they can be user devices with communication and control functions, and the above user settings can communicate with the server 105. When the terminal devices 101, 102, 103 are software, they can be installed in the above user devices; the terminal devices 101, 102, 103 can be implemented as multiple software or software modules (for example, software or software modules used to provide distributed services), or they can be implemented as a single software or software module. No specific limitation is made here.

[0038] The server 105 may be a server that provides various services, such as a background server that analyzes log data on the terminal devices 101, 102, and 103. The background server may receive the log data sent by the terminal devices 101, 102, and 103, and obtain an abnormal keyword based on the log data, filter out a target call stack through the abnormal keyword, determine a root cause analysis result based on the target call stack, and obtain the cause of the abnormality based on the root cause analysis result and the target call stack.

[0039] It should be noted that the server can be hardware or software. When the server is hardware, it can be implemented as a distributed server cluster consisting of multiple servers, or it can be implemented as a single server. When the server is software, it can be implemented as multiple software or software modules (for example, software or software modules used to provide distributed services), or it can be implemented as a single software or software module. No specific limitation is made here.

[0040] It should be noted that the abnormal data processing method provided in the embodiments of the present disclosure is generally executed by the server 105 , and the abnormal data processing device provided in the embodiments of the present disclosure is generally set in the server 105 .

[0041] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is only for illustration. Any number of terminal devices, networks and servers may be provided according to implementation requirements.

[0042] In traditional technologies, algorithms that combine operation and maintenance data with big data and machine learning technologies to locate the root cause of anomalies have gradually become mainstream. Most algorithms are based on clustering, association analysis, or Bayesian algorithms to locate the root cause of anomalies. For example, root cause location solutions based on call chains, root cause analysis based on abnormal range search, and algorithms based on time series correlation analysis. Logs contain a lot of valid information, but due to the large volume of logs, valid information is often submerged in massive logs. At present, most technologies for fault analysis on logs are based on log tracing, but this method requires that log data must be standardized, which is not very realistic in actual application scenarios.

[0043] In view of the above defects, the present disclosure provides a method for processing abnormal data, which can effectively locate the root cause of the abnormality and assist in quickly and accurately locating the fault. Figure 2 , shows a process 200 of an embodiment of an abnormal data processing method according to the present disclosure, the abnormal data processing method comprises the following steps:

[0044] Step 201: perform anomaly detection on the acquired log data to obtain anomaly keywords.

[0045] In this embodiment, the log data may be information recorded by the system when some system has problems or program errors. The content of the log data may be different due to different aspects. For example, for code errors, the log data is in the form of a call stack, and each call stack includes information such as the method called by the code and the number of lines corresponding to the code. Optionally, the log data may also be system detection data collected in real time. By analyzing the amount of collected log data, it can be determined whether the log is reflecting system abnormalities. When reflecting system abnormalities, the target call stack can be obtained through the log data.

[0046] In this embodiment, the execution subject on which the abnormal data processing method runs can obtain log data from multiple places, for example, from Figure 1 Log data is obtained from the terminal devices 101, 102, and 103 shown.

[0047] In this embodiment, the log data may be real-time data or historical data. After obtaining the log data, the above step 201 includes: performing keyword processing on the log data to obtain initial keywords in the log data; filtering normal keywords in the initial keywords to obtain abnormal keywords.

[0048] Step 202: Filter out the target call stack from the log data based on the abnormal keyword.

[0049] In this embodiment, the call stack is a method stack. The method stack is an independent memory space allocated by the system for each method call of an object. The method stack is not unique to the object. If a method of the same object is called multiple times, the method stack may be different each time.

[0050] In this embodiment, each piece of data in the log data may include multiple words and a call stack. After obtaining the exception keyword, the log data is searched using the exception keyword to determine the call stack corresponding to the exception keyword, and the call stack corresponding to the exception keyword is the target call stack.

[0051] Step 203: Determine the root cause analysis result based on the target call stack.

[0052] In this embodiment, the root cause analysis result is a root cause failure result obtained through the target call stack, and the root cause failure result is a failure instance that causes the abnormality of the log data. Specifically, the root cause analysis result includes: the abnormal IP and the root cause method included in the target call stack. The root cause method is a function method indicated by the stack in the call stack, for example, a line of the call stack is a root cause method.

[0053] The above step 203 includes: performing similarity analysis on the root cause methods in the target call stack, collecting all the root cause methods whose similarity is greater than a similarity threshold (eg, 90%), and using the collected root cause methods as the root cause analysis results.

[0054] like Figure 3 As shown, after obtaining the exception keyword, the target call stack is determined, and the root cause analysis result is determined through the target call stack.

[0055] Step 204, based on the target call stack and the root cause analysis result, obtain the cause of the exception.

[0056] In this embodiment, the abnormal cause is the cause of the log abnormality, and the abnormal cause can be determined in a variety of ways, such as using a variety of rules to perform rule analysis on the target call stack and the root cause analysis results, or bringing the target call stack and the root cause analysis results into a pre-established correspondence table for table analysis. Among them, the multiple rules are the tendency results of the abnormal cause obtained after the tendency analysis of the call stack and the root cause analysis results in a large amount of abnormal log data. For example, one of the multiple rules is that a root cause method in the call stack corresponds to an abnormal cause. Among them, the correspondence table is used to characterize the correspondence between the call stack, the root cause method in the call stack, and the abnormal cause. The target call stack and the root cause method in the root cause analysis result are simultaneously matched with the call stack and the root cause method in the correspondence table. The abnormal cause of the call stack and the root cause method in the correspondence table that are most similar to the target call stack and the root cause method in the root cause analysis result is selected to obtain the common abnormal cause.

[0057] The abnormal data processing method provided by the embodiment of the present disclosure first performs abnormal detection on the acquired log data to obtain abnormal keywords; secondly, based on the abnormal keywords, the target call stack is screened out from the log data; thirdly, based on the target call stack, the root cause analysis result is determined; finally, based on the target call stack and the root cause analysis result, the abnormal cause is obtained. Thus, the target call stack is obtained based on the abnormal keywords, and the abnormal cause is obtained based on the target call stack and the root cause analysis result, which effectively locates the abnormal root cause, can assist in quickly and accurately locating the fault, and improves the efficiency of abnormal data analysis.

[0058] In some optional implementations disclosed above, obtaining the cause of the exception based on the target call stack and the root cause analysis results includes: obtaining a pre-trained large language model; inputting the target call stack and the root cause analysis results into the large language model to obtain the cause of the exception output by the large language model.

[0059] In this optional implementation, the large language model can be ChatGPT (Chat Generative Pre-trained Transformer), which is a large language model based on deep learning technology. As a new chatbot, ChatGPT can not only chat, but also write papers, poems, scripts, and codes, showing a high degree of intelligence and possessing primary thinking logic. Similarly, the technical advantages and potential of large language models in the field of intelligent operation and maintenance are unlimited.

[0060] The large language model can be used as the core technology of log analysis to understand and generate natural language for operation and maintenance-related log data, and provide functions such as anomaly detection, root cause analysis, and repair suggestions. The large language model is used to further analyze and provide suggestions for the initially located root cause analysis results and target call stack information. It can accurately return the root cause method of the exception in seconds and give the cause of the exception to assist R&D personnel in fault recovery.

[0061] like Figure 3 As shown, after obtaining the root cause analysis result, the root cause analysis result is input into ChatGPT to obtain the abnormal cause output by ChatGPT.

[0062] The method for obtaining the cause of the exception provided in this embodiment analyzes the target call stack and the root cause analysis results through a pre-trained large language model to obtain the cause of the exception, thereby improving the reliability of obtaining the cause of the exception.

[0063] Optionally, after the target call stack and the root cause analysis results are input into the large language model, the exception cause and recovery suggestions output by the large language model can also be obtained. The abnormal data processing method provided in this embodiment triggers log analysis by detecting log data. When an application exception occurs, the original information of the abnormal log corresponding to the abnormal keyword is detected to obtain the root cause analysis results, and then the abnormal root cause results are further analyzed through ChatGPT, and detailed analysis and recovery suggestions are given, thereby improving the reliability of abnormal data processing.

[0064] In some optional implementations disclosed, the above-mentioned log data includes streaming real-time data, and anomaly detection is performed on the acquired log data to obtain anomaly keywords, including: in response to the data volume of the streaming real-time data being greater than a data threshold, data preprocessing is performed on the streaming real-time data to obtain preprocessed log data; based on the preprocessed log data, filtering keywords are obtained; based on the filtering keywords, anomaly keywords are obtained.

[0065] In this optional implementation, the streaming real-time data is the data in the log data that needs to be monitored in real time. When the amount of the streaming real-time data increases abnormally, it can reflect the abnormality of the streaming real-time data. The data threshold is a limit value used to reflect the data abnormality, and the data threshold can be obtained through multiple experiments. For example, the data threshold is 5MB.

[0066] In this optional implementation, the data preprocessing includes: filling missing values, data format conversion, and data normalization for streaming real-time data. Missing value filling refers to filling the blank areas in the log data with fixed values ​​(such as 0), data format conversion refers to converting all data in different formats in the log data into a unified format, and data normalization refers to converting all data in the log data into the same numerical range.

[0067] In this optional implementation, the above-mentioned obtaining the filtering keywords based on the preprocessed log data includes: filtering the preprocessed log data using a fixed threshold method to filter out obviously normal keywords to obtain the filtering keywords. Optionally, the above-mentioned obtaining the filtering keywords based on the preprocessed log data includes: obtaining the filtering keywords from the preprocessed log data based on the features of the pre-calibrated abnormal keywords.

[0068] In this optional implementation, the above-mentioned obtaining abnormal keywords based on the filtered keywords includes: performing abnormal detection on the filtered keywords by using mean shift and isolated peak filtering methods to detect abnormal keywords with abnormal surge.

[0069] The abnormal data processing method provided in this embodiment performs data preprocessing on the streaming real-time data when the log data includes streaming real-time data and when the data volume of the streaming real-time data is greater than a data threshold, and filters the preprocessed log data to obtain abnormal keywords, thereby providing an optional implementation method for obtaining the abnormal keywords and improving the reliability of obtaining the abnormal keywords.

[0070] Optionally, when the log data includes streaming real-time data, the above-mentioned anomaly detection is performed on the acquired log data to obtain abnormal keywords, including: in response to the data volume of the streaming real-time data being greater than a data threshold, using the Drain (DepthTree Online Log Parsing Image.png fixed depth parsing tree) algorithm to perform offline template extraction to perform keyword template extraction, and using massive data to offline count the upper and lower data of the keywords; performing data preprocessing on the upper and lower data to obtain preprocessed log data, using a fixed threshold method to filter the preprocessed log data, filtering out obviously normal keywords, and obtaining filtered keywords; performing anomaly detection on the filtered keyword data through mean shift and isolated peak filtering methods to detect abnormal keywords with abnormal surges.

[0071] In some optional implementations of the present disclosure, the above-mentioned filtering out the target call stack from the log data based on the exception keyword includes: determining at least one original call stack corresponding to the exception keyword in the log data; determining the frequency of occurrence of each original call stack, and taking the original call stack with the highest frequency of occurrence in the log data as the target call stack.

[0072] In this optional implementation, for streaming real-time data, based on the exception keyword obtained through the streaming real-time data, the target call stack can be determined through the original call stack corresponding to the exception keyword and the frequency of occurrence of the original call stack.

[0073] In this optional implementation, each exception keyword corresponds to one or more original call stacks, and based on the exception keyword, at least one original call stack corresponding to the exception keyword can be obtained.

[0074] In this optional implementation, the above-mentioned determination of the occurrence frequency of each original call stack and taking the original call stack with the highest occurrence frequency in the log data as the target call stack includes: determining the occurrence frequency of each original call stack, sorting the occurrence frequency of all original call stacks in ascending or descending order, and taking the original call stack with the highest occurrence frequency in the ascending or descending results as the target call stack.

[0075] The method for obtaining the target call stack provided by this optional implementation method determines at least one original call stack corresponding to the exception keyword in the log data; determines the frequency of occurrence of each original call stack, and uses the original call stack with the highest frequency of occurrence in the log data as the target call stack, providing a reliable implementation method for obtaining the target call stack.

[0076] In some optional implementations of the present disclosure, determining the root cause analysis result based on the target call stack includes: sorting the occurrence frequencies of IPs corresponding to the target call stack, and taking the IP with the highest occurrence frequency as the root cause analysis result.

[0077] In this optional implementation, the target call stack includes at least one call stack, and there are multiple IPs corresponding to the target call stack. The IPs corresponding to the target call stack are sorted from small to large in frequency of occurrence to obtain the IP with the highest frequency of occurrence, and the IP with the highest frequency of occurrence is used as the root cause analysis result.

[0078] The abnormal data processing method provided in this implementation sorts the IP occurrence frequencies corresponding to the target call stack, and uses the IP with the highest occurrence frequency as the root cause analysis result, thereby providing a reliable implementation method for obtaining the root cause analysis result.

[0079] In some optional implementations of the present disclosure, determining the root cause analysis result based on the target call stack includes: responding to a key indicator corresponding to the target call stack in a database; obtaining the number of occurrences of the target call stack; and determining the root cause analysis result based on the number of occurrences, the target call stack and the semantic similarity of the key indicator.

[0080] In this optional implementation, key indicators refer to indicators that consider the abnormal effects of log data. When it is streaming real-time data, the streaming real-time data is stored in a database, and the key indicators of the streaming real-time data are analyzed accordingly, and the streaming real-time data and the key indicators of the streaming real-time data are stored accordingly.

[0081] In this optional implementation, the root cause analysis result is determined based on the semantic similarity of the number of occurrences, the target call stack and the key indicator, including: determining the number of occurrences of each root cause method in the target call stack based on the number of occurrences; calculating the similarity of each root cause method with the key indicator respectively to obtain the similarity value of each root cause method; multiplying the number of occurrences of each root cause method with its similarity value to obtain a total similarity value; sorting the total similarity values ​​of all root cause methods in the target call stack to obtain the root cause method with the largest total similarity value, and using the root cause method as the root cause analysis result.

[0082] Optionally, the above-mentioned determination of the root cause analysis result based on the number of occurrences, the target call stack and the semantic similarity of the key indicator includes: determining the number of occurrences of each root cause method in the target call stack based on the number of occurrences; calculating the similarity of each root cause method with the key indicator respectively to obtain the similarity value of each root cause method; multiplying the number of occurrences of each root cause method with its similarity value to obtain a total similarity value; sorting the total similarity values ​​of all root cause methods in the target call stack in ascending or descending order to obtain the root cause method in the last N (N>1) or first N places of the total similarity value, and taking the root cause method as the root cause analysis result.

[0083] The abnormal data processing method provided in this embodiment responds to the key indicators corresponding to the target call stack in the database; obtains the number of occurrences of the target call stack; determines the root cause analysis result based on the number of occurrences, the target call stack and the semantic similarity of the key indicators, and provides another reliable implementation method for obtaining the root cause analysis result.

[0084] Optionally, the above-mentioned determination of the root cause analysis result based on the target call stack includes: in response to the database not having a key indicator corresponding to the target call stack; obtaining the number of occurrences of the target call stack; based on the number of occurrences, determining the IP of the call stack that appears most times in the target call stack, and using the IP of the call stack that appears most times as the root cause analysis result.

[0085] In some optional implementations of the present embodiment, the above-mentioned log data includes alarm data in a preset time period after the key indicator alarm, and the above-mentioned abnormality detection on the acquired log data to obtain abnormal keywords includes: data cleaning and data filtering on the alarm data to obtain processed data; based on the processed data, obtaining filtered keywords; based on the filtered keywords, obtaining abnormal keywords.

[0086] In this optional implementation, the key indicator is an indicator that indicates the abnormal amount of log data. For example, the key indicators include: number of errors, tp99; the number of errors refers to the number of errors in a unit of information, and tp99 means that after configuring the alarm threshold corresponding to the monitoring indicator, it is necessary to ensure that at least 99% of the time consumed by all calls of the method within a certain time period is less than the word threshold, otherwise the system will alarm.

[0087] In this optional implementation, the abnormal keyword includes the abnormal keyword. When the log data only contains alarm data, the abnormal keyword obtained by performing abnormal detection on the log data is the abnormal keyword; when the log data does not only contain alarm data, the abnormal keyword obtained by performing abnormal detection on the log data includes the abnormal keyword.

[0088] In this optional implementation, when the key indicator exceeds the preset threshold, real-time collection is performed after the key indicator alarm is triggered, which triggers the process of obtaining the alarm data in the preset time period. When the abnormal keyword is obtained, data collection is first performed to pull the log data of the abnormal application or system at the preset time from the data storage system; then data cleaning and data filtering are performed to obtain processed data.

[0089] The above-mentioned method of obtaining filtering keywords based on processed data includes: counting the number of keywords from log data; then performing preliminary data processing, using a fixed threshold method (the fixed threshold method means that the total number of data items does not exceed a certain threshold, which means that there is no problem by default) to filter out obviously normal keywords and obtain filtering keywords.

[0090] The above-mentioned obtaining of abnormal keywords based on the filtered keywords includes: performing anomaly detection on the filtered keyword data by using a mean shift and a peak filtering method to detect abnormal keywords with abnormal surge.

[0091] The abnormal data processing method provided in this embodiment performs data cleaning and data filtering on the alarm data to obtain processed data; based on the processed data, a filtering keyword is obtained; based on the filtering keyword, an abnormal keyword is obtained, which provides another optional implementation method for obtaining the abnormal keyword and improves the reliability of obtaining the abnormal keyword.

[0092] Optionally, the above-mentioned log data includes: streaming real-time data and alarm data in a preset time period after the key indicator alarm, and the above-mentioned abnormality detection on the acquired log data to obtain abnormal keywords includes: in response to the data volume of the streaming real-time data being greater than the data threshold, preprocessing the streaming real-time data to obtain preprocessed log data; obtaining a filtering keyword based on the preprocessed log data; obtaining a first abnormal keyword based on the filtering keyword; cleaning and filtering the alarm data to obtain processed data; obtaining a filtering keyword based on the processed data; obtaining a second abnormal keyword based on the filtering keyword; and using the set of the first abnormal keyword and the second abnormal keyword as the abnormal keyword. Specifically, as Figure 3 As shown, the log data includes: alarm data and streaming real-time data, and abnormal keywords are obtained through the alarm data and streaming real-time data.

[0093] In some optional implementations of the present disclosure, the above-mentioned filtering out the target call stack from the log data based on the abnormal keyword includes: determining at least one original call stack corresponding to the abnormal keyword in the log data; determining the frequency of occurrence of each original call stack; and determining the target call stack based on the frequency of occurrence, the semantic similarity of the root cause method and the key indicator in each original call stack.

[0094] In this optional implementation, the above-mentioned determination of the target call stack based on the frequency of occurrence, the semantic similarity of the root cause method in each original call stack and the key indicator includes: calculating the similarity between each original call stack and the key indicator to obtain at least one semantic similarity of each original call stack, selecting the maximum semantic similarity of each original call stack, multiplying the frequency of occurrence of each original call stack by the respective maximum semantic similarity, and obtaining the similarity probability value of each original call stack, wherein the similarity probability value is obtained by combining the semantic similarity probability calculated by the keyword and the frequency of occurrence to obtain the final similarity probability value of the call stack method under the keyword. The similarity probability values ​​of all original call stacks are sorted in ascending or descending order to obtain the original call stack with the largest similarity probability value, and the original call stack with the largest similarity probability value is used as the target call stack.

[0095] The method for obtaining the target call stack provided by this optional implementation method determines at least one original call stack corresponding to the abnormal keyword in the log data; determines the frequency of occurrence of each original call stack; and determines the target call stack based on the frequency of occurrence, the semantic similarity of the root cause method and the key indicator in each original call stack, thereby providing a reliable implementation method for obtaining the target call stack under the alarm data.

[0096] In some optional implementations of the present disclosure, determining the root cause analysis result based on the target call stack includes: taking the root cause method with the greatest semantic similarity to the key indicator in the target call stack as the root cause analysis result.

[0097] In this optional implementation, a root cause method with the largest similarity probability value in the target call stack is selected as the root cause method under the abnormal keyword, and the root cause method is used as the root cause analysis result. Optionally, the root cause method with the largest semantic similarity to the key indicator in the target call stack and the similarity probability value corresponding to the root cause method can also be used as the root cause analysis result. The similarity probability values ​​less than the similarity threshold are pruned to finally obtain the similarity probability value corresponding to the root cause method.

[0098] The method for obtaining the root cause analysis result provided by this optional implementation takes the root cause method with the greatest semantic similarity to the key indicator in the target call stack as the root cause analysis result, providing a reliable implementation method for obtaining the root cause analysis result.

[0099] Optionally, the similarity probability value corresponding to the root cause method can also be obtained by weighting the numerical similarity probability and the semantic similarity probability, where similarity is the numerical similarity probability, as shown in formula (1) to formula (3):

[0100] similarity=1 / e -dist (1)

[0101]

[0102] Where x is the key indicator, y is the time series data of the number of errors of abnormal keywords, α is the smoothing coefficient, and DTW(x,y) is the distance between the two time series obtained by the DTW (Dynamic Time Warping) algorithm.

[0103] SimCSE (Simple Contrastive Learning of SentenceEmbeddings) semantic similarity probability, as follows:

[0104] errorkey n =ω1*Similarity+ω2*max(SimCSE(X,Y)) (3)

[0105] In formula (3), ω1 and ω2 are the weights of numerical similarity probability and semantic similarity probability, which are set to 0.2 and 0.8 respectively. SimCSE(X,Y) returns the semantic similarity probability of each call stack corresponding to the key indicator and the abnormal keyword.

[0106] Further references Figure 4 As an implementation of the abnormal data processing method shown in the above figures, the present disclosure provides an embodiment of an abnormal data processing device. Figure 2 Corresponding to the method embodiment shown, the device can be specifically applied to various electronic devices.

[0107] like Figure 4 As shown, an embodiment of the present disclosure provides an abnormal data processing device 400, and the device 400 includes: a word obtaining unit 401, a screening unit 402, a determination unit 403, and a cause obtaining unit 404. Among them, the above-mentioned word obtaining unit 401 can be configured to perform abnormality detection on the acquired log data to obtain abnormal keywords. The above-mentioned screening unit 402 can be configured to filter out the target call stack from the log data based on the abnormal keywords. The above-mentioned determination unit 403 can be configured to determine the root cause analysis result based on the target call stack. The above-mentioned cause obtaining unit 404 can be configured to obtain the cause of the exception based on the target call stack and the root cause analysis result.

[0108] In this embodiment, the specific processing of the word obtaining unit 401, the screening unit 402, the determination unit 403, and the cause obtaining unit 404 and the technical effects thereof can be referred to in the respective Figure 2 Corresponding to step 201, step 202, step 203, and step 204 in the embodiment.

[0109] In some embodiments, the cause obtaining unit 404 is configured to: obtain a pre-trained large language model; input the target call stack and the root cause analysis result into the large language model to obtain the abnormal cause output by the large language model.

[0110] In some embodiments, the above-mentioned log data includes streaming real-time data, and the above-mentioned word obtaining unit 401 can be configured to: in response to the data volume of the streaming real-time data being greater than a data threshold, perform data preprocessing on the streaming real-time data to obtain preprocessed log data; obtain filtering keywords based on the preprocessed log data; and obtain abnormal keywords based on the filtering keywords.

[0111] In some embodiments, the screening unit 402 is configured to: determine at least one original call stack corresponding to the abnormal keyword in the log data; determine the frequency of occurrence of each original call stack, and use the original call stack with the highest frequency in the log data as the target call stack.

[0112] In some embodiments, the determination unit 403 may be configured to sort the occurrence frequencies of the IPs corresponding to the target call stack, and use the IP with the highest occurrence frequency as the root cause analysis result.

[0113] In some embodiments, the determination unit 403 may be configured to: respond to a database having a key indicator corresponding to a target call stack; obtain the number of occurrences of the target call stack; and determine the root cause analysis result based on the number of occurrences, the target call stack, and the semantic similarity of the key indicator.

[0114] In some embodiments, the above-mentioned log data includes alarm data in a preset time period after the key indicator alarm, and the above-mentioned word obtaining unit 401 can be configured to: perform data cleaning and data filtering on the alarm data to obtain processed data; obtain filtering keywords based on the processed data; obtain abnormal keywords based on the filtering keywords.

[0115] In some embodiments, the above-mentioned screening unit 402 is configured to: determine at least one original call stack corresponding to the abnormal keyword in the log data; determine the frequency of occurrence of each original call stack; and determine the target call stack based on the frequency of occurrence, the semantic similarity of the root cause method and the key indicator in each original call stack.

[0116] In some embodiments, the determination unit 403 may be configured to: take the root cause method in the target call stack that has the greatest semantic similarity with the key indicator as the root cause analysis result.

[0117] The abnormal data processing device provided by the embodiment of the present disclosure, first, the word obtaining unit 401 performs abnormal detection on the acquired log data to obtain abnormal keywords; secondly, the screening unit 402 screens out the target call stack from the log data based on the abnormal keywords; thirdly, the determination unit 403 determines the root cause analysis result based on the target call stack; finally, the cause obtaining unit 404 obtains the abnormal cause based on the target call stack and the root cause analysis result. Thus, the target call stack is obtained based on the abnormal keywords, and the abnormal cause is obtained based on the target call stack and the root cause analysis result, which effectively locates the abnormal root cause, can assist in quickly and accurately locating the fault, and improves the efficiency of abnormal data analysis.

[0118] Reference below Figure 5 , which shows a structural schematic diagram of an electronic device 500 suitable for implementing an embodiment of the present disclosure.

[0119] like Figure 5 As shown, the electronic device 500 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 501, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 502 or a program loaded from a storage device 508 into a random access memory (RAM) 503. In the RAM 503, various programs and data required for the operation of the electronic device 500 are also stored. The processing device 501, the ROM 502, and the RAM 503 are connected to each other via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.

[0120] Typically, the following devices may be connected to the I / O interface 505: an input device 506 including, for example, a touch screen, a touch pad, a keyboard, a mouse, etc.; an output device 507 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 508 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 509. The communication device 509 may allow the electronic device 500 to communicate with other devices wirelessly or by wire to exchange data. Figure 5 The electronic device 500 is shown with various devices, but it should be understood that it is not required to implement or possess all the devices shown. More or fewer devices may be implemented or possessed instead. Figure 5 Each block shown in the figure may represent one device, or may represent multiple devices as required.

[0121] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through a communication device 509, or installed from a storage device 508, or installed from a ROM 502. When the computer program is executed by the processing device 501, the above-mentioned functions defined in the method of the embodiment of the present disclosure are executed.

[0122] It should be noted that the computer-readable medium of the embodiment of the present disclosure may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the embodiment of the present disclosure, the computer-readable storage medium may be any tangible medium containing or storing a program, which may be used by or in combination with an instruction execution system, device or device. In the embodiment of the present disclosure, the computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which a computer-readable program code is carried. This propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer readable signal medium may also be any computer readable medium other than a computer readable storage medium, which may send, propagate or transmit a program for use by or in conjunction with an instruction execution system, apparatus or device. The program code contained on the computer readable medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination of the above.

[0123] The computer-readable medium may be included in the server; or it may exist independently without being installed in the server. The computer-readable medium carries one or more programs. When the one or more programs are executed by the server, the server: performs anomaly detection on the acquired log data to obtain anomaly keywords; based on the anomaly keywords, filters out the target call stack from the log data; based on the target call stack, determines the root cause analysis result; based on the target call stack and the root cause analysis result, obtains the cause of the anomaly.

[0124] Computer program code for performing the operations of embodiments of the present disclosure may be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages, such as Java, Smalltalk, C++, and conventional procedural programming languages, such as "C" or similar programming languages. The program code may be executed entirely on a user's computer, partially on a user's computer, as a separate software package, partially on a user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0125] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present disclosure. In this regard, each box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0126] The units involved in the embodiments described in the present disclosure may be implemented by software or by hardware. The described units may also be arranged in a processor, for example, may be described as: a processor, comprising a word obtaining unit, a screening unit, a determination unit, and a cause obtaining unit. Among them, the names of these units do not constitute a limitation on the unit itself under certain circumstances. For example, the word obtaining unit may also be described as a unit "configured to perform anomaly detection on the acquired log data and obtain anomaly keywords".

[0127] The above description is only a preferred embodiment of the present disclosure and an explanation of the technical principles used. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by a specific combination of the above-mentioned technical features, but should also cover other technical solutions formed by any combination of the above-mentioned technical features or their equivalent features without departing from the above-mentioned inventive concept. For example, the above-mentioned features are replaced with the technical features with similar functions disclosed in the embodiments of the present disclosure (but not limited to) to form a technical solution.

Claims

1. A method for processing abnormal data, the method comprising: Perform anomaly detection on the acquired log data to obtain abnormal keywords; Based on the abnormal keyword, filter out a target call stack from the log data; Determine a root cause analysis result based on the target call stack; Based on the target call stack and the root cause analysis result, the cause of the exception is obtained.

2. The method according to claim 1, wherein: The abnormal cause obtained based on the target call stack and the root cause analysis result includes: Get a pre-trained large language model; The target call stack and the root cause analysis result are input into the large language model to obtain the abnormal cause output by the large language model.

3. The method according to claim 1 or 2, wherein: The log data includes streaming real-time data, and the abnormality detection is performed on the acquired log data to obtain abnormal keywords including: In response to the data volume of the streaming real-time data being greater than a data threshold, performing data preprocessing on the streaming real-time data to obtain preprocessed log data; Based on the preprocessed log data, obtaining filtering keywords; Based on the filtering keywords, abnormal keywords are obtained.

4. The method according to claim 3, wherein: The step of filtering out a target call stack from the log data based on the abnormal keyword includes: In the log data, determining at least one original call stack corresponding to the abnormal keyword; The occurrence frequency of each original call stack is determined, and the original call stack with the highest occurrence frequency in the log data is used as the target call stack.

5. The method according to claim 3, wherein: Determining the root cause analysis result based on the target call stack includes: The IP occurrence frequencies corresponding to the target call stack are sorted, and the IP with the highest occurrence frequency is used as the root cause analysis result.

6. The method according to claim 3, wherein: Determining the root cause analysis result based on the target call stack includes: In response to a key indicator corresponding to the target call stack being present in a database; Obtaining the number of occurrences of the target call stack; A root cause analysis result is determined based on the number of occurrences, the target call stack, and the semantic similarity of the key indicator.

7. The method according to claim 1 or 2, wherein the log data includes alarm data in a preset time period after a key indicator alarm, and the abnormality detection is performed on the acquired log data to obtain abnormal keywords including: Performing data cleaning and data filtering on the alarm data to obtain processed data; Based on the processed data, a filtering keyword is obtained; Based on the filtering keywords, abnormal keywords are obtained.

8. The method according to claim 7, wherein: The step of filtering out a target call stack from the log data based on the abnormal keyword includes: In the log data, determining at least one original call stack corresponding to the abnormal keyword; Determine the frequency of occurrence of each original call stack; Based on the occurrence frequency, the semantic similarity between the root cause method in each original call stack and the key indicator, a target call stack is determined.

9. The method according to claim 8, wherein: Determining the root cause analysis result based on the target call stack includes: The root cause method in the target call stack with the greatest semantic similarity to the key indicator is taken as the root cause analysis result.

10. An abnormal data processing device, the device comprising: The word obtaining unit is configured to perform anomaly detection on the acquired log data and obtain anomaly keywords; A screening unit, configured to screen out a target call stack from the log data based on the abnormal keyword; a determination unit configured to determine a root cause analysis result based on the target call stack; The cause obtaining unit is configured to obtain the exception cause based on the target call stack and the root cause analysis result.

11. An electronic device, comprising: one or more processors; a storage device having one or more programs stored thereon; When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 9.

12. A computer readable medium having a computer program stored thereon, wherein: When the program is executed by a processor, the method according to any one of claims 1 to 9 is implemented.