Log data processing method and device, electronic equipment and readable storage medium
By obtaining the key field information of the log, determining the processing priority, associating the log and extracting exception information, the problem of log exception detection in international business systems is solved, and the accuracy of exception determination is achieved.
Patent Information
- Application Number
- CN202510219769.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-02-26
AI Technical Summary
In international business systems, log data formats are diverse and have strong cross-system liquidity, which leads to difficulty in detecting and capturing log exceptions. Traditional manual processing has problems such as time-consuming, omissions and false alarms, resulting in low accuracy in determining exceptions.
By obtaining the key field information of the pending log, the processing priority is determined based on the preset black, white and gray list rule base, the associated related logs are merged into a joint log, exception information is extracted and classified, and the target work ticket is generated to improve the accuracy of exception determination.
Through the black, white and gray list priority classification and log association mechanism, the exception classification results are accurately determined and work orders are generated, which improves the accuracy of log exception determination.
Smart Images

Figure CN120144738A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical fields of log analysis and anomaly monitoring, and particularly to a method, apparatus, electronic device, and readable storage medium for processing log data. Background Art
[0002] In international business systems, especially in cross-border transaction processing scenarios, log data involving multiple transaction types and multiple languages is involved. The log data has diverse formats and strong cross-system mobility, which makes it more difficult to detect and capture log anomalies.
[0003] Traditional log analysis and anomaly monitoring usually rely on manual inspection and processing of log data one by one to determine and report anomalies. Since manual processing takes a long time and there are problems such as missed anomalies and false alarms, the accuracy of log anomaly determination is relatively low. Summary of the Invention
[0004] In view of this, embodiments of this application at least provide a method, apparatus, electronic device, and readable storage medium for processing log data, which improves the accuracy of log anomaly determination.
[0005] This application mainly includes the following aspects:
[0006] In a first aspect, an embodiment of this application provides a method for processing log data, the method including:
[0007] Obtain multiple logs to be processed, and extract keyword field information from each log to be processed;
[0008] For any one of the logs to be processed, determine the processing priority of the log to be processed based on the keyword field information of the log to be processed and a preset black-white-gray list rule library;
[0009] Based on the keyword field information and processing priority of the log to be processed, determine associated logs having an association relationship with the log to be processed, and merge the associated logs and the log to be processed to obtain a combined log;
[0010] Based on the anomaly information extracted from the combined log, determine the anomaly classification result of the combined log;
[0011] For the log to be processed with the first priority, generate a target work order corresponding to the combined log based on the keyword field information and anomaly classification result of the combined log corresponding to the log to be processed.
[0012] In a second aspect, an embodiment of this application further provides a log data processing apparatus, the log data processing apparatus including:
[0013] A collection and parsing module, configured to obtain a plurality of logs to be processed, and extract key field information from each log to be processed;
[0014] A log filtering module, configured to determine the processing priority of any log to be processed based on the key field information of the log to be processed and a preset black / white / grey list rule library;
[0015] A log association module, configured to determine associated logs having an association relationship with the log to be processed based on the key field information and the processing priority of the log to be processed, and merge the associated logs and the log to be processed to obtain a combined log;
[0016] An exception extraction module, configured to determine an exception classification result of the combined log based on the exception information extracted from the combined log;
[0017] A work order generation module, configured to generate a target work order corresponding to the combined log based on the key field information and the exception classification result of the combined log corresponding to the log to be processed with the first priority of the processing priority.
[0018] In a third aspect, an embodiment of the present application further provides an electronic device, including: a processor, a memory, and a bus, where the memory stores machine-readable instructions executable by the processor. When the electronic device runs, communication is performed between the processor and the memory through the bus, and when the machine-readable instructions are run by the processor, the steps of the log data processing method described above are executed.
[0019] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium, where a computer program is stored on the computer-readable storage medium, and when the computer program is run by a processor, the steps of the log data processing method described above are executed.
[0020] A log data processing method, apparatus, electronic device, and readable storage medium provided by an embodiment of the present application obtain a plurality of logs to be processed, and extract key field information from each log to be processed; for any log to be processed, based on the key field information of the log to be processed and a preset black / white / grey list rule library, determine the processing priority of the log to be processed; based on the key field information and the processing priority of the log to be processed, determine associated logs that have an association relationship with the log to be processed, and merge the associated logs and the log to be processed to obtain a combined log; based on the abnormal information extracted from the combined log, determine the abnormal classification result of the combined log; for the log to be processed with the first priority, based on the key field information and the abnormal classification result of the combined log corresponding to the log to be processed, generate a target work order corresponding to the combined log. In this way, through the combination of black / white / grey list priority classification and a log association mechanism, the abnormal classification result can be accurately determined and a work order can be generated, improving the accuracy of log anomaly determination.
[0021] To make the above objects, features, and advantages of the present application more obvious and understandable, the following specifically gives preferred embodiments and, in conjunction with the accompanying drawings, provides a detailed description as follows. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] To more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments. It should be understood that the following drawings only show some embodiments of the present application and should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.
[0023] Figure 1 Shows a flowchart of a log data processing method provided by an embodiment of the present application;
[0024] Figure 2 Shows a functional module diagram of a log data processing apparatus provided by an embodiment of the present application;
[0025] Figure 3 Shows a second functional module diagram of a log data processing apparatus provided by an embodiment of the present application;
[0026] Figure 4 Shows a structural schematic diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0027] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part rather than all of the embodiments of this application. Components of the embodiments of this application described and illustrated in the accompanying drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely represents selected embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative efforts fall within the scope of protection of this application.
[0028] To enable those skilled in the art to use the content of this application, the following implementation manners are given in combination with a specific application scenario, "Log Analysis and Anomaly Monitoring of the International Business System". For those skilled in the art, the general principles defined here can be applied to other embodiments and application scenarios without departing from the spirit and scope of this application.
[0029] To facilitate the understanding of this application, the technical solutions provided by this application will be described in detail below in combination with specific embodiments.
[0030] Please refer to Figure 1 , Figure 1 which is a flowchart of a log data processing method provided by an embodiment of this application. As Figure 1 shown, the log data processing method provided by an embodiment of this application includes the following steps:
[0031] S101: Obtain multiple logs to be processed, and extract key field information from each log to be processed.
[0032] Here, first, the international business system needs to collect and record the original log data of anomalies in real time from a variety of different log sources, that is, obtain multiple logs to be processed; then, after the log collection is completed, log parsing is performed on each log to be processed, and key field information is extracted from each log to be processed. Among them, a variety of different log sources include API logs, SWIFT messages, etc.; SWIFT messages are standardized message formats used for transmitting financial transaction information and other services among international banks. The key field information is the field information that must be included in log data processing. For example, API logs must include the field information of the field "transaction_id"; SWIFT messages must include the field information of "reference_id", etc.
[0033] In the embodiments of this application, first, streaming log collection tools such as the distributed message queue Kafka and the log collection tool Fluentd are used to access multi-source log streams. According to different log sources and types, the logs are routed to independent message classification units (Topics) to ensure efficient and independent processing of the logs. Specifically, different Topics are configured in Kafka for each log source; for example: "topic_api_logs" is used to store API logs; "topic_swift_logs" is used to store SWIFT message logs; Fluentd is configured to collect logs, and the tags and routing rules of the logs are defined to ensure accurate diversion of logs from different sources. For example: API logs are pushed to "topic_api_logs" in Kafka through Fluentd; SWIFT messages are collected to "topic_swift_logs" in Kafka after being recorded in the system logs.
[0034] After accessing the log stream, by consuming the log stream in Kafka, the log format (such as JSON, SWIFT message, plain text, etc.) is judged according to the log characteristics, and it is routed to the corresponding processing logic. Specifically, by trying to parse the log data into JSON format to confirm whether it is valid JSON; by judging whether specific SWIFT message fields (such as: 20: or 32A:) are included in the log to determine that it is a SWIFT message; if the log is neither in JSON format nor contains the characteristic fields of a SWIFT message, it is classified as a plain text log.
[0035] Then, according to the log type, the corresponding parsing method is adopted to extract the key field information. Specifically, for JSON-formatted logs, a JSON parser is used to extract the key field information; for SWIFT message-formatted logs, the log content is parsed based on the rules of SWIFT message fields (such as using the delimiter ":" to extract fields) to extract the key field information; for plain text-formatted logs, regular expressions are used to match the key field information, such as timestamps, transaction IDs, and error messages.
[0036] After extracting the key field information, it is also possible to perform a preliminary verification on the extracted key field information to ensure that the obtained key field information is complete and conforms to the expected format. For logs with parsing failures or missing fields, record the error information and send the logs to the exception queue. Specifically, the verification rules include mandatory field checks and field format verification. Among them, the mandatory field check is used to ensure the existence of key fields. For example, API logs must include the "transaction_id" field; SWIFT logs must include the "reference_id" field. The field format verification is used to ensure that the formats of each key field are correct. For example, the timestamp field format is valid, and the amount field format is a valid number, etc.
[0037] Finally, combine the extracted key field information with the original log content of the log to be processed for storage, providing a unified input format for subsequent filtering, correlation, and exception analysis. Specifically, source information can be added to each log record, such as source = "API" or source = "SWIFT", and the original log content is retained to ensure that it can be traced back during subsequent processing or troubleshooting. Moreover, retaining the original log data can also be used for auditing and secondary analysis.
[0038] This step collects logs in real time from multiple log sources (such as API logs, SWIFT messages, etc.), automatically identifies their formats, and extracts key field information. By cleaning and structuring the data, it provides basic data for subsequent log data processing. Among them, efficient log stream access, format recognition, dynamic field parsing, and exception fault tolerance processing can ensure the efficiency, accuracy, and scalability of log collection and parsing.
[0039] S102. For any of the logs to be processed, determine the processing priority of the log to be processed based on the key field information of the log to be processed and a preset black / white / grey list rule library.
[0040] Here, after extracting key field information from each log to be processed, for any log to be processed, the system will match the rules in the preset black / white / grey list library based on the key field information of the log to be processed, and classify the priority of the log to be processed through the rules in the preset black / white / grey list rule library to determine the processing priority of the log to be processed. Among them, the rules in the preset black / white / grey list library are stored in the Remote Dictionary Server (Redis) in the form of key-value pairs. Each rule includes four core fields for defining the rule, including the rule unique identifier "rule_id", the log keyword to be matched "keyword", the processing priority "priority", and the rule exception trigger frequency "frequency". The processing priorities are divided into the first priority, the second priority, and the third priority in descending order. The rules in the preset black / white / grey list library are divided into three categories according to the processing priority. The rules with the first priority "high" are blacklist rules, the rules with the second priority "medium" are grey list rules, and the rules with the third priority "low" are white list rules.
[0041] The log to be processed will enter different priority classifications according to the processing priority determined by the matched rules and undergo subsequent processing. In this way, determining the processing order of the logs through priority classification can help reduce ineffective alarms and improve the response efficiency of the system. At the same time, classifying and screening the logs to be processed first can also ensure that subsequent log association and anomaly determination are more accurate and efficient.
[0042] S103. Based on the key field information and processing priority of the log to be processed, determine the associated logs that have an association relationship with the log to be processed, and merge the associated logs and the log to be processed to obtain a combined log.
[0043] Here, the system will associate the log to be processed with the associated logs that have an association relationship with the log to be processed from different sources according to the key field information and processing priority of the log to be processed, and merge the associated logs and the log to be processed to obtain a combined log to generate a complete transaction link. Among them, the associated logs are the other logs in the multiple logs to be processed except this log to be processed. In this way, generating a complete transaction link through log association can provide sufficient context data for subsequent extraction of abnormal information, determination of abnormal classification results, and work order generation, ensuring that anomalies can be identified more accurately.
[0044] S104. Based on the abnormal information extracted from the combined log, determine the abnormal classification result of the combined log.
[0045] Here, the log to be processed is parsed, classified, and associated to obtain a combined log. The system extracts abnormal information from the combined log and determines the abnormal classification result of the combined log based on the extracted abnormal information, so as to provide accurate guidance for subsequent work order generation and abnormal handling based on the generated work order.
[0046] S105. For the log to be processed with the first priority of processing priority, based on the keyword field information and abnormal classification result of the combined log corresponding to the log to be processed, generate a target work order corresponding to the combined log.
[0047] Here, for the log to be processed with the first priority of processing priority, based on the keyword field information and abnormal classification result of the combined log corresponding to the log to be processed, generate a target work order corresponding to the combined log. This step enables the system to automatically generate and distribute work orders corresponding to the abnormalities in the combined log of the logs to be processed with the first priority, which can ensure that the abnormalities corresponding to the first priority, that is, the blacklist rules, can be quickly located and processed, so as to improve the accuracy and efficiency of system abnormal handling.
[0048] Furthermore, for any of the logs to be processed, determining the processing priority of the log to be processed based on the keyword field information of the log to be processed and a preset black / white / gray list rule library includes:
[0049] Step a1. Query at least one first rule that matches the keyword field information of the log to be processed from the preset black / white / gray list rule library.
[0050] Here, query at least one first rule that matches the keyword field information of the log to be processed from the preset black / white / gray list rule library. For example, assume that the keyword field information of the log to be processed is {"transaction_id":"TX123","status":"failed","reason":"payment failure"}, and the rule in the preset black / white / gray list rule library is {"rule_id":"R001","keyword":"payment failure","priority":"high"}. Since "payment failure" in the rule matches "payment failure" in the key log information, this rule is determined as the first rule.
[0051] Step a2. Determine the highest processing priority among the processing priorities corresponding to the at least one first rule as the processing priority of the log to be processed; where multiple rules are set in the preset black / white / gray list rule library, and any one of the rules is configured with a processing priority.
[0052] Here, multiple keyword field data of the same log may simultaneously match the rules of multiple processing priorities. In this case, the highest processing priority among the processing priorities corresponding to at least one first rule is determined as the processing priority of the log to be processed; wherein, multiple rules are set in the preset black-and-white-and-gray list rule library, and any rule is configured with a processing priority.
[0053] Further, determining the associated log having an association relationship with the log to be processed based on the keyword field information and the processing priority of the log to be processed includes:
[0054] Step b1, determining whether there is a second log in at least one first log that has the same keyword field information as the log to be processed; the first log is other logs among the multiple logs to be processed that have the same processing priority as the log to be processed.
[0055] Here, determining whether there is a second log in at least one first log that has the same keyword field information as the log to be processed; wherein, the first log is other logs among the multiple logs to be processed that have the same processing priority as the log to be processed.
[0056] Step b2, if it exists, determining the second log as the associated log having an association relationship with the log to be processed.
[0057] Here, if it exists, determining the second log as the associated log having an association relationship with the log to be processed. In the embodiments of the present application, through the direct matching of keyword field information, the first-layer initial association result is generated. The matching priority is centered on field matching to ensure high precision. Specifically, match the "transaction_id" of the API log with the "reference_id" of the SWIFT message. If the field matching is successful, store the record in the initial association intermediate table to generate an association record. Among them, strict field matching is used preferentially to avoid redundant matching.
[0058] Step b3, if it does not exist, determining the associated log having an association relationship with the log to be processed based on the dynamic time window and the fuzzy matching algorithm.
[0059] Here, for the log to be processed whose keyword field information cannot be directly matched, the initial association intermediate table is complementarily matched through the dynamic time window and the fuzzy matching algorithm to ensure the coverage rate of cross-source association.
[0060] Specifically, the dynamic time window dynamically adjusts the time window threshold according to the business scenario. If the timestamp of a log differs from the timestamp of the log to be processed within the time window threshold, they are considered likely to belong to the same transaction link, and the log is determined as an associated log having an association relationship with the log to be processed. The fuzzy matching algorithm performs fuzzy matching based on the similarity of the prefix and suffix of the keyword field information. For example, "TX123" and "TX12345" are matched, and the string similarity algorithm, such as the Levenshtein algorithm, is used to improve the accuracy of fuzzy matching.
[0061] Furthermore, for the dynamic adjustment requirements, it is also possible to support the real-time update of the associated intermediate table, automatically correct incorrect associations or add new fields for supplementation. Specifically, regularly scan the associated intermediate table to supplement and correct records with missing fields or matching conflicts; when adding new log fields (such as API response time), dynamically update the intermediate table structure to ensure the integrity of the associated logs; use verification rules to detect redundant or incorrect association records and clean them up.
[0062] Furthermore, all associated logs can be stored in a unified intermediate table, providing a flexible query interface to support log link analysis. Specifically, the associated logs adopt a unified storage structure, enabling each combined log to record the complete information including the API and SWIFT message. The fields of the associated logs include: transaction ID "transaction_id", SWIFT message reference number "swift_reference", timestamp "timestamp", status "status", response time, error code, and other extended fields.
[0063] Furthermore, the associated logs also provide a query interface based on "transaction_id" or "swift_reference" to support transaction diagnosis at the link level.
[0064] Furthermore, determining the exception classification result of the combined log based on the exception information extracted from the combined log includes:
[0065] Step c1, perform semantic extraction on the combined log according to the trained natural language processing model to obtain the exception information of the combined log.
[0066] Here, using natural language processing (NLP) technology, semantic extraction is performed on the combined logs according to the trained natural language processing model to obtain the abnormal information of the combined logs. In the embodiments of the present application, the natural language processing model uses the BERT model, which supports in-depth semantic understanding of context, is suitable for parsing long texts and complex structures, can perform semantic analysis on the abnormal messages recorded in the combined logs and the context information of the abnormal messages, and can accurately extract the abnormal information of the combined logs. Moreover, the BERT model supports adaptation in multiple languages and multiple domains, and can be trained with diverse business log samples to ensure that it can adapt to different types of abnormal messages and be fine-tuned for multiple languages (such as English, French, etc.) and specific domain terms (such as keyword fields in SWIFT messages).
[0067] Step c2, generate an abnormal classification label for the combined logs based on the abnormal information, and determine the confidence level of the abnormal classification label.
[0068] Here, semantic extraction is performed on the combined logs using the trained natural language processing model, and the extracted semantic information is converted into an abnormal classification label, such as network timeout, payment rejection, etc. And a confidence level is attached to the abnormal classification label, and the value range of the confidence level is [0,1].
[0069] Step c3, determine the abnormal classification result of the combined logs according to the confidence level of the abnormal classification label.
[0070] Here, the abnormal classification result of the combined logs is determined according to the confidence level of the abnormal classification label.
[0071] Furthermore, the generated abnormal classification label and the corresponding confidence level can also be stored in a database, such as an Elasticsearch database, to provide real-time data query support services.
[0072] Furthermore, for the abnormal conditions where the abnormal classification labels frequently appear, the black / white / grey list rule library can also be automatically optimized to increase the recognition ability for such abnormal conditions. Specifically, the system will automatically generate new rules according to the abnormal classification label data of the historical logs.
[0073] Furthermore, the determining the abnormal classification result of the combined logs according to the confidence level of the abnormal classification label includes:
[0074] Step d1, if the confidence level of the abnormal classification label is less than the first threshold, then determine the classification result of the abnormal information manually as the abnormal classification result of the combined logs.
[0075] Here, if the confidence level of the abnormal classification label is less than the first threshold, it indicates that the confidence level of the abnormal classification label is relatively low, and the abnormality in the combined log is an unknown abnormality. It is necessary to push it for manual review for abnormal classification, and determine the classification result of the abnormal information feedback by the manual review as the abnormal classification result of the combined log. In the embodiment of the present application, the first threshold is 60%.
[0076] Step d2, if the confidence level of the abnormal classification label is within the target threshold range, determine the abnormal classification result of the combined log according to the number of times that the confidence level of the abnormal classification label in the historical log is within the target threshold range; the target threshold range is an interval greater than or equal to the first threshold and less than or equal to the second threshold.
[0077] Here, if the confidence level of the abnormal classification label is within the target threshold range, determine the abnormal classification result of the combined log according to the number of times that the confidence level of the abnormal classification label in the historical log is within the target threshold range; wherein, the target threshold range is an interval greater than or equal to the first threshold and less than or equal to the second threshold. In the embodiment of the present application, the second threshold is 90%, that is, the target threshold range is the interval of [60%, 90%]. Specifically, if the number of times that the confidence level of the abnormal classification label in the historical log is within the target threshold range is less than the preset number threshold, determine the abnormal classification label as the abnormal classification result of the combined log; if the number of times that the confidence level of the abnormal classification label in the historical log is within the target threshold range reaches the preset number threshold, it is necessary to push it for manual review for abnormal classification, and determine the classification result of the abnormal information feedback by the manual review as the abnormal classification result of the combined log.
[0078] Step d3, if the confidence level of the abnormal classification label is greater than the second threshold, determine the abnormal classification label as the abnormal classification result of the combined log.
[0079] Here, if the confidence level of the abnormal classification label is greater than the second threshold, that is, the confidence level is greater than 90%, it indicates that the confidence level of the abnormal classification label is very high, and the abnormal classification label can be directly determined as the abnormal classification result of the combined log.
[0080] In the embodiment of the present application, the natural language processing (NLP) technology is used to perform semantic analysis on the abnormal information in the log, extract the key abnormal information and generate an accurate classification label for each log to determine the abnormal classification result. It can ensure the accuracy of abnormal classification and further optimize the processing flow of log abnormalities.
[0081] Further, after generating the target work order corresponding to the combined log based on the keyword field information and the abnormal classification result of the combined log corresponding to the log to be processed, the method further includes:
[0082] Step e1: Adjust the preset black / white / grey list rule library based on the feedback information of the target work order and / or the rule exception trigger frequency to obtain an adjusted preset black / white / grey list rule library.
[0083] Here, the preset black / white / grey list rule library is dynamically adjusted according to the feedback information of the target work order and / or the rule exception trigger frequency to obtain an adjusted preset black / white / grey list rule library.
[0084] Step e2: Determine the processing priority of the log to be processed based on the keyword field information of the log to be processed and the preset black / white / grey list rule library, including:
[0085] Determine the processing priority of the log to be processed based on the keyword field information of the log to be processed and the adjusted preset black / white / grey list rule library.
[0086] Here, the adjusted rules are updated to Redis in real time to ensure that the system always uses the latest rules. Specifically, the processing priority of the log to be processed is determined based on the keyword field information of the log to be processed and the adjusted preset black / white / grey list rule library. To further improve the system's exception handling efficiency and reduce the repetition of work order generation and invalid warnings.
[0087] Furthermore, the adjustment of the preset black / white / grey list rule library based on the feedback information of the target work order and / or the rule exception trigger frequency to obtain an adjusted preset black / white / grey list rule library includes:
[0088] Step f1: If the feedback information of the target work order indicates that the exception corresponding to any second rule is resolved, adjust the second rule from the rule corresponding to the blacklist to the rule corresponding to the white list.
[0089] Here, the rules in the preset black / white / grey list rule library are adjusted according to the work order status feedback, and the rule corresponding to the blacklist with the corresponding exception repaired is downgraded to the rule corresponding to the white list. Specifically, if the feedback information of the target work order indicates that the exception corresponding to any second rule is resolved, adjust the second rule from the rule corresponding to the blacklist to the rule corresponding to the white list; where the second rule is the rule corresponding to the blacklist in the preset black / white / grey list rule library. For example, taking the rule corresponding to the blacklist {"rule_id":"R001","keyword":"payment failure","priority":"high"} as an example, if the feedback information of the target work order indicates that the exception corresponding to this rule is resolved, then adjust this rule to {"rule_id":"R001","keyword":"payment failure","priority":"low"}, that is, the rule corresponding to the white list.
[0090] Step f2, if the rule exception trigger frequency of any third rule reaches the preset trigger frequency threshold, adjust the third rule from a gray list rule to a black list rule.
[0091] Here, by monitoring the to-be-processed logs, the rule exception trigger frequency is counted, and the rules in the preset black / white / gray list rule library are adjusted according to the rule exception trigger frequency. Specifically, if the rule exception trigger frequency of any third rule reaches the preset trigger frequency threshold, the third rule is adjusted from a gray list rule to a black list rule; among them, the third rule is the rule corresponding to the gray list in the preset black / white / gray list rule library. For example, taking the rule corresponding to the gray list {"rule_id":"R003","keyword":"timeout","priority":"medium","frequency":10} as an example, if the rule exception trigger frequency of this rule reaches the preset trigger frequency threshold of 15, then this rule is adjusted to {"rule_id":"R003","keyword":"timeout","priority":"high"}, that is, the rule corresponding to the black list.
[0092] Furthermore, all rule adjustment operations can be recorded in the log system for auditing and tracing. Taking the example in step f2 as an example, the rule adjustment operation can be recorded as:
[0093] “[2025-01-08 12:00:00]Rule R003 priority updated from medium to high due to 15 triggers in 1 hour”.
[0094] Furthermore, the target work order includes the required fields of the target work order and the exception resolution suggestions dynamically generated based on the exception classification result and the preset exception resolution suggestion rule library. Among them, the required fields of the target work order include Transaction ID: used to identify the abnormal business transaction; Exception classification result: used to indicate the exception classification result; Exception occurrence time: used to locate the problem location; and Context log information: used to provide more background support for the responsible team. Among them, the rules in the preset exception resolution suggestion rule library can be dynamically and continuously optimized. Examples of exception resolution suggestions are as follows: Payment failure suggestion: Contact the payment gateway team to check the service availability; Database timeout suggestion: Check slow queries or database lock problems.
[0095] Furthermore, the target work order can be formatted into JSON format for easy storage and distribution.
[0096] Furthermore, the method further includes:
[0097] For the to-be-processed logs with the first priority and the second priority of processing priority, no work orders are generated, and only the rule exception trigger frequency is recorded.
[0098] Specifically, for the to-be-processed logs with the second priority, record the rule exception trigger frequency and observe whether the priority needs to be adjusted; for the to-be-processed logs with the third priority, only record the trigger frequency to reduce resource occupancy.
[0099] Furthermore, after generating the target work order, the method further includes:
[0100] According to the exception classification result, use the dynamic distribution algorithm and the load balancing mechanism to push the work order to the queue of the corresponding responsible team to avoid overloading a single team. Specifically, establish the mapping relationship between the exception type and the responsible team. Use the round-robin or priority-based load balancing algorithm to dynamically adjust the priority of the distribution queue. For high-priority exceptions, give priority to distributing to the queue with a lower current load. Use message queue tools (such as RabbitMQ, Kafka) to push the work order to the team queue in real time. Each responsible team maintains an independent queue, and the queue messages support the confirmation and retry mechanisms. After receiving the work order, the responsible team confirms the reception status through the callback interface, and the system marks the work order as "distributed". Provide a REST API interface for the responsible team to update the work order status and processing feedback in real time. And design a status field for each work order, including: Unprocessed: The work order has been generated but not yet distributed; Processing: The responsible team has received the work order and is processing it; Resolved: The responsible team has completed the processing and the problem has been solved; Unable to solve: Upgrade processing is required.
[0101] An embodiment of the present application provides a log data processing method, including: obtaining a plurality of to-be-processed logs, and extracting keyword field information from each to-be-processed log; for any to-be-processed log, based on the keyword field information and the preset black / white / gray list rule library, determine the processing priority of the to-be-processed log; based on the keyword field information and the processing priority of the to-be-processed log, determine the associated log having an associated relationship with the to-be-processed log, and merge the associated log and the to-be-processed log to obtain a combined log; based on the exception information extracted from the combined log, determine the exception classification result of the combined log; for the to-be-processed log with the first priority, generate a target work order based on the keyword field information and the exception classification result of the corresponding combined log. In this way, through the combination of black / white / gray list priority classification and the log association mechanism, the exception classification result can be accurately determined and the work order can be generated, improving the accuracy of log exception determination.
[0102] Based on the same application concept, in the embodiments of the present application, a log data processing device corresponding to the log data processing method provided in the above embodiments is further provided. Since the principle of problem-solving of the device in the embodiments of the present application is similar to that of the log data processing method in the above embodiments of the present application, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be described again.
[0103] Please refer to Figure 2 , Figure 2 which is one of the functional module diagrams of a log data processing device provided in the embodiments of the present application. As Figure 2 shown, the log data processing device 200 includes:
[0104] An acquisition and parsing module 210, configured to obtain a plurality of logs to be processed, and extract key field information from each log to be processed.
[0105] A log screening module 220, configured to determine the processing priority of any one of the logs to be processed based on the key field information of the log to be processed and a preset black / white / gray list rule library.
[0106] A log association module 230, configured to determine an associated log having an association relationship with the log to be processed based on the key field information and the processing priority of the log to be processed, and merge the associated log and the log to be processed to obtain a combined log.
[0107] An exception extraction module 240, configured to determine an exception classification result of the combined log based on the exception information extracted from the combined log.
[0108] A work order generation module 250, configured to generate a target work order corresponding to the combined log based on the key field information and the exception classification result of the combined log corresponding to the log to be processed with the first priority of the processing priority.
[0109] Further, when the log screening module 220 is used to determine the processing priority of any one of the logs to be processed based on the key field information of the log to be processed and a preset black / white / gray list rule library, the log screening module 220 is specifically configured to:
[0110] Query at least one first rule matching the key field information of the log to be processed from the preset black / white / gray list rule library;
[0111] Determine the highest processing priority among the processing priorities corresponding to the at least one first rule as the processing priority of the log to be processed;
[0112] Among them, multiple rules are set in the preset black-and-white-gray list rule library, and any one of the rules is configured with a processing priority.
[0113] Further, when the log association module 230 is used to determine the associated log having an association relationship with the to-be-processed log based on the keyword field information and processing priority of the to-be-processed log, the log association module 230 specifically is used for:
[0114] Judging whether there is a second log in at least one first log that has the same keyword field information as the to-be-processed log; the first log is other logs among the multiple to-be-processed logs that have the same processing priority as the to-be-processed log;
[0115] If there is, determining the second log as the associated log having an association relationship with the to-be-processed log;
[0116] If not, determining the associated log having an association relationship with the to-be-processed log based on a dynamic time window and a fuzzy matching algorithm.
[0117] Further, when the anomaly extraction module 240 is used to determine the anomaly classification result of the combined log based on the anomaly information extracted from the combined log, the anomaly extraction module 240 specifically is used for:
[0118] Performing semantic extraction on the combined log according to the trained natural language processing model to obtain the anomaly information of the combined log;
[0119] Generating an anomaly classification label for the combined log based on the anomaly information and determining the confidence level of the anomaly classification label;
[0120] Determining the anomaly classification result of the combined log according to the confidence level of the anomaly classification label.
[0121] Further, when the anomaly extraction module 240 is used to determine the anomaly classification result of the combined log according to the confidence level of the anomaly classification label, the anomaly extraction module 240 specifically is used for:
[0122] If the confidence level of the anomaly classification label is less than the first threshold, determining the classification result of the manual classification of the anomaly information as the anomaly classification result of the combined log;
[0123] If the confidence level of the anomaly classification label is within the target threshold interval, determining the anomaly classification result of the combined log according to the number of times that the confidence level of the anomaly classification label in the historical log is within the target threshold interval; the target threshold interval is an interval greater than or equal to the first threshold and less than or equal to the second threshold;
[0124] If the confidence level of the abnormal classification label is greater than the second threshold, then determine the abnormal classification label as the abnormal classification result of the combined log.
[0125] Further, please refer to Figure 3 , Figure 3 which is the second functional module diagram of a log data processing device provided by an embodiment of the present application. As Figure 3 shown, the log data processing device 200 further includes:
[0126] A feedback optimization module 260, configured to adjust the preset black-and-white-and-gray list rule library based on the feedback information of the target work order and / or the rule exception trigger frequency, so as to obtain an adjusted preset black-and-white-and-gray list rule library.
[0127] The log screening module 220 is further configured to:
[0128] Determine the processing priority of the to-be-processed log based on the keyword field information of the to-be-processed log and the adjusted preset black-and-white-and-gray list rule library.
[0129] Further, when the feedback optimization module 260 is configured to adjust the preset black-and-white-and-gray list rule library based on the feedback information of the target work order and / or the rule exception trigger frequency to obtain an adjusted preset black-and-white-and-gray list rule library, the feedback optimization module 260 is specifically configured to:
[0130] If the feedback information of the target work order indicates that the exception corresponding to any second rule is resolved, adjust the second rule from the rule corresponding to the blacklist to the rule corresponding to the white list;
[0131] If the rule exception trigger frequency of any third rule reaches the preset trigger frequency threshold, adjust the third rule from the gray list rule to the blacklist rule.
[0132] An embodiment of the present application provides a log data processing device, including a collection and parsing module, configured to obtain a plurality of logs to be processed and extract key field information from each log to be processed; a log screening module, configured to determine the processing priority of any log to be processed based on the key field information of the log to be processed and a preset black / white / grey list rule library; a log association module, configured to determine associated logs having an association relationship with the log to be processed based on the key field information and the processing priority of the log to be processed, and merge the associated logs and the log to be processed to obtain a combined log; an exception extraction module, configured to determine an exception classification result of the combined log based on the exception information extracted from the combined log; and a work order generation module, configured to generate a target work order corresponding to the combined log based on the key field information and the exception classification result of the combined log corresponding to the log to be processed with the first priority. In this way, through the combination of black / white / grey list priority classification and a log association mechanism, the exception classification result can be accurately determined and a work order can be generated, improving the accuracy of log exception determination.
[0133] Based on the same inventive concept, please refer to Figure 4 , Figure 4 which is a schematic structural diagram of an electronic device provided by an embodiment of the present application. As Figure 4 shown, the electronic device 400 includes: a processor 410, a memory 420, and a bus 430.
[0134] The memory 420 stores machine-readable instructions executable by the processor 410. When the electronic device 400 runs, the processor 410 communicates with the memory 420 through the bus 430. When the machine-readable instructions are run by the processor 410, the steps of the log data processing method provided in the above embodiment are executed. The specific implementation manner can refer to the method embodiment and will not be elaborated here.
[0135] Based on the same inventive concept, an embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, the steps of the log data processing method provided in the above embodiment are executed. The specific implementation manner can refer to the method embodiment and will not be elaborated here.
[0136] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the above-described device and unit can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated here.
[0137] In the embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For another example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some communication interfaces. The indirect coupling or communication connection of the devices or units can be in electrical, mechanical or other forms.
[0138] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0139] In addition, each functional unit in the embodiments provided in the present application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.
[0140] If the functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks or optical discs that can store program codes.
[0141] It should be noted that similar reference numerals and letters represent similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. In addition, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.
[0142] Finally, it should be noted that the above-described embodiments are only specific implementation manners of the present application, used to illustrate the technical solutions of the present application, rather than limiting it. The protection scope of the present application is not limited thereto. Although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that any person skilled in the art within the technical scope disclosed by the present application can still modify the technical solutions described in the foregoing embodiments, or can easily think of changes, or perform equivalent replacements on some of the technical features; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application. All should be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims described.
Claims
1. A log data processing method, characterized in that: The method comprises: Obtain multiple logs to be processed, and extract key field information from each log to be processed; For any of the logs to be processed, based on the key field information of the logs to be processed and a preset black, white and grey list rule library, determine the processing priority of the logs to be processed; Based on the key field information and processing priority of the log to be processed, determining an associated log associated with the log to be processed, and merging the associated log and the log to be processed to obtain a joint log; Determining an abnormality classification result of the joint log based on the abnormality information extracted from the joint log; For the to-be-processed log whose processing priority is the first priority, a target work order corresponding to the joint log is generated based on key field information and anomaly classification results of the joint log corresponding to the to-be-processed log.
2. The log data processing method according to claim 1, characterized in that: For any of the logs to be processed, determining the processing priority of the logs to be processed based on the key field information of the logs to be processed and a preset black, white and gray list rule base includes: Query at least one first rule matching the key field information of the log to be processed from the preset black, white and gray list rule library; Determine the highest processing priority among the processing priorities corresponding to the at least one first rule as the processing priority of the log to be processed; There are multiple rules set in the preset black, white and grey list rule library, and any of the rules is configured with a processing priority.
3. The log data processing method according to claim 1, characterized in that: The determining, based on the key field information and the processing priority of the log to be processed, an associated log having an associated relationship with the log to be processed includes: Determine whether there is a second log in at least one first log that has the same key field information as the log to be processed; the first log is another log in the multiple logs to be processed that has the same processing priority as the log to be processed; If so, determining the second log as an associated log associated with the log to be processed; If not, then based on the dynamic time window and fuzzy matching algorithm, determine the associated log that has an associated relationship with the log to be processed.
4. The log data processing method according to claim 1, characterized in that: The determining, based on the abnormal information extracted from the joint log, the abnormal classification result of the joint log includes: Performing semantic extraction on the joint log according to the trained natural language processing model to obtain abnormal information of the joint log; generating an anomaly classification label for the joint log based on the anomaly information, and determining a confidence level of the anomaly classification label; An anomaly classification result of the joint log is determined according to the confidence of the anomaly classification label.
5. The log data processing method according to claim 4, characterized in that: Determining the abnormality classification result of the joint log according to the confidence of the abnormality classification label includes: If the confidence of the abnormal classification label is less than a first threshold, determining the manual classification result of the abnormal information as the abnormal classification result of the joint log; If the confidence of the abnormal classification label is within the target threshold interval, the abnormal classification result of the joint log is determined according to the number of times the confidence of the abnormal classification label in the historical log is within the target threshold interval; the target threshold interval is an interval greater than or equal to the first threshold and less than or equal to the second threshold; If the confidence of the abnormal classification label is greater than the second threshold, the abnormal classification label is determined as the abnormal classification result of the joint log.
6. The log data processing method according to claim 1, characterized in that: After generating a target work order corresponding to the joint log based on key field information and anomaly classification results of the joint log corresponding to the log to be processed, the method further includes: Based on the feedback information of the target work order and / or the rule abnormality triggering frequency, the preset black, white and gray list rule library is adjusted to obtain an adjusted preset black, white and gray list rule library; Determining the processing priority of the log to be processed based on the key field information of the log to be processed and the preset black, white and gray list rule base includes: Based on the key field information of the log to be processed and the adjusted preset black, white and gray list rule base, the processing priority of the log to be processed is determined.
7. The log data processing method according to claim 6, characterized in that: The step of adjusting the preset black, white and gray list rule base based on the feedback information of the target work order and / or the rule abnormality triggering frequency to obtain the adjusted preset black, white and gray list rule base includes: If the feedback information of the target work order indicates that the exception corresponding to any second rule is resolved, the second rule is adjusted from a rule corresponding to the blacklist to a rule corresponding to the whitelist; If the rule abnormal triggering frequency of any third rule reaches a preset triggering frequency threshold, the third rule is adjusted from a gray list rule to a black list rule.
8. A log data processing device, characterized in that: The log data processing device comprises: The collection and analysis module is used to obtain multiple logs to be processed and extract key field information from each log to be processed; A log screening module, for determining the processing priority of any of the logs to be processed based on key field information of the logs to be processed and a preset black, white and gray list rule library; A log association module, configured to determine an associated log associated with the log to be processed based on key field information and processing priority of the log to be processed, and merge the associated log and the log to be processed to obtain a joint log; An anomaly extraction module, used to determine an anomaly classification result of the joint log based on the anomaly information extracted from the joint log; The work order generation module is used to generate a target work order corresponding to the joint log for the to-be-processed log with a processing priority of the first priority based on key field information of the joint log corresponding to the to-be-processed log and an abnormal classification result.
9. An electronic device, characterized in that: include: A processor, a memory and a bus, wherein the memory stores machine-readable instructions executable by the processor, and when the electronic device is running, the processor and the memory communicate through the bus, and the machine-readable instructions are executed by the processor to execute the steps of the log data processing method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the log data processing method according to any one of claims 1 to 7 are executed.
Citation Information
Patent Citations
Equipment failure detection method, system and related components
CN109388623A
Log storage method and device, computer equipment and storage medium
CN114116614A
Abnormal log detection method and device
CN114968633A
Log analysis method and device, equipment and storage medium
CN115329748A
Abnormal log processing method and device
CN117648214A