Fault plan recommendation method and device based on real-time log stream association matching
By collecting and processing the log data of large-scale distributed systems in real time, identifying the fault event sequence and searching matching fault plans, the problem that the existing technology cannot automatically identify fault events is solved, and efficient and accurate fault handling is achieved.
Patent Information
- Application Number
- CN202311555871.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-21
- Publication Date
- 2025-05-23
AI Technical Summary
The existing technology cannot automatically identify fault events and recommend fault plans, making it difficult to meet the needs of large-scale distributed systems.
By collecting and preprocessing the log data of the target system in real time, Flink is used to process large-scale concurrent log data flow, using preset complex event processing algorithms to identify the fault event sequence, and using the Doris database to retrieve the matching fault plan, and finally present the fault event sequence and plan through visual interaction.
It realizes efficient and automated fault detection of large-scale data, improves the efficiency and accuracy of fault handling, and can identify and resolve faults more quickly.
Smart Images

Figure CN120029796A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing technology, and in particular to a method and device for recommending a fault plan based on real-time log stream association matching. Background Art
[0002] In large-scale distributed systems, failures are inevitable. In order to identify and resolve failures in a timely manner, maintenance personnel usually need to spend a lot of time and energy analyzing system logs, finding the cause of the failure, and taking corresponding corrective measures. However, with the continuous expansion of system scale and the rapid growth of log data, traditional manual analysis methods can no longer meet the needs of real-time fault handling.
[0003] Therefore, a new fault plan active recommendation method and system is needed to automatically identify potential fault events and recommend appropriate solutions based on the existing fault plans in the system, thereby improving the efficiency and accuracy of fault handling. Summary of the invention
[0004] In view of this, an embodiment of the present invention provides a fault plan recommendation method and device based on real-time log stream association matching to eliminate or improve one or more defects existing in the prior art and solve the problem that the prior art cannot automatically identify fault events and recommend fault plans.
[0005] One aspect of the present invention provides a method for recommending a fault plan based on real-time log stream association matching, the method comprising the following steps:
[0006] Collect log data of the target system from one or more data sources, format and pre-process it through the data interceptor, and then store it in the Kafka message queue;
[0007] Read the log data in the Kafka message queue and convert it into a log data stream in a first set format for Flink processing and calling;
[0008] Using a preset complex event processing algorithm to monitor the log data stream and identify a preset fault event sequence in the log data stream;
[0009] Using the Doris database to retrieve a fault plan that matches the fault event sequence;
[0010] The fault event sequence and the fault contingency plan are visually presented, and the handling process of the fault event sequence is tracked and recorded.
[0011] In some embodiments, collecting log data of a target system in one or more data sources includes: using a Flume component to collect the log data of the target system in the data source.
[0012] In some embodiments, the method further includes: using a FlinkCDC component to read the log data in the Kafka message queue, and the preset complex event processing algorithm is a FlinkCEP component.
[0013] In some embodiments, after collecting the log data of the target system from one or more data sources, the method further includes: using a data interceptor to perform format normalization preprocessing on the log data.
[0014] In some embodiments, identifying a preset fault event sequence in the log data stream includes:
[0015] Setting corresponding time windows for multiple types of fault event sequences to identify corresponding fault event sequences within the time windows;
[0016] And / or, setting a value range constraint, a time interval constraint between events, and a time sequence constraint to detect and identify the fault event sequence.
[0017] In some embodiments, the fault prevention plan includes a handling process and tracking requirements for the fault event sequence.
[0018] In some embodiments, the Doris database is used to retrieve the fault plan that matches the fault event sequence, and the Doris database selects and outputs the optimal fault plan corresponding to the fault event sequence based on fault characteristics, historical matching results, and practical feedback, including:
[0019] The Doris database selects a plurality of candidate fault plans according to the fault characteristics of the fault event sequence, wherein the fault characteristics include numerical parameters of a plurality of attributes;
[0020] The Doris database queries the historical matching results, and uses the number of times each candidate fault plan is selected for the fault event sequence as a first score;
[0021] The Doris database queries the practice feedback of each candidate fault plan, and calculates a second score based on the practice feedback result;
[0022] Taking a weighted sum of the first score and the second score corresponding to each candidate fault contingency plan to obtain a third score;
[0023] The candidate fault pre-plan with the third highest score is used as the optimal fault pre-plan for the fault event sequence.
[0024] In some embodiments, the method further includes: viewing, selecting, interacting with and updating the fault plan through visual presentation.
[0025] On the other hand, the present invention also provides a fault plan recommendation device based on real-time log stream association matching, including a processor and a memory, wherein the memory stores computer instructions, and the processor is used to execute the computer instructions stored in the memory. When the computer instructions are executed by the processor, the device implements the steps of the above method.
[0026] On the other hand, the present invention also provides a computer-readable storage medium having a computer program stored thereon, characterized in that the program implements the steps of the above method when executed by a processor.
[0027] The beneficial effects of the present invention are at least:
[0028] The fault plan recommendation method and device based on real-time log stream association matching of the present invention collects the log data of the target system and stores it in the Kafka distributed streaming media platform for management, processes large-scale concurrent log data streams based on Flink, monitors the log data streams through a preset complex event processing algorithm to automatically identify fault event sequences, and uses the high-speed retrieval and matching capabilities of the Doris database to accurately query the fault plans corresponding to the fault event sequences, and ensures the viewing, selection and operation requirements of the operation and maintenance personnel for the fault plans through visual interaction, thereby achieving efficient and accurate fault detection and disposal.
[0029] Furthermore, based on the Flume component to collect log data, the FlinkCDC component is used to read the log data, and the FlinkCEP component is used to analyze the fault time series, thus achieving efficient and automatic fault detection for large-scale data.
[0030] Furthermore, by setting time windows, value range constraints, time interval constraints between events and time sequence constraints for different fault event sequence detection, more complex and systematic fault detection capabilities can be achieved.
[0031] Additional advantages, purposes, and features of the present invention will be described in part in the following description, and will become apparent to those skilled in the art after studying the following, or may be learned from the practice of the present invention. The purposes and other advantages of the present invention may be achieved and obtained by the structures specifically indicated in the specification and the accompanying drawings.
[0032] Those skilled in the art will appreciate that the objectives and advantages that can be achieved with the present invention are not limited to the above specific description, and the above and other objectives that can be achieved by the present invention will be more clearly understood from the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] The drawings described herein are used to provide a further understanding of the present invention, constitute a part of the present application, and do not constitute a limitation of the present invention. In the drawings:
[0034] Figure 1 The present invention is a flowchart of a method for recommending a fault plan based on real-time log stream correlation matching according to an embodiment of the present invention.
[0035] Figure 2 The present invention is a flowchart of matching fault plans in a fault plan recommendation method based on real-time log stream association matching according to an embodiment of the present invention.
[0036] Figure 3 The present invention is a flowchart of a method for proactively recommending fault plans based on real-time log stream correlation matching according to an embodiment of the present invention.
[0037] Figure 4 The present invention is a flowchart of a method for proactively recommending fault plans based on real-time log stream correlation matching according to another embodiment of the present invention.
[0038] Figure 5 for Figure 4 Specific flow chart of step 110 in FIG.
[0039] Figure 6 for Figure 4 Specific flow chart of step 120 in FIG. DETAILED DESCRIPTION
[0040] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the embodiments and the accompanying drawings. Here, the illustrative embodiments of the present invention and their descriptions are used to explain the present invention, but are not intended to limit the present invention.
[0041] It should also be noted that, in order to avoid obscuring the present invention due to unnecessary details, only structures and / or processing steps closely related to the solutions according to the present invention are shown in the accompanying drawings, while other details that are not closely related to the present invention are omitted.
[0042] It should be emphasized that the term “include / comprises” when used herein refers to the presence of features, elements, steps or components, but does not exclude the presence or addition of one or more other features, elements, steps or components.
[0043] It should also be noted that, unless otherwise specified, the term “connection” herein may refer not only to a direct connection but also to an indirect connection involving an intermediate.
[0044] In existing large-scale distributed systems, fault detection and handling require operation and maintenance personnel to spend a lot of time and energy to achieve. As the amount of data continues to increase, relying solely on manpower can no longer meet the needs of fault detection and handling.
[0045] Therefore, the present invention provides a fault plan recommendation method based on real-time log stream association matching, such as Figure 1 As shown, the method includes the following steps S101 to S105:
[0046] Step S101: log data of the target system from one or more data sources is collected, formatted and preprocessed by a data interceptor, and then stored in a Kafka message queue.
[0047] Step S102: read the log data in the Kafka message queue and convert it into a log data stream in a first set format for Flink processing and calling.
[0048] Step S103: using a preset complex event processing algorithm to monitor the log data stream and identify a preset fault event sequence in the log data stream.
[0049] Step S104: using the Doris database to retrieve a fault plan that matches the fault event sequence.
[0050] Step S105: Visually present the fault event sequence and the fault contingency plan, and track and record the handling process of the fault event sequence.
[0051] In step S101, the data source may be multiple, and the format of the data of each system in the data source may also be diversified, including text format, binary format, table and database format, image multimedia format, etc. In the operation of these systems, an important mechanism for recording database operations through log data is used for recovery, backup and transaction management. In this embodiment, the log data is collected through the Flume component. Flume is a distributed system for collecting, aggregating and moving a large amount of log data, which is designed to provide reliable and scalable log data transmission. The log data collected by Flume is stored based on Kafka, which is a distributed stream processing platform designed to process large-scale data streams. As a message queue, Kafka supports producers to publish messages to topics, and consumers can subscribe to these topics to receive messages. Its uniqueness lies in providing a persistent storage mechanism to save messages on disk to ensure that data will not be lost in consumer processing, thereby meeting the system with high reliability requirements. Kafka is also known for its high throughput, horizontal scalability and real-time performance. It can process thousands of messages, support horizontal expansion under growing data loads, and realize instant consumption after producers generate messages.
[0052] In some embodiments, after collecting the log data of the target system from one or more data sources, the process further includes: using a data interceptor to pre-process the log data in a standardized format. Data interceptors are generally used for data cleaning and pre-processing to ensure that the input data meets specific requirements and standards. Data interceptors can intercept, check, and correct data in the data stream to prevent erroneous or non-compliant data from entering the system.
[0053] In step S102, the FlinkCDC component is used to read the log data in the Kafka message queue. Flink CDC is used to achieve the integrated reading capability of the full and incremental data from the database, and with the help of Flink (a distributed stream processing engine that can efficiently perform real-time data processing tasks on large-scale data sets) excellent pipeline capabilities and rich upstream and downstream ecology, it supports capturing changes in multiple databases and synchronizing these changes to downstream storage in real time.
[0054] First, set the format to meet the processing requirements of Flink. Specifically, to use FlinkCDC to read data from Kafka, you first need to configure Kafka as the data source of the Flink application, including parameters such as the address, topic, and consumer group of the specified Kafka. FlinkCDC uses different Change Data Capture methods to capture changes from Kafka. CDC may obtain changes in the form of incremental pulls or by subscribing to Kafka topics. Define the format of log data in Kafka, understand the structure, field names, and data types of log messages, and choose a serialization and deserialization method based on actual conditions, such as using Avro, JSON, or other common serialization formats. Once the log data is read from Kafka, you need to implement data conversion logic to convert the raw data in Kafka into a data stream that Flink can process, including steps such as data cleaning, field selection, and data type conversion to ensure that the data meets the requirements of the Flink application.
[0055] In step S103, the complex event processing algorithm is preset as the FlinkCEP component. The FlinkCEP component is specially designed to process and analyze real-time streaming data of complex event patterns. The FlinkCEP component allows users to detect complex event patterns in data streams by defining patterns, rules, and time windows. These patterns may include specific sequences or conditions that occur in event streams. Specifically, by establishing identification rules to detect certain characteristic parameters in the log data stream, the occurrence of certain specific events is monitored, and the occurrence sequence of specific events is limited as a fault event sequence that can characterize the fault state.
[0056] In some embodiments, identifying a preset fault event sequence in a log data stream includes: setting corresponding time windows for multiple types of fault event sequences to identify the corresponding fault event sequence within the time window. The goal is to consider a fault only when the fault event sequence occurs within a specific time window, which provides a more diversified system fault detection capability and can adapt to more application scenarios.
[0057] And / or, set a value range constraint, a time interval constraint between events, and a time sequence constraint to detect and identify the fault event sequence. In this embodiment, for the detection of the fault event sequence, not only the occurrence sequence of the event is considered, but also more detailed limitations are made to accurately regulate the fault elements to achieve more accurate fault detection.
[0058] In steps S104 and S105, the fault plan is matched to the identified fault event sequence. In some embodiments, the fault plan includes a handling process and tracking requirements for the fault event sequence. Among them, the handling process may include parameter adjustment and repair, version rollback, data restoration, replay and recalculation, data center switching, cross-regional backup and business switching, etc. Tracking requirements include monitoring and tracking the main functions involved in the fault, and establishing logs to record the handling operations and related business status of the fault.
[0059] Doris uses a distributed storage and computing architecture that can be horizontally expanded to process large-scale data. This architecture enables Doris to have high throughput and low latency, and can support fast data retrieval. Doris's data model supports multi-dimensional analysis and is suitable for OLAP scenarios (Online Analytical Processing). Multiple dimensions can be defined as needed, and data can be slicing and dicing can be performed flexibly to support complex analytical queries.
[0060] Visualization is mainly used to achieve efficient interaction between the system and operation and maintenance personnel, feedback detection results and receive operation decision instructions. In addition, it can realize feedback on the operating status, deployment, modification and update iteration of the fault plan. In some embodiments, the method also includes: viewing, selecting, interacting and updating the fault plan through visual display.
[0061] In some embodiments, in step S10, i.e., using the Doris database to retrieve the fault plan that matches the fault event sequence, the Doris database selects and outputs the optimal fault plan corresponding to the fault event sequence based on the fault characteristics, historical matching results and practical feedback, such as Figure 2 As shown, steps S201 to S205 are included:
[0062] Step S201: the Doris database selects a plurality of candidate fault plans according to the fault characteristics of the fault event sequence, where the fault characteristics include numerical parameters of a plurality of attributes.
[0063] Step S202: query the Doris database for historical matching results, and use the number of times each candidate fault plan is selected for the fault event sequence as the first score.
[0064] Step S203: the Doris database queries the practical feedback of each candidate fault contingency plan, and calculates a second score based on the practical feedback result.
[0065] Step S204: performing weighted summation of the first score and the second score corresponding to each candidate fault contingency plan to obtain a third score.
[0066] Step S205: taking the candidate fault contingency plan with the third highest score as the optimal fault contingency plan for the fault event sequence.
[0067] In steps S201 to S205, for the same fault time sequence, multiple candidate fault plans that meet the requirements may be matched. In the process of selecting candidate fault plans, the optimal fault plan is selected based on the number of times it has been adopted in the past and the effect finally achieved by implementing the plan. Among them, the first score can be the number of times the fault plan has been selected and executed for a specific fault event sequence. The second score is obtained based on the results of practical feedback. The fault elimination rate can be directly used as the second score, or the feedback score of the operation and maintenance personnel can be used as the second score.
[0068] On the other hand, the present invention also provides a fault plan recommendation device based on real-time log stream association matching, including a processor and a memory, wherein the memory stores computer instructions, and the processor is used to execute the computer instructions stored in the memory. When the computer instructions are executed by the processor, the device implements the steps of the above method.
[0069] On the other hand, the present invention also provides a computer-readable storage medium having a computer program stored thereon, characterized in that the program implements the steps of the above method when executed by a processor.
[0070] The present invention is described below in conjunction with a specific embodiment:
[0071] This embodiment provides a fault plan active recommendation method based on real-time log stream association matching. First, the Flume component is used to collect the log data of each system, and the interceptor is used for formatting preprocessing to ensure reliable log data collection and storage. The preprocessed log data is sent to the Kafka message queue for subsequent processing. Next, FlinkCDC is used to read the log data from Kafka and convert it into a data stream that can be processed by Flink to achieve real-time log data processing. During the data stream processing process, the Flink component is used to clean and format the data to ensure the accuracy and consistency of the data. For the real-time log data stream, FlinkCEP technology is used for complex event processing to identify event sequences that may cause faults. By defining appropriate event patterns and rules, the system can effectively match the event sequence related to the fault. Once the event sequence that may cause the fault is matched, the real-time search function is immediately triggered, and Doris is used to quickly retrieve the fault plan related to the matching event. The system can recommend appropriate solutions based on the fault characteristics and the existing fault knowledge base. In order to intuitively display the fault plan, the found fault plan is displayed to the operation and maintenance personnel through a visual interface to help them understand the cause and solution of the fault. Operations and maintenance personnel can interact with the system through the interface and track the progress of troubleshooting in real time.
[0072] See also Figure 3 The method for actively recommending fault plans based on real-time log stream association matching specifically includes the following contents:
[0073] Step 100: Format and preprocess the collected log data, and store the preprocessed data in Kafka.
[0074] In step 100, the data of different platform systems can be connected to the alarm processing device by configuring Flume to monitor the data source, and format and clean it according to the set interceptor to form the corresponding valid original data. The log data pre-processed by Flume will be sent to the Kafka message queue for subsequent processing. Kafka provides persistent storage and efficient message delivery mechanism, which can handle large-scale data streams. The pre-processing can at least include data formatting and data processing, etc.
[0075] Step 200: Read the data in Kafka for data cleaning and conversion, and use the complex event processing algorithm for correlation matching.
[0076] In step 200, FlinkCDC is used to read log data from Kafka and convert it into a data stream that can be processed by Flink. With Flink's stream processing capabilities, data streams can be cleaned, converted, and analyzed in real time. FlinkCEP is used to identify and process complex event patterns. In the fault plan recommendation system, FlinkCEP is used to identify event sequences that may cause faults. By defining appropriate event patterns and rules, the system can effectively match event sequences related to faults.
[0077] Step 300: The retrieved fault contingency plan is intuitively presented to the operation and maintenance personnel through a visual interface through active recommendation.
[0078] It is understandable that Doris is an open source distributed column storage system. With its high performance, high reliability and real-time data processing capabilities, it provides fast query, low latency and flexible SQL interface for large-scale data storage and multi-dimensional analysis, meeting the needs of data storage, analysis and mining. In the fault plan recommendation system, Doris is used to quickly retrieve fault plans related to matching events. The system recommends appropriate solutions based on fault characteristics and existing fault knowledge base. The fault plan recommendation system displays the found fault plans and related information to the operation and maintenance personnel through a visual interface. The interface design is intuitive and friendly, allowing operation and maintenance personnel to easily view and select fault plans. The operation and maintenance personnel can also interact with the system through the interface, such as selecting fault plans, updating fault handling progress, etc., and the system will respond in real time and provide corresponding feedback.
[0079] After the real-time alarm information is stored in the preset Doris database, the alarm data and analysis results can be uniformly accessed through restAPI and other collection methods to form a unified dispatch and unified operation and maintenance effect. It unifies the alarm rules and sending and scheduling settings of various monitoring performance indicator parameters through the platform, and supports multiple alarm notification methods such as SMS, email, and APP push. Operation and maintenance personnel only need to use one platform to visually view and manage all alarm information of the IT system. It provides a complete alarm monitoring system with a complete set of work processes from alarm access discovery, setting deployment, abnormal alarm to closed-loop summary, constructing monitoring, work order, and automated operation and maintenance support system, which can achieve the effect of synchronous follow-up of monitoring means, timely warning of abnormal situations, complete tracking of fault alarms, and rapid handling of faults.
[0080] As can be seen from the above description, the fault plan active recommendation method based on real-time log stream association matching provided by the embodiment of the present application collects the log data of each system through data collection and uses the interceptor to format and preprocess it, and stores it in the Kafka message queue. Through Flink technology, data cleaning, format conversion, and complex time processing are performed to identify the event sequence that may cause the fault. By defining appropriate event patterns and rules, the system can effectively match the event sequence related to the fault. Once the event sequence that may cause the fault is matched, the system immediately triggers the real-time search function and uses Doris to quickly retrieve the fault plan related to the matching event. The system actively recommends the fault plan found through a visual interface to the operation and maintenance personnel, so that they can better understand the cause and solution of the fault. The operation and maintenance personnel can interact with the system through the interface and track the progress of the fault handling in real time. Without manual screening and marking, it can effectively improve the automation, accuracy and reliability of the platform business system alarm monitoring, and can also improve the efficiency of the operation and maintenance personnel to operate and maintain the platform business system according to the alarm monitoring results, thereby improving the operation and maintenance efficiency and operation stability of the platform business system.
[0081] In order to further improve the automation, accuracy and reliability of alarm monitoring of the fault plan recommendation system, in a fault plan active recommendation method based on real-time log stream correlation matching provided in an embodiment of the present application, see Figure 4 The method 100 for proactively recommending a fault plan based on real-time log stream correlation matching may specifically include the following contents:
[0082] Step 110, use appropriate Flume interceptors to perform data preprocessing, such as formatting and cleaning, on the collected raw data. Custom interceptors can be written to handle specific log formats or data normalization requirements to ensure data consistency and accuracy.
[0083] Step 120, the log data pre-processed by Flume will be sent to the specified topic of the Kafka message queue for subsequent processing. Kafka provides persistent storage and efficient message delivery mechanism, which can handle large-scale data streams. It effectively caches large amounts of log data and provides accurate pre-processed data for real-time log stream data processing.
[0084] The steps 110 and 120 complete the preprocessing of the platform system log data, in order to reduce the subsequent task loss as much as possible and provide guarantee for data integrity and consistency.
[0085] In order to further improve the automation, accuracy and reliability of alarm monitoring of the fault plan recommendation system, in a fault plan active recommendation method based on real-time log stream correlation matching provided in an embodiment of the present application, see Figure 4 The method step 200 for proactively recommending a fault plan based on real-time log stream correlation matching may specifically include the following contents:
[0086] Step 210, use Flink's Kafka connector to read the pre-processed log data from the topic specified by Kafka. Through Flink's stream processing capabilities, the data stream can be cleaned, converted and analyzed in real time. Use the window operator to divide the data stream into windows, and perform analysis operations such as aggregation and statistics within the window.
[0087] Step 220, by defining appropriate event patterns and rules in the Flink job, the FlinkCEP library can be used to achieve pattern matching and sequence recognition of event streams. Event patterns define the rules and order of a series of events, which are used to describe the characteristics of fault event sequences. Rules can include event types, time constraints, conditional constraints, etc. Using FlinkCEP, the system can effectively match event sequences related to faults. The identified fault event sequence can be used as a trigger to further trigger the system's fault plan recommendation mechanism. The fault plan recommendation system can achieve real-time monitoring and analysis of real-time event streams, improving the efficiency of fault detection and resolution.
[0088] Steps 210 and 220 are a set of data processing fault event correlation matching procedures. By specifying complex event processing rules, the fault event sequence is effectively and quickly matched and identified, triggering the fault recommendation mechanism. At the same time, Flink's transactional writing and checkPoint mechanism are used to maximize the guarantee that no data will be lost during real-time data processing, thereby improving the efficiency and security of data processing.
[0089] In order to further improve the efficiency of the data processing process and the accuracy of the recommendation, in a fault plan active recommendation method based on real-time log stream correlation matching provided in an embodiment of the present application, see Figure 5 The method 220 for proactively recommending a fault plan based on real-time log stream correlation matching may specifically include the following contents:
[0090] Step 221, based on known fault scenarios and experience, a series of event patterns are defined through the FlinkCEP module. These patterns cover the sequence of events, time intervals, and constraints on other attributes. The defined event patterns can be matched in real-time data streams. Event sequences that meet specific conditions are captured to identify potential fault conditions. In order to better capture time correlations, time windows are introduced into the event pattern to ensure that only event sequences that match within a specific time range are considered fault events. When defining event patterns, various constraints are set, such as the numerical range of event attributes, the temporal relationship between events, etc. These constraints help improve the accuracy of fault events.
[0091] Step 222, FlinkCEP is a complex event processing library based on a stream processing framework. Its core function is to extract event sequences that meet specific conditions from a continuous event stream through pattern definition and matching algorithms. When the real-time event stream passes through the FlinkCEP module, the matching algorithm starts to execute. The algorithm compares the event stream with the predefined event pattern to find matches that appear in the event stream. The matching process takes into account factors such as the order, time interval, and attribute constraints of the event sequence.
[0092] Step 223, through the matching algorithm of the FlinkCEP module, we can identify complex events that appear in the real-time event stream and meet the predefined event pattern. These complex events may be composed of multiple events that occur in succession and meet specific time windows, timing relationships, and attribute constraints. When the FlinkCEP matching algorithm successfully finds an event sequence that matches the predefined event pattern, the system identifies these complex events as potential fault candidates. These candidate events may reflect a series of signs of abnormal operation of the system or equipment behind them. This identification helps to predict possible faults in advance. Once a complex event that may cause a fault is identified, the fault plan recommendation process is triggered to support the decision-making process of system administrators or operation and maintenance personnel when dealing with faults.
[0093] In order to further improve the automation, accuracy and reliability of the fault plan recommendation system, in a fault plan active recommendation method based on real-time log stream association matching provided in an embodiment of the present application, see Figure 4 The method step 300 for proactively recommending a fault plan based on real-time log stream correlation matching may specifically include the following contents:
[0094] Step 310, the fault plan recommendation system uses Doris, an efficient big data storage and analysis engine, to quickly retrieve fault plans related to the event, and intelligently recommend appropriate solutions based on the fault characteristics and the existing fault knowledge base. The system first performs rapid data retrieval and matching through Doris, and uses its distributed storage and indexing technology to achieve high-performance query response. The system compares and matches the existing fault knowledge base based on fault characteristics, such as event type, parameters, etc., to accurately find relevant fault plans. Based on the matching results, the system can intelligently recommend solutions, including fault handling guides, operating manuals, etc., to provide maintenance personnel or users with accurate and practical solutions. By utilizing Doris's rapid retrieval and matching capabilities, the fault plan recommendation system can improve fault handling efficiency, help quickly restore the normal operation of the system, and provide a better user experience.
[0095] In order to further improve the efficiency of the data processing process and the accuracy of the recommendation, in a fault plan active recommendation method based on real-time log stream correlation matching provided in an embodiment of the present application, see Figure 6 The method step 310 for proactively recommending a fault plan based on real-time log stream correlation matching may specifically include the following contents:
[0096] Step 311, make full use of the column storage characteristics of Doris to achieve efficient and scalable large-scale fault plan data storage and ensure fast query capabilities. During the upload of existing fault plan data, by interacting with the Doris system, the data is persisted in a column storage format to improve storage efficiency. The patent application emphasizes the optimization of data storage and retrieval to ensure data integrity. This innovation provides efficient and reliable data support for fault management and improves the efficiency and accuracy of fault plan processing.
[0097] Step 312, extract the features related to the fault by identifying the key information. Based on the extracted fault features, make full use of Doris's high-performance retrieval capabilities to perform query operations on fault plan data. In the data processing stage, by analyzing the important details in the log data, the fault features are extracted and stored in the Doris distributed column storage system. Subsequently, through specially designed query operations, the fault features are quickly accessed and analyzed to achieve efficient retrieval of fault plan data. This method gives full play to the advantages of the Doris storage system in high-performance retrieval and provides an efficient and accurate solution for fault management.
[0098] Step 313, by analyzing the query results and the extracted fault features, a specific matching algorithm is applied to match the fault plan. The matching algorithm is based on data mining technology, considers multi-dimensional features, and compares the correlation between real-time log data and the plan. Through intelligent matching, the system can quickly identify the plan applicable to the actual fault situation and provide an efficient emergency response plan. This innovation has a wide range of application value in the field of fault management and effectively improves the efficiency and accuracy of fault handling.
[0099] Step 314: After the fault plan is successfully matched, the system intelligently recommends an applicable solution based on the matching result. The recommendation is based on fault characteristics, historical data and best practices, providing operation and maintenance personnel with targeted emergency measures, maintenance processes, etc. The recommended solution takes actual conditions into consideration to ensure practicality and effectiveness of the operation. Patented innovation has significant value in efficiently solving fault problems and optimizing operation and maintenance processes, improving operation and maintenance efficiency and fault handling quality.
[0100] Step 320, the fault plan recommendation system displays the fault plan and related information to the operation and maintenance personnel through an intuitive visual interface, and provides easy viewing, selection and interaction functions. The operation and maintenance personnel can browse multiple fault plans through the interface, obtain key information, and select suitable solutions. They can also interact with the system in real time, such as selecting a plan, updating the processing progress, etc. The system will respond immediately and provide real-time feedback to improve the efficiency and accuracy of fault handling. The interface may also provide additional functions such as search, filtering and statistics to help operation and maintenance personnel quickly locate, analyze and evaluate fault conditions. Through the design and interactivity of the visual interface, the fault plan recommendation system provides operation and maintenance personnel with intuitive and convenient tools to meet their needs for viewing, selecting and operating fault plans, and promote the smooth progress of fault handling.
[0101] Steps 310 and 320 are a set of fault plan recommendation procedures, which use Doris to quickly retrieve and match fault plans, and display them to operation and maintenance personnel through an intuitive visual interface, providing easy viewing and selection functions, real-time interaction and feedback. The advantage is that by using Doris's high-performance query and distributed storage technology, the system can quickly and accurately recommend solutions, help improve fault handling efficiency, restore system normal operation, and provide a better user experience.
[0102] An embodiment of the present invention further provides a computer device, which may include a processor and a memory, wherein the processor and the memory may be connected via a bus or in other ways.
[0103] The processor may be a central processing unit (CPU). The processor may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, or a combination of the above chips.
[0104] The memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs, non-transitory computer executable programs and modules, such as the program instructions / modules corresponding to the key shielding method of the vehicle-mounted display device in the embodiment of the present invention. The processor executes various functional applications and data processing of the processor by running the non-transitory software programs, instructions and modules stored in the memory.
[0105] The memory may include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function; the data storage area may store data created by the processor, etc. In addition, the memory may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other non-volatile solid-state storage device. In some embodiments, the memory may optionally include a memory remotely arranged relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0106] The one or more modules are stored in the memory, and when executed by the processor, the method described in this embodiment is performed.
[0107] The embodiment of the present invention also provides a computer-readable storage medium on which a computer program is stored, and when the computer program is executed by a processor, the steps of the aforementioned edge computing server deployment method are implemented. The computer-readable storage medium can be a tangible storage medium, such as a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a floppy disk, a hard disk, a removable storage disk, a CD-ROM, or any other form of storage medium known in the technical field.
[0108] In summary, the fault plan recommendation method and device based on real-time log stream association matching described in the present invention collects the log data of the target system and stores it in the Kafka distributed streaming media platform for management, processes large-scale concurrent log data streams based on Flink, monitors the log data streams through a preset complex event processing algorithm to automatically identify fault event sequences, and utilizes the high-speed retrieval and matching capabilities of the Doris database to accurately query the fault plans corresponding to the fault event sequences, and through visual interaction, ensures that the operation and maintenance personnel can view, select and operate the fault plans, thereby achieving efficient and accurate fault detection and disposal.
[0109] Furthermore, based on the Flume component to collect log data, the FlinkCDC component is used to read the log data, and the FlinkCEP component is used to analyze the fault time series, thus achieving efficient and automatic fault detection for large-scale data.
[0110] Furthermore, by setting time windows, value range constraints, time interval constraints between events and time sequence constraints for different fault event sequence detection, more complex and systematic fault detection capabilities can be achieved.
[0111] It should be understood by those skilled in the art that the exemplary components, systems and methods described in conjunction with the embodiments disclosed herein can be implemented in hardware, software or a combination of the two. Whether it is performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention. When implemented in hardware, it can be, for example, an electronic circuit, an application specific integrated circuit (ASIC), appropriate firmware, a plug-in, a function card, etc. When implemented in software, the elements of the present invention are programs or code segments used to perform the required tasks. The program or code segment can be stored in a machine-readable medium, or transmitted on a transmission medium or a communication link via a data signal carried in a carrier.
[0112] It should be clear that the present invention is not limited to the specific configuration and processing described above and shown in the figures. For the sake of simplicity, a detailed description of the known method is omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present invention is not limited to the specific steps described and shown, and those skilled in the art can make various changes, modifications and additions, or change the order between the steps after understanding the spirit of the present invention.
[0113] In the present invention, features described and / or illustrated for one embodiment may be used in the same or similar manner in one or more other embodiments, and / or combined with features of other embodiments or replace features of other embodiments.
[0114] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the embodiments of the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A fault plan recommendation method based on real-time log stream correlation matching, It is characterized in that The method comprises the following steps: Collect log data of the target system from one or more data sources, format and pre-process it through the data interceptor, and then store it in the Kafka message queue; Read the log data in the Kafka message queue and convert it into a log data stream in a first set format for Flink processing and calling; Using a preset complex event processing algorithm to monitor the log data stream and identify a preset fault event sequence in the log data stream; Using the Doris database to retrieve a fault plan that matches the fault event sequence; The fault event sequence and the fault contingency plan are visually presented, and the handling process of the fault event sequence is tracked and recorded.
2. According to claim 1, the fault plan recommendation method based on real-time log stream correlation matching, It is characterized in that Collecting log data of a target system in one or more data sources includes: using a Flume component to collect the log data of the target system in the data source.
3. According to claim 1, the fault plan recommendation method based on real-time log stream correlation matching, It is characterized in that The method further includes: using a FlinkCDC component to read the log data in the Kafka message queue, and the preset complex event processing algorithm is a FlinkCEP component.
4. According to claim 1, the fault plan recommendation method based on real-time log stream correlation matching, It is characterized in that After the log data of the target system in one or more data sources are collected, the method further includes: using a data interceptor to perform format normalization preprocessing on the log data.
5. According to claim 1, the fault plan recommendation method based on real-time log stream correlation matching, It is characterized in that Identifying a preset fault event sequence in the log data stream, including: Setting corresponding time windows for multiple types of fault event sequences to identify corresponding fault event sequences within the time windows; And / or, setting a value range constraint, a time interval constraint between events, and a time sequence constraint to detect and identify the fault event sequence.
6. According to claim 1, the fault plan recommendation method based on real-time log stream correlation matching, It is characterized in that The fault prevention plan includes the handling process and tracking requirements for the fault event sequence.
7. The fault plan recommendation method based on real-time log stream correlation matching according to claim 1, It is characterized in that The Doris database is used to retrieve the fault contingency plan matching the fault event sequence. The Doris database selects and outputs the optimal fault contingency plan corresponding to the fault event sequence based on fault characteristics, historical matching results and practical feedback, including: The Doris database selects a plurality of candidate fault plans according to the fault characteristics of the fault event sequence, wherein the fault characteristics include numerical parameters of a plurality of attributes; The Doris database queries the historical matching results, and uses the number of times each candidate fault plan is selected for the fault event sequence as a first score; The Doris database queries the practice feedback of each candidate fault plan, and calculates a second score based on the practice feedback result; Taking a weighted sum of the first score and the second score corresponding to each candidate fault contingency plan to obtain a third score; The candidate fault pre-plan with the third highest score is used as the optimal fault pre-plan for the fault event sequence.
8. According to claim 1, the fault plan recommendation method based on real-time log stream correlation matching, It is characterized in that The method further comprises: The fault contingency plan is viewed, selected, interacted with and updated through visual presentation.
9. A fault plan recommendation device based on real-time log stream association matching, comprising a processor and a memory, It is characterized in that The memory stores computer instructions, and the processor is used to execute the computer instructions stored in the memory. When the computer instructions are executed by the processor, the device implements the steps of the method as claimed in any one of claims 1 to 8.
10. A computer-readable storage medium having a computer program stored thereon, It is characterized in that When the program is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.
Citation Information
Cited By
Equipment data storage Dores method based on industrial internet
CN120386780A