Log processing system and method based on Kafka and Flink

By adopting the combination solution of Kafka and Flink in the log processing system, the consistency and accuracy problems in log data processing in the prior art are solved, and efficient and accurate log data processing is achieved.

CN119988138APending Publication Date: 2025-05-13SHANDONG LANGCHAO YUNTOU INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510092706.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The existing log processing system has problems such as data loss, duplicate processing and inaccurate processing results when processing massive log data, and lacks sufficient data consistency and fault tolerance guarantees.

Method used

The log processing system based on Kafka and Flink is adopted to process message middleware through the Kafka cluster layer, and the stream processing capabilities of the Flink computing layer are used for real-time calculation and analysis to ensure data consistency and accuracy.

Benefits of technology

It realizes the consistency and accuracy of log data processing in a distributed environment, avoids the problems of data loss or repeated processing, and improves the stability and reliability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119988138A_ABST
    Figure CN119988138A_ABST
Patent Text Reader

Abstract

The invention discloses a log processing system and method based on Kafka and Flink, and relates to the technical field of real-time stream processing, the system adopts a distributed architecture and comprises a data source layer, a Kafka cluster layer, an Flink calculation layer and an analysis result storage layer, carrying out standardization processing on the log data of different data sources; the Kafka cluster layer serves as message middleware and is responsible for receiving log data from the data source layer and caching and distributing the log data; the Flink calculation layer is responsible for reading log data from a Kafka cluster through a Kafka connector, and performing real-time calculation and analysis by utilizing the stream processing capability of the Flink calculation layer; and the result storage layer is responsible for storing the data processed by the Flink calculation layer into a corresponding storage system so as to carry out subsequent data analysis and mining. According to the method, a Kafka + Flink combined mechanism is utilized, so that the consistency and accuracy of log data processing in a distributed environment are ensured, and the problem of data loss or repeated processing is avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of real-time stream processing, and specifically to a log processing system and method based on Kafka and Flink. Background Art

[0002] With the advent of the information age, enterprises and organizations generate massive amounts of log data in the process of running their application systems. These log data not only record the system's operating status, user behavior and other information, but also contain potential attack clues and performance bottlenecks. Therefore, efficient and accurate processing of log data is crucial to ensuring the stable operation of the system, improving the system's reliability and optimizing system performance.

[0003] However, traditional log processing systems generally face many challenges. First, due to the huge amount of log data, the traditional batch processing mode can no longer meet the real-time requirements. Secondly, in distributed systems, due to network delays, node failures and other reasons, data loss and duplicate processing have become common problems. In addition, existing log processing systems often lack sufficient data consistency and fault tolerance guarantees, resulting in the inability to guarantee the accuracy of processing results. Summary of the invention

[0004] In view of the problems of data loss, repeated processing and inaccurate processing results in the existing log processing system, the present invention provides a log processing system and method based on Kafka and Flink to realize efficient and accurate processing of massive log data.

[0005] In the first aspect, the present invention provides a log processing system based on Kafka and Flink. The technical solution adopted to solve the above technical problems is as follows:

[0006] A log processing system based on Kafka and Flink, which adopts a distributed architecture and includes a data source layer, a Kafka cluster layer, a Flink computing layer, and an analysis result storage layer, wherein:

[0007] The data source layer is responsible for collecting log data from various business systems in real time and standardizing log data from different data sources;

[0008] As a message middleware, the Kafka cluster layer is responsible for receiving log data from the data source layer and caching and distributing it.

[0009] The Flink computing layer is responsible for reading log data from the Kafka cluster through the Kafka Connector and using its stream processing capabilities for real-time computing and analysis.

[0010] The result storage layer is responsible for storing the data processed by the Flink computing layer into the corresponding storage system for subsequent data analysis and mining.

[0011] Optionally, the data source layer involved is configured with data collection tools and log collection strategies;

[0012] The data source layer involved is based on the log collection strategy and uses data collection tools to collect log data from various business systems in real time. The collected log data is standardized and pushed to the Kafka cluster in real time.

[0013] Optionally, before the data source layer sends the standardized log data to the Kafka cluster layer, set the acks parameter of the Kafka cluster to 1. At the same time, the Kafka cluster uses Kafka 0.11.0.0 and later versions, and uses idempotence and transaction mechanisms to solve possible duplicate sending problems.

[0014] Optionally, pre-configure the scale, partitioning strategy, number of replicas, and persistent storage of the Kafka cluster to meet the processing requirements of real-time data. The specific configuration contents are as follows:

[0015] (a1) Kafka cluster size: Configure the number of nodes in the Kafka cluster based on the estimated business needs and data volume;

[0016] (a2) Partition strategy of Kafka cluster: Set the number of topic partitions of Kafka cluster according to the amount of log data;

[0017] (a3) Number of replicas of the Kafka cluster: Set the number of replicas based on the estimated business needs and data volume;

[0018] (a4) Persistent storage of Kafka cluster: Select high-performance and reliable storage media and customize storage strategies to achieve persistent storage of log data.

[0019] Optionally, before the Flink computing layer reads log data from the Kafka cluster through the Kafka Connector, you need to configure the following parameters:

[0020] (b1) Connection information of the Kafka cluster: including the address and port number of the Kafka cluster. This connection information is used to establish a connection with the Kafka cluster.

[0021] (b2) Topic name: specifies the Kafka topic to be read or written. A Kafka topic is a logical grouping of messages in a Kafka cluster. Each Kafka topic contains at least one partition.

[0022] (b3) Partition strategy: defines how to distribute log data to different partitions;

[0023] Before the Flink computing layer uses its stream processing capabilities to perform real-time computing and analysis on log data, the following contents are pre-configured:

[0024] (c1) Job management and scheduling at the Flink computing layer: Configure the JobManager and TaskManager at the Flink computing layer, customize the parallelism and resource allocation strategies, and customize the job scheduling strategies;

[0025] (c2) State management of the Flink computing layer: Select a high-performance, reliable state backend and configure a checkpoint strategy;

[0026] (c3) Data processing logic of the Flink computing layer: Consider the characteristics of the data and business requirements, select appropriate algorithms and data structures, and write the Flink job data processing logic;

[0027] (c4) Exception handling and fault tolerance mechanism of Flink computing layer: Configure retry mechanism and backup mechanism in Flink job to ensure timely processing and recovery of job status when failure occurs.

[0028] In the second aspect, the present invention provides a log processing method based on Kafka and Flink. The technical solution adopted to solve the above technical problems is as follows:

[0029] A log processing method based on Kafka and Flink includes the following steps:

[0030] S1. Collect log data from various business systems in real time and standardize log data from different data sources;

[0031] S2, receiving the standardized log data through the Kafka cluster, caching and distributing it;

[0032] S3 and Flink read log data from the Kafka cluster through the Kafka Connector and use its stream processing capabilities for real-time computing and analysis;

[0033] S4: Store the data after Flink's real-time calculation and analysis in the corresponding storage system for subsequent data analysis and mining.

[0034] Optionally, execute step S1, pre-select a data collection tool, and configure a log collection strategy;

[0035] Based on the log collection strategy, data collection tools are used to collect log data from various business systems in real time, and the collected log data is standardized and pushed to the Kafka cluster in real time.

[0036] Optionally, before the involved Kafka cluster receives log data, set the acks parameter of the Kafka cluster to 1. At the same time, the Kafka cluster uses Kafka 0.11.0.0 and later versions, and uses idempotence and transaction mechanisms to solve possible duplicate sending problems.

[0037] Optionally, before executing step S2, pre-configure the scale, partitioning strategy, number of replicas, and persistent storage of the Kafka cluster, as follows:

[0038] (a1) Kafka cluster size: Configure the number of nodes in the Kafka cluster based on the estimated business needs and data volume;

[0039] (a2) Partition strategy of Kafka cluster: Set the number of topic partitions of Kafka cluster according to the amount of log data;

[0040] (a3) Number of replicas of the Kafka cluster: Set the number of replicas based on the estimated business needs and data volume;

[0041] (a4) Persistent storage of Kafka cluster: Select high-performance and reliable storage media and customize storage strategies to achieve persistent storage of log data.

[0042] Optionally, before Flink reads log data from the Kafka cluster through the Kafka Connector, you need to configure the following parameters:

[0043] (b1) Connection information of the Kafka cluster: including the address and port number of the Kafka cluster. This connection information is used to establish a connection with the Kafka cluster.

[0044] (b2) Topic name: specifies the Kafka topic to be read or written. A Kafka topic is a logical grouping of messages in a Kafka cluster. Each Kafka topic contains at least one partition.

[0045] (b3) Partition strategy: defines how to distribute log data to different partitions;

[0046] Before Flink uses its stream processing capabilities to perform real-time calculations and analysis on log data, pre-configure the following:

[0047] (c1) Job management and scheduling at the Flink computing layer: Configure the JobManager and TaskManager at the Flink computing layer, customize the parallelism and resource allocation strategies, and customize the job scheduling strategies;

[0048] (c2) State management of the Flink computing layer: Select a high-performance, reliable state backend and configure a checkpoint strategy;

[0049] (c3) Data processing logic of the Flink computing layer: Consider the characteristics of the data and business requirements, select appropriate algorithms and data structures, and write the Flink job data processing logic;

[0050] (c4) Exception handling and fault tolerance mechanism of Flink computing layer: Configure retry mechanism and backup mechanism in Flink job to ensure timely processing and recovery of job status when failure occurs.

[0051] The log processing system and method based on Kafka and Flink of the present invention have the following beneficial effects compared with the prior art:

[0052] 1. The present invention uses the combined mechanism of Kafka+Flink to ensure the consistency and accuracy of log data processing in a distributed environment, avoiding the problem of data loss or repeated processing;

[0053] 2. The present invention solves the problems of data loss, duplicate processing and inaccurate processing results in existing log processing systems by integrating Kafka's high throughput and low latency characteristics with Flink's exactly-once processing semantics. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] Attached Figure 1 It is a system architecture diagram of the present invention. DETAILED DESCRIPTION

[0055] In order to make the technical solution, the technical problem solved and the technical effect of the present invention more clearly understood, the technical solution of the present invention is clearly and completely described below in conjunction with specific embodiments.

[0056] Embodiment 1:

[0057] Combined with Figure 1 This embodiment proposes a log processing system based on Kafka and Flink. The system adopts a distributed architecture, including a data source layer, a Kafka cluster layer, a Flink computing layer, and an analysis result storage layer, wherein:

[0058] The data source layer is responsible for collecting log data from various business systems in real time and standardizing log data from different data sources;

[0059] As a message middleware, the Kafka cluster layer is responsible for receiving log data from the data source layer and caching and distributing it.

[0060] The Flink computing layer is responsible for reading log data from the Kafka cluster through the Kafka Connector and using its stream processing capabilities for real-time computing and analysis.

[0061] The result storage layer is responsible for storing the data processed by the Flink computing layer into the corresponding storage system for subsequent data analysis and mining.

[0062] In this embodiment, the data source layer involved is configured with a data collection tool and a log collection strategy. The data collection tool can be Flume or Logstash. Based on the log collection strategy, the data source layer uses the data collection tool to collect the log data of each business system in real time, and performs standardized processing on the collected log data, such as unifying the log format and parsing the fields, and then pushes it to the Kafka cluster in real time.

[0063] In this embodiment, before the data source layer sends the standardized log data to the Kafka cluster layer, the acks parameter of the Kafka cluster is set to 1. At the same time, the Kafka cluster uses Kafka 0.11.0.0 and later versions, and uses idempotence and transaction mechanisms to solve possible duplicate sending problems.

[0064] It should be added that the acks parameter of the Kafka cluster has the following three optional values:

[0065] 1 (default value): indicates that the Leader node needs to reply to the received message, so that the producer will send the next data;

[0066] -1 (i.e. all): means that all Leader + replica nodes need to reply to the received message (acks = -1), so that the producer will send the next data;

[0067] 0: Indicates that no node response is required and the producer will continue to send the next piece of data.

[0068] When the acks parameter of the Kafka cluster is set to 1, there is an unexpected situation when sending messages to the Kafka cluster: the message is successfully written, but at this time the producer does not receive a response to the successful writing due to network problems. The producer will start the retry operation until the network is restored, and the message will be sent multiple times. This becomes at least once. Therefore, when deploying Kafka, you need to select a version later than Kafka 0.11.0.0. After Kafka 0.11.0.0, the Kafka cluster implements exactly once semantics through two mechanisms, idempotence and transaction.

[0069] In this embodiment, in order to meet the real-time data processing requirements, it is necessary to pre-configure the scale, partition strategy, number of replicas, and persistent storage of the Kafka cluster. The specific configuration contents are as follows:

[0070] (a1) Kafka cluster size: Configure the number of nodes in the Kafka cluster based on the estimated business needs and data volume to ensure system stability and scalability;

[0071] (a2) Partition strategy of Kafka cluster: According to the amount of log data, set the number of topic partitions of Kafka cluster to improve the parallel processing capability and load balancing of data;

[0072] (a3) Number of replicas of the Kafka cluster: Set an appropriate number of replicas based on the estimated business needs and data volume to improve data reliability and fault tolerance and prevent data loss;

[0073] (a4) Persistent storage of Kafka cluster: Select high-performance and reliable storage media and configure reasonable storage strategies, such as RAID and distributed storage, to achieve persistent storage of log data and ensure data reliability and recoverability.

[0074] In this embodiment, before the Flink computing layer reads log data from the Kafka cluster through the Kafka Connector, the following parameters need to be configured:

[0075] (b1) Connection information of the Kafka cluster: including the address and port number of the Kafka cluster. This connection information is used to establish a connection with the Kafka cluster.

[0076] (b2) Topic name: specifies the Kafka topic to be read or written. A Kafka topic is a logical grouping of messages in a Kafka cluster. Each Kafka topic contains at least one partition.

[0077] (b3) Partition strategy: defines how to distribute log data to different partitions.

[0078] Before the Flink computing layer uses its stream processing capabilities to perform real-time computing and analysis on log data, you need to pre-configure the following:

[0079] (c1) Job management and scheduling at the Flink computing layer: Configure the JobManager and TaskManager of the Flink computing layer, set reasonable parallelism and resource allocation strategies, and set reasonable job scheduling strategies to ensure stable operation and efficient processing of jobs;

[0080] (c2) State management of the Flink computing layer: Select a high-performance and reliable state backend, such as RocksDB, HDFS, etc., and configure a reasonable checkpoint strategy to ensure that the job status can be restored in the event of a failure. In order to ensure that the entire process of log analysis is done exactly once, this embodiment configures the checkpoint as "exactly-once";

[0081] (c3) Data processing logic of the Flink computing layer: Write efficient Flink job data processing logic based on business needs. When writing the processing logic, you need to fully consider the characteristics of the data and business needs, and use appropriate algorithms and data structures to improve processing efficiency and accuracy;

[0082] (c4) Exception handling and fault tolerance mechanism of Flink computing layer: Configure exception handling and fault tolerance mechanisms in Flink jobs, such as retry mechanism and backup mechanism, to ensure timely processing and recovery of job status when failure occurs.

[0083] Embodiment 2:

[0084] Reference Figure 1 This embodiment proposes a log processing method based on Kafka and Flink, which includes the following steps:

[0085] S1. Collect log data from various business systems in real time and standardize log data from different data sources.

[0086] Before executing step S1, pre-select a data collection tool, such as Flume or Logstash, and configure a log collection policy;

[0087] Based on the log collection strategy, data collection tools are used to collect log data from various business systems in real time, and the collected log data is standardized such as unifying the log format and parsing fields, and then pushed to the Kafka cluster in real time.

[0088] S2: Receive the standardized log data through the Kafka cluster, and cache and distribute it.

[0089] Before the Kafka cluster receives log data, set the acks parameter of the Kafka cluster to 1. At the same time, the Kafka cluster uses Kafka 0.11.0.0 and later versions, and uses idempotence and transaction mechanisms to solve possible duplicate sending problems.

[0090] It should be added that the acks parameter of the Kafka cluster has the following three optional values:

[0091] 1 (default value): indicates that the Leader node needs to reply to the received message, so that the producer will send the next data;

[0092] -1 (i.e. all): means that all Leader + replica nodes need to reply to the received message (acks = -1), so that the producer will send the next data;

[0093] 0: Indicates that no node response is required and the producer will continue to send the next piece of data.

[0094] When the acks parameter of the Kafka cluster is set to 1, there is an unexpected situation when sending messages to the Kafka cluster: the message is successfully written, but at this time the producer does not receive a response to the successful writing due to network problems. The producer will start the retry operation until the network is restored, and the message will be sent multiple times. This becomes at least once. Therefore, when deploying Kafka, you need to select a version later than Kafka 0.11.0.0. After Kafka 0.11.0.0, the Kafka cluster implements exactly once semantics through two mechanisms, idempotence and transaction.

[0095] In order to meet the real-time data processing requirements, it is necessary to pre-configure the scale, partitioning strategy, number of replicas, and persistent storage of the Kafka cluster. The specific configuration contents are as follows:

[0096] (a1) Kafka cluster size: Configure the number of nodes in the Kafka cluster based on the estimated business needs and data volume to ensure system stability and scalability;

[0097] (a2) Partition strategy of Kafka cluster: According to the amount of log data, set the number of topic partitions of Kafka cluster to improve the parallel processing capability and load balancing of data;

[0098] (a3) Number of replicas of the Kafka cluster: Set an appropriate number of replicas based on the estimated business needs and data volume to improve data reliability and fault tolerance and prevent data loss;

[0099] (a4) Persistent storage of Kafka cluster: Select high-performance and reliable storage media and configure reasonable storage strategies, such as RAID and distributed storage, to achieve persistent storage of log data and ensure data reliability and recoverability.

[0100] S3 and Flink read log data from the Kafka cluster through the Kafka Connector and use its stream processing capabilities for real-time computing and analysis.

[0101] S3.1. Before Flink reads log data from the Kafka cluster through the Kafka Connector, you need to configure the following parameters:

[0102] (b1) Connection information of the Kafka cluster: including the address and port number of the Kafka cluster. This connection information is used to establish a connection with the Kafka cluster.

[0103] (b2) Topic name: specifies the Kafka topic to be read or written. A Kafka topic is a logical grouping of messages in a Kafka cluster. Each Kafka topic contains at least one partition.

[0104] (b3) Partition strategy: defines how to distribute log data to different partitions;

[0105] S3.2 Before the computing layer uses its stream processing capabilities to perform real-time computing and analysis on log data, it is necessary to pre-configure the following:

[0106] (c1) Job management and scheduling at the Flink computing layer: Configure the JobManager and TaskManager of the Flink computing layer, set reasonable parallelism and resource allocation strategies, and set reasonable job scheduling strategies to ensure stable operation and efficient processing of jobs;

[0107] (c2) State management of the Flink computing layer: Select a high-performance and reliable state backend, such as RocksDB, HDFS, etc., and configure a reasonable checkpoint strategy to ensure that the job status can be restored in the event of a failure. In order to ensure that the entire process of log analysis is done exactly once, this embodiment configures the checkpoint as "exactly-once";

[0108] (c3) Data processing logic of the Flink computing layer: Write efficient Flink job data processing logic based on business needs. When writing the processing logic, you need to fully consider the characteristics of the data and business needs, and use appropriate algorithms and data structures to improve processing efficiency and accuracy;

[0109] (c4) Exception handling and fault tolerance mechanism of Flink computing layer: Configure exception handling and fault tolerance mechanisms in Flink jobs, such as retry mechanism and backup mechanism, to ensure timely processing and recovery of job status when failure occurs.

[0110] S4: Store the data after Flink's real-time calculation and analysis in the corresponding storage system for subsequent data analysis and mining.

[0111] In summary, the log processing system and method based on Kafka and Flink of the present invention solves the problems of data loss, repeated processing, and inaccurate processing results existing in the existing log processing system by integrating the high throughput and low latency characteristics of Kafka with the exactly-once processing semantics of Flink.

[0112] The above specific examples are used to explain the principles and implementation methods of the present invention in detail. These examples are only used to help understand the core technical content of the present invention. Based on the above specific embodiments of the present invention, any improvements and modifications made to the present invention by technicians in this technical field without departing from the principles of the present invention should fall within the scope of patent protection of the present invention.

Claims

1. A log processing system based on Kafka and Flink, characterized in that: The system adopts a distributed architecture, including the data source layer, Kafka cluster layer, Flink computing layer, and analysis result storage layer, where: The data source layer is responsible for collecting log data from various business systems in real time and standardizing log data from different data sources; As a message middleware, the Kafka cluster layer is responsible for receiving log data from the data source layer and caching and distributing it. The Flink computing layer is responsible for reading log data from the Kafka cluster through the Kafka Connector and using its stream processing capabilities for real-time computing and analysis. The result storage layer is responsible for storing the data processed by the Flink computing layer into the corresponding storage system for subsequent data analysis and mining.

2. According to claim 1, a log processing system based on Kafka and Flink is characterized in that: The data source layer is configured with data collection tools and log collection strategies; The data source layer collects log data from various business systems in real time based on the log collection strategy using data collection tools, and pushes the collected log data to the Kafka cluster in real time after standardization.

3. According to the log processing system based on Kafka and Flink according to claim 1, it is characterized in that: Before the data source layer sends the standardized log data to the Kafka cluster layer, the acks parameter of the Kafka cluster is set to 1. At the same time, the Kafka cluster uses Kafka 0.11.0.0 and later versions, and uses idempotence and transaction mechanisms to solve possible duplicate sending problems.

4. According to the log processing system based on Kafka and Flink according to claim 1, it is characterized in that: Pre-configure the scale, partitioning strategy, number of replicas, and persistent storage of the Kafka cluster to meet the processing requirements of real-time data. The specific configuration contents are as follows: (a1) Kafka cluster size: Configure the number of nodes in the Kafka cluster based on the estimated business needs and data volume; (a2) Partition strategy of Kafka cluster: Set the number of topic partitions of Kafka cluster according to the amount of log data; (a3) Number of replicas of the Kafka cluster: Set the number of replicas based on the estimated business needs and data volume; (a4) Persistent storage of Kafka cluster: Select high-performance and reliable storage media and customize storage strategies to achieve persistent storage of log data.

5. According to the log processing system based on Kafka and Flink according to claim 1, it is characterized in that: Before the Flink computing layer reads log data from the Kafka cluster through the Kafka Connector, you need to configure the following parameters: (b1) Connection information of the Kafka cluster: including the address and port number of the Kafka cluster. This connection information is used to establish a connection with the Kafka cluster. (b2) Topic name: specifies the Kafka topic to be read or written. A Kafka topic is a logical grouping of messages in a Kafka cluster. Each Kafka topic contains at least one partition. (b3) Partition strategy: defines how to distribute log data to different partitions; Before the Flink computing layer uses its stream processing capabilities to perform real-time computing and analysis on log data, the following contents are pre-configured: (c1) Job management and scheduling at the Flink computing layer: Configure the JobManager and TaskManager at the Flink computing layer, customize the parallelism and resource allocation strategies, and customize the job scheduling strategies; (c2) State management of the Flink computing layer: Select a high-performance, reliable state backend and configure a checkpoint strategy; (c3) Data processing logic of the Flink computing layer: Consider the characteristics of the data and business requirements, select appropriate algorithms and data structures, and write the Flink job data processing logic; (c4) Exception handling and fault tolerance mechanism of Flink computing layer: Configure retry mechanism and backup mechanism in Flink job to ensure timely processing and recovery of job status when failure occurs.

6. A log processing method based on Kafka and Flink, characterized in that: It includes the following steps: S1. Collect log data from various business systems in real time and standardize log data from different data sources; S2, receiving the standardized log data through the Kafka cluster, caching and distributing it; S3 and Flink read log data from the Kafka cluster through the Kafka Connector and use its stream processing capabilities for real-time computing and analysis; S4: Store the data after Flink's real-time calculation and analysis in the corresponding storage system for subsequent data analysis and mining.

7. The log processing method based on Kafka and Flink according to claim 6 is characterized in that: Execute step S1, pre-select a data collection tool, and configure a log collection strategy; Based on the log collection strategy, data collection tools are used to collect log data from various business systems in real time, and the collected log data is standardized and pushed to the Kafka cluster in real time.

8. The log processing method based on Kafka and Flink according to claim 6, characterized in that: Before the Kafka cluster receives log data, set the acks parameter of the Kafka cluster to 1. At the same time, the Kafka cluster uses Kafka0.11.0.0 and later versions, and uses idempotence and transaction mechanisms to solve possible duplicate sending problems.

9. The log processing method based on Kafka and Flink according to claim 6, characterized in that: Before executing step S2, pre-configure the scale, partitioning strategy, number of replicas, and persistent storage of the Kafka cluster as follows: (a1) Kafka cluster size: Configure the number of nodes in the Kafka cluster based on the estimated business needs and data volume; (a2) Partition strategy of Kafka cluster: Set the number of topic partitions of Kafka cluster according to the amount of log data; (a3) Number of replicas of the Kafka cluster: Set the number of replicas based on the estimated business needs and data volume; (a4) Persistent storage of Kafka cluster: Select high-performance and reliable storage media and customize storage strategies to achieve persistent storage of log data.

10. The log processing method based on Kafka and Flink according to claim 6, characterized in that: Before Flink reads log data from the Kafka cluster through the Kafka Connector, you need to configure the following parameters: (b1) Connection information of the Kafka cluster: including the address and port number of the Kafka cluster. This connection information is used to establish a connection with the Kafka cluster. (b2) Topic name: specifies the Kafka topic to be read or written. A Kafka topic is a logical grouping of messages in a Kafka cluster. Each Kafka topic contains at least one partition. (b3) Partition strategy: defines how to distribute log data to different partitions; Before Flink uses its stream processing capabilities to perform real-time calculations and analysis on log data, pre-configure the following: (c1) Job management and scheduling at the Flink computing layer: Configure the JobManager and TaskManager at the Flink computing layer, customize the parallelism and resource allocation strategies, and customize the job scheduling strategies; (c2) State management of the Flink computing layer: Select a high-performance, reliable state backend and configure a checkpoint strategy; (c3) Data processing logic of the Flink computing layer: Consider the characteristics of the data and business requirements, select appropriate algorithms and data structures, and write the Flink job data processing logic; (c4) Exception handling and fault tolerance mechanism of Flink computing layer: Configure retry mechanism and backup mechanism in Flink job to ensure timely processing and recovery of job status when failure occurs.