Data processing system and method based on Kafka and Elasticsearch, terminal and storage medium

By combining Kafka and Elasticsearch to build a data processing system, the traditional data processing methods are solved, and the problems of slow speed, high latency and poor reliability in massive data processing are achieved, and efficient and reliable data processing is achieved.

CN120540778APending Publication Date: 2025-08-26SHENZHEN COOCAA NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510624242.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-15
Publication Date
2025-08-26

AI Technical Summary

Technical Problem

When traditional data processing methods face massive data, there are problems such as slow processing speed, high latency and poor reliability.

Method used

Combining Kafka and Elasticsearch, a data processing system is built, including the data reporting module, the Kafka message queue module and the Elasticsearch consumption module. Through Kafka's high throughput and distributed processing capabilities, as well as the efficient data storage and retrieval capabilities of Elasticsearch, efficient data processing is achieved.

Benefits of technology

It realizes efficient data processing, enhances data reliability and security, and meets current data processing needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120540778A_ABST
    Figure CN120540778A_ABST
Patent Text Reader

Abstract

The invention discloses a data processing system and method based on Kafka and Elasticsearch, a terminal and a storage medium, and the system comprises a data reporting module which is used for reporting data from a plurality of data sources to a Kafka message queue; the Kafka message queue module is used for receiving the data reported by the data reporting module and classifying and storing the data in a Kafka cluster in a Topic form; and the Elasticsearch consumption module is used for reading the data from the Kafka message queue module, processing the read data and then indexing the processed data into Elasticsearch so as to carry out data analysis and query. According to the method, efficient data processing is realized through the high throughput and the distributed processing capability of Kafka and the efficient data storage and retrieval capability of Elasticsearch.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing technology, and in particular to a data processing system, method, terminal and computer-readable storage medium based on Kafka and Elasticsearch. Background Art

[0002] With the continuous development of big data technology, the demand for data processing is increasing. For example, batch processing involves processing data in batches on a fixed cycle (e.g., daily / weekly). This is typically used for offline analysis and relies on structured data storage (e.g., tables). A strict schema must be defined before processing. This approach is suitable for scenarios requiring high consistency and low timeliness (e.g., generating financial statements). However, it suffers from high latency, inability to respond in real time, and poor scalability. For example, processing based on relational databases uses ACID transactions to ensure data consistency, supports complex queries, and relies on vertical scaling (upgrading single-machine hardware) to improve performance. However, performance degrades significantly as data volume increases (e.g., joining tables with billions of data points). This approach makes it difficult to process unstructured data (e.g., text and images).

[0003] In other words, when faced with massive amounts of data, traditional data processing methods have problems such as slow processing speed, high latency, and poor reliability.

[0004] Therefore, the existing technology still needs to be improved and developed. Summary of the Invention

[0005] The main purpose of the present invention is to provide a data processing system, method, terminal and computer-readable storage medium based on Kafka and Elasticsearch, aiming to solve the problems of slow processing speed, high latency and poor reliability in traditional data processing methods when processing massive data in the existing technology.

[0006] To achieve the above objectives, the present invention provides a data processing system based on Kafka and Elasticsearch, wherein the data processing system based on Kafka and Elasticsearch comprises:

[0007] Data reporting module, used to report data from multiple data sources to the Kafka message queue;

[0008] A Kafka message queue module is used to receive the data reported by the data reporting module and store the data in a Kafka cluster in the form of topics;

[0009] The Elasticsearch consumption module is used to read data from the Kafka message queue module and index the read data into Elasticsearch after processing for data analysis and query.

[0010] Optionally, in the data processing system based on Kafka and Elasticsearch, the data reporting module includes:

[0011] A data acquisition unit, configured to acquire data from multiple data sources through the Kafka Producer SDK or a collection tool, including log data, event stream data, and user behavior data;

[0012] The data transmission unit is used to configure the data compression method and confirmation mechanism, and report the data obtained by the data acquisition unit to the Kafka message queue through the Kafka protocol.

[0013] Optionally, in the data processing system based on Kafka and Elasticsearch, the Kafka message queue module includes:

[0014] a data receiving unit, configured to receive the data reported by the data transmission unit;

[0015] The data storage unit is used to classify and store the data in the form of Topics in the Kafka cluster. The Kafka cluster is composed of multiple Brokers, each Broker represents a message middleware processing node, and each Topic contains one or more Partitions.

[0016] Optionally, in the data processing system based on Kafka and Elasticsearch, the Elasticsearch consumption module includes:

[0017] A data reading unit, used to subscribe to a specified Topic or Partition from the Kafka message queue module through the Kafka Consumer SDK and read data in batches;

[0018] a data processing unit, configured to process the batch data read by the data reading unit to obtain processed data;

[0019] The data indexing unit is used to index the processed data into Elasticsearch in batches for data analysis and query.

[0020] In addition, the present invention also provides a data processing method based on Kafka and Elasticsearch, which includes the following steps:

[0021] Report data from multiple data sources to the Kafka message queue;

[0022] The data is classified and stored in the Kafka cluster in the form of topics;

[0023] Read data from the Kafka cluster, process the data, and index it into Elasticsearch for data analysis and query.

[0024] Optionally, the data processing method based on Kafka and Elasticsearch, wherein the step of reporting data from multiple data sources to a Kafka message queue, specifically includes:

[0025] Use the Kafka Producer SDK or collection tools to obtain data from multiple data sources, including log data, event stream data, and user behavior data.

[0026] Configure the data compression method and confirmation mechanism, and report the acquired data to the Kafka message queue through the Kafka protocol.

[0027] Optionally, the data processing method based on Kafka and Elasticsearch, wherein the data is classified and stored in the Kafka cluster in the form of topics, specifically includes:

[0028] Receive the data reported to the Kafka message queue;

[0029] The data is classified and stored in the form of Topics in a Kafka cluster. The Kafka cluster consists of multiple Brokers, each Broker represents a message middleware processing node, and each Topic contains one or more Partitions.

[0030] Optionally, the data processing method based on Kafka and Elasticsearch, wherein the data is read from the Kafka cluster and indexed into Elasticsearch after processing for data analysis and query, specifically includes:

[0031] Use the Kafka Consumer SDK to subscribe to a specified Topic or Partition from the Kafka cluster and read data in batches.

[0032] Process the read batch data to obtain processed data;

[0033] The processed data is batch-indexed into Elasticsearch for data analysis and query.

[0034] In addition, to achieve the above-mentioned purpose, the present invention also provides a terminal, wherein the terminal includes: a memory, a processor, and a data processing program based on Kafka and Elasticsearch stored on the memory and runnable on the processor, and when the data processing program based on Kafka and Elasticsearch is executed by the processor, the steps of the data processing method based on Kafka and Elasticsearch as described above are implemented.

[0035] In addition, to achieve the above-mentioned purpose, the present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a data processing program based on Kafka and Elasticsearch, and when the data processing program based on Kafka and Elasticsearch is executed by the processor, the steps of the data processing method based on Kafka and Elasticsearch as described above are implemented.

[0036] In the present invention, the system includes: a data reporting module for reporting data from multiple data sources to a Kafka message queue; a Kafka message queue module for receiving the data reported by the data reporting module and storing the data in a Kafka cluster in the form of topics; and an Elasticsearch consumption module for reading data from the Kafka message queue module and indexing the read data into Elasticsearch after processing for data analysis and query. The present invention combines Kafka and Elasticsearch to construct an efficient data processing system. By leveraging Kafka's high throughput and distributed processing capabilities and Elasticsearch's efficient data storage and retrieval capabilities, it achieves efficient data processing, enhances data reliability and security, and can meet current data processing needs. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 This is a structural diagram of a preferred embodiment of the data processing system based on Kafka and Elasticsearch of the present invention;

[0038] Figure 2This is a schematic diagram of the system architecture in a preferred embodiment of the data processing system based on Kafka and Elasticsearch of the present invention;

[0039] Figure 3 This is a flow chart of a preferred embodiment of the data processing method based on Kafka and Elasticsearch of the present invention;

[0040] Figure 4 FIG. 4 is a structural diagram of a preferred embodiment of the terminal of the present invention. DETAILED DESCRIPTION

[0041] In order to make the purpose, technical solutions and advantages of the present invention more clear and distinct, the present invention is further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0042] The data processing system based on Kafka and Elasticsearch described in the preferred embodiment of the present invention is as follows: Figure 1 and Figure 2 As shown, the data processing system based on Kafka and Elasticsearch includes:

[0043] Data reporting module, used to report data from multiple data sources to the Kafka message queue;

[0044] A Kafka message queue module is used to receive the data reported by the data reporting module and store the data in a Kafka cluster in the form of topics;

[0045] The Elasticsearch consumption module is used to read data from the Kafka message queue module and index the read data into Elasticsearch after processing for data analysis and query.

[0046] Specifically, the data reporting module includes:

[0047] A data acquisition unit, configured to acquire data from multiple data sources through the Kafka Producer SDK or a collection tool, including log data, event stream data, and user behavior data;

[0048] The data transmission unit is used to configure the data compression method and confirmation mechanism, and report the data obtained by the data acquisition unit to the Kafka message queue through the Kafka protocol.

[0049] The data reporting module is responsible for reporting data from various data sources to the Kafka message queue. Data sources can be various logs (such as Nginx access logs, server error logs, microservice call chain logs), event streams, user behavior data (such as clickstreams, page dwell time, app button click events), etc. In other words, Kafka's data sources can be very diverse. Whether it is structured logs, semi-structured event streams, or unstructured user behavior data, they can all be efficiently collected, transmitted, and distributed through Kafka (Kafka is a high-throughput, distributed message queue based on the publish / subscribe model that supports partitioned, distributed real-time message processing and has a fault-tolerant guarantee mechanism). The data reporting module reports data to the Kafka cluster using the Kafka protocol (the Kafka protocol is the underlying rules for communication between clients (Producers / Consumers) and Brokers in the Kafka cluster, defining the message format, request / response model, and error handling mechanism. It is the core foundation of Kafka's high performance and scalability) through the Kafka Producer SDK or collection tool. During the reporting process, you can configure parameters such as data compression method and confirmation mechanism to improve the efficiency and reliability of data transmission.

[0050] The Kafka Producer SDK is a software development kit (SDK) for sending messages to Kafka clusters. It encapsulates the APIs, configuration, and underlying communication logic associated with Kafka producers, enabling developers to easily integrate Kafka messaging into their applications. The Kafka Producer SDK is a key component of real-time data pipelines. Its core value lies in: simplified development, encapsulating complex network communication and retry logic; high performance, supporting optimizations such as asynchrony, batching, and compression; and flexibility, enabling business needs to be adapted through custom serialization and partitioning. By selecting an SDK suitable for your language and configuring appropriate parameters, you can build a highly reliable, high-throughput message producer.

[0051] Specifically, the Kafka message queue module includes:

[0052] a data receiving unit, configured to receive the data reported by the data transmission unit;

[0053] The data storage unit is used to classify and store the data in the form of Topics in the Kafka cluster. The Kafka cluster is composed of multiple Brokers, each Broker represents a message middleware processing node, and each Topic contains one or more Partitions.

[0054] The Kafka message queue module is the core of the entire data processing system, responsible for receiving data sent by the data reporting module and storing it in the Kafka cluster. A Kafka cluster consists of multiple brokers, each of which is a message middleware processing node. (Each broker is an independent Kafka server (node) that works together to achieve high availability, high performance, and scalable data stream processing. Each broker is responsible for storing data for a partition and processing read and write requests from producers and consumers.) Data is stored in the Kafka cluster in the form of topics (a core abstraction in Kafka that distinguishes different services or types of data streams). Each topic can contain one or more partitions (each topic is divided into multiple partitions and distributed across different brokers), enhancing parallel data processing capabilities. Furthermore, the Kafka cluster supports data replication (Kafka's data replication mechanism is a core design for ensuring high data availability and fault tolerance. By replicating the data of each partition across multiple brokers, it ensures that data remains securely accessible even if some nodes fail), improving data reliability and fault tolerance.

[0055] The core functions of topics include:

[0056] Logical classification: Each Topic represents a type of data stream (for example: order_events, user_logs, sensor_data).

[0057] Physical storage: Topic data is actually stored on the distributed broker's disk and organized in the form of partitions.

[0058] Multi-subscription support: allows different consumer groups to independently consume data from the same topic.

[0059] Specifically, the Elasticsearch consumption module includes:

[0060] A data reading unit, used to subscribe to a specified Topic or Partition from the Kafka message queue module through the Kafka Consumer SDK and read data in batches;

[0061] a data processing unit, configured to process the batch data read by the data reading unit to obtain processed data;

[0062] The data indexing unit is used to index the processed data into Elasticsearch in batches for data analysis and query.

[0063] The Elasticsearch consumer module is responsible for reading data from the Kafka message queue module and indexing it into Elasticsearch. Elasticsearch is a distributed search and analysis engine with efficient data storage and retrieval capabilities. The Elasticsearch consumer module subscribes to the specified Topic or Partition from the Kafka cluster through the Kafka Consumer SDK and reads data in batches. After processing, the read data is indexed into Elasticsearch in batches for subsequent data analysis and query. The Elasticsearch consumer module can be configured with multiple consumer threads. Multiple threads can consume data from multiple partitions at the same time (or process batches of documents in parallel), significantly improving throughput to achieve parallel consumption of Kafka topic partitions and improve data processing efficiency and throughput.

[0064] Elasticsearch (abbreviated as ES, such as Figure 2 (As shown in the figure) is a distributed search and analysis engine based on Lucene, designed for scenarios such as real-time search, log analysis, and index storage for massive data. Through reasonable sharding design, hardware resources, and query optimization, Elasticsearch can support search and analysis needs of tens of millions of queries per second.

[0065] like Figure 2 As shown, after display and analysis through ES, the data analysis results are transferred to Kibana, a visualization tool in the Elastic Stack (ELK). Designed specifically for Elasticsearch, Kibana is used for data exploration, analysis, and visualization. It provides interactive dashboards, charts, maps, and a management interface to help users extract insights from massive amounts of data. All Kibana features rely on data stored in Elasticsearch, and Kibana queries and visualizations reflect data changes in ES in real time.

[0066] The data processing method based on Kafka and Elasticsearch described in the preferred embodiment of the present invention is as follows: Figure 3 As shown, the data processing method based on Kafka and Elasticsearch includes the following steps:

[0067] Step S10: reporting data from multiple data sources to the Kafka message queue;

[0068] Step S20: Classify and store the data in the form of topics in the Kafka cluster;

[0069] Step S30: Read data from the Kafka cluster, and index the read data into Elasticsearch after processing for data analysis and query.

[0070] Specifically, data from multiple data sources is obtained through the Kafka Producer SDK or collection tools, including log data, event stream data, and user behavior data; the data compression method and confirmation mechanism are configured, and the obtained data is reported to the Kafka message queue through the Kafka protocol.

[0071] Specifically, the data reported to the Kafka message queue is received; the data is classified and stored in the form of a Topic in the Kafka cluster, where the Kafka cluster is composed of multiple Brokers, each Broker represents a message middleware processing node, and each Topic contains one or more Partitions.

[0072] Specifically, the Kafka Consumer SDK is used to subscribe to a specified Topic or Partition from the Kafka cluster and read data in batches; the read batch data is processed to obtain processed data; and the processed data is indexed in batches into Elasticsearch for data analysis and query.

[0073] The technical effects that the present invention can bring are as follows:

[0074] (1) System integration: Combine Kafka (a high-throughput, distributed message queue based on publish / subscribe mode) and Elasticsearch (a distributed search and analysis engine) to build an efficient data reporting and consumption system.

[0075] (2) Efficient data processing: Through Kafka's high throughput and distributed processing capabilities, and Elasticsearch's efficient data storage and retrieval capabilities, efficient data reporting and consumption are achieved.

[0076] (3) Reliable data transmission: During the data reporting process, you can configure parameters such as data compression method and confirmation mechanism to improve the reliability and stability of data transmission. At the same time, Kafka's replication mechanism and Elasticsearch's data persistence storage further enhance data reliability and security.

[0077] (4) Flexible configuration and expansion: The system supports flexible configuration and expansion. Users can adjust the connection parameters, log configuration, etc. of Kafka and Elasticsearch according to actual needs. They can also customize message processing and indexing logic to meet the needs of different scenarios.

[0078] Furthermore, the present invention can also add the following processing methods:

[0079] (1) Enhanced data security: During data transmission and storage, data encryption and access control can be strengthened to ensure data security. For example, SSL / TLS protocols can be used to encrypt data transmission between Kafka and Elasticsearch, while strict access control policies can be set to limit access to data.

[0080] (2) Optimize data storage and retrieval: You can optimize the Elasticsearch indexing strategy to improve the efficiency of data retrieval. For example, you can design a reasonable index structure and sharding strategy based on the data access pattern and query requirements.

[0081] (3) Support for multiple data formats: The system can be extended to support multiple data formats, such as JSON, XML, CSV, etc. This can be achieved by adding corresponding parsers and converters in the data reporting module so that the system can process data from different data sources.

[0082] (4) Enhance the scalability and fault tolerance of the system: A more flexible system architecture can be designed to easily add more Kafka Brokers or Elasticsearch nodes when needed. At the same time, fault tolerance mechanisms such as automatic retry and failover can be introduced to improve the stability and availability of the system.

[0083] (5) Integration with other data analysis and visualization tools: The system can be integrated with other data analysis and visualization tools (such as Kibana, Grafana, etc.) so that users can visualize and analyze data more conveniently.

[0084] (6) Support for distributed deployment and load balancing: The system can be designed to support distributed deployment and achieve load balancing across multiple nodes. This can improve the system's processing power and scalability while reducing the load and failure risk of a single node.

[0085] (7) Optimize data reporting and consumption performance: You can improve system performance by adjusting the configuration parameters of Kafka and Elasticsearch, as well as optimizing the logic of data reporting and consumption. For example, you can increase the number of Kafka partitions and adjust the batch index size of Elasticsearch.

[0086] Further, if Figure 4 As shown, based on the above-mentioned data processing method and system based on Kafka and Elasticsearch, the present invention also provides a terminal, which includes a processor 10, a memory 20 and a display 30. Figure 4 Only some of the components of the terminal are shown, but it should be understood that implementation of all of the shown components is not required, and more or fewer components may be implemented instead.

[0087] In some embodiments, the memory 20 may be an internal storage unit of the terminal, such as a hard disk or memory of the terminal. In other embodiments, the memory 20 may also be an external storage device of the terminal, such as a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc. equipped on the terminal. Furthermore, the memory 20 may also include both an internal storage unit of the terminal and an external storage device. The memory 20 is used to store application software and various types of data installed on the terminal, such as the program code of the installation terminal. The memory 20 may also be used to temporarily store data that has been output or is to be output. In one embodiment, a data processing program 40 based on Kafka and Elasticsearch is stored on the memory 20, and the data processing program 40 based on Kafka and Elasticsearch can be executed by the processor 10, thereby realizing the data processing method based on Kafka and Elasticsearch in this application.

[0088] In some embodiments, the processor 10 may be a central processing unit (CPU), a microprocessor, or other data processing chip, configured to execute program code stored in the memory 20 or process data, such as executing the data processing method based on Kafka and Elasticsearch.

[0089] In some embodiments, the display 30 may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, or an OLED (Organic Light-Emitting Diode) touchscreen. The display 30 is used to display information on the terminal and to display a visual user interface. The processor 10, memory 20, and display 30 of the terminal communicate with each other via a system bus.

[0090] In one embodiment, when the processor 10 executes the data processing program 40 based on Kafka and Elasticsearch in the memory 20, the following steps are implemented:

[0091] Report data from multiple data sources to the Kafka message queue;

[0092] The data is classified and stored in the Kafka cluster in the form of topics;

[0093] Read data from the Kafka cluster, process the data, and index it into Elasticsearch for data analysis and query.

[0094] The step of reporting data from multiple data sources to the Kafka message queue specifically includes:

[0095] Use the Kafka Producer SDK or collection tools to obtain data from multiple data sources, including log data, event stream data, and user behavior data.

[0096] Configure the data compression method and confirmation mechanism, and report the acquired data to the Kafka message queue through the Kafka protocol.

[0097] The data is classified and stored in the Kafka cluster in the form of topics, specifically including:

[0098] Receive the data reported to the Kafka message queue;

[0099] The data is classified and stored in the form of Topics in a Kafka cluster. The Kafka cluster consists of multiple Brokers, each Broker represents a message middleware processing node, and each Topic contains one or more Partitions.

[0100] The process of reading data from the Kafka cluster and indexing the processed data into Elasticsearch for data analysis and querying includes:

[0101] Use the Kafka Consumer SDK to subscribe to a specified Topic or Partition from the Kafka cluster and read data in batches.

[0102] Process the read batch data to obtain processed data;

[0103] The processed data is batch-indexed into Elasticsearch for data analysis and query.

[0104] The present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a data processing program based on Kafka and Elasticsearch, and when the data processing program based on Kafka and Elasticsearch is executed by a processor, the steps of the data processing method based on Kafka and Elasticsearch as described above are implemented.

[0105] In summary, the present invention provides a data processing system, method, terminal and computer-readable storage medium based on Kafka and Elasticsearch. The system includes: a data reporting module for reporting data from multiple data sources to a Kafka message queue; a Kafka message queue module for receiving the data reported by the data reporting module and storing the data in a Kafka cluster in the form of a Topic; and an Elasticsearch consumption module for reading data from the Kafka message queue module and indexing the read data into Elasticsearch after processing for data analysis and query. The present invention combines Kafka and Elasticsearch to construct an efficient data processing system. Through Kafka's high throughput and distributed processing capabilities and Elasticsearch's efficient data storage and retrieval capabilities, efficient data processing is achieved, data reliability and security are enhanced, and current data processing needs can be met.

[0106] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or terminal comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or terminal. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or terminal comprising the element.

[0107] Of course, those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing related hardware (such as a processor, controller, etc.) through a computer program. The program can be stored in a computer-readable storage medium that can be read by a computer. When the program is executed, it can include the processes in the above-described method embodiments. The computer-readable storage medium can be a memory, a magnetic disk, an optical disk, etc.

[0108] It should be understood that the application of the present invention is not limited to the above examples. For those skilled in the art, improvements or changes can be made based on the above description. All these improvements and changes should fall within the scope of protection of the claims attached to the present invention.

Claims

1. A data processing system based on Kafka and Elasticsearch, characterized in that: The data processing system based on Kafka and Elasticsearch includes: Data reporting module, used to report data from multiple data sources to the Kafka message queue; A Kafka message queue module is used to receive the data reported by the data reporting module and store the data in a Kafka cluster in the form of topics; The Elasticsearch consumption module is used to read data from the Kafka message queue module and index the read data into Elasticsearch after processing for data analysis and query.

2. The data processing system based on Kafka and Elasticsearch according to claim 1, characterized in that The data reporting module includes: A data acquisition unit, configured to acquire data from multiple data sources through the Kafka Producer SDK or an acquisition tool, including log data, event stream data, and user behavior data; The data transmission unit is used to configure the data compression method and confirmation mechanism, and report the data obtained by the data acquisition unit to the Kafka message queue through the Kafka protocol.

3. The data processing system based on Kafka and Elasticsearch according to claim 2, characterized in that The Kafka message queue module includes: a data receiving unit, configured to receive the data reported by the data transmission unit; The data storage unit is used to classify and store the data in the form of Topics in the Kafka cluster. The Kafka cluster is composed of multiple Brokers, each Broker represents a message middleware processing node, and each Topic contains one or more Partitions.

4. The data processing system based on Kafka and Elasticsearch according to claim 3, characterized in that The Elasticsearch consumption module includes: A data reading unit, used to subscribe to a specified Topic or Partition from the Kafka message queue module through the Kafka Consumer SDK and read data in batches; a data processing unit, configured to process the batch data read by the data reading unit to obtain processed data; The data indexing unit is used to index the processed data into Elasticsearch in batches for data analysis and query.

5. A data processing method based on Kafka and Elasticsearch, characterized in that: The data processing method based on Kafka and Elasticsearch includes: Report data from multiple data sources to the Kafka message queue; The data is classified and stored in the Kafka cluster in the form of topics; Read data from the Kafka cluster, process the data, and index it into Elasticsearch for data analysis and query.

6. The data processing method based on Kafka and Elasticsearch according to claim 5, characterized in that: Reporting data from multiple data sources to the Kafka message queue specifically includes: Use the Kafka Producer SDK or collection tools to obtain data from multiple data sources, including log data, event stream data, and user behavior data. Configure the data compression method and confirmation mechanism, and report the acquired data to the Kafka message queue through the Kafka protocol.

7. The data processing method based on Kafka and Elasticsearch according to claim 6, characterized in that: The data is classified and stored in the Kafka cluster in the form of topics, specifically including: Receive the data reported to the Kafka message queue; The data is classified and stored in the form of Topics in a Kafka cluster. The Kafka cluster consists of multiple Brokers, each Broker represents a message middleware processing node, and each Topic contains one or more Partitions.

8. The data processing method based on Kafka and Elasticsearch according to claim 7, characterized in that: The process of reading data from the Kafka cluster and indexing the processed data into Elasticsearch for data analysis and querying includes: Use the Kafka Consumer SDK to subscribe to a specified Topic or Partition from the Kafka cluster and read data in batches. Process the read batch data to obtain processed data; The processed data is batch-indexed into Elasticsearch for data analysis and query.

9. A terminal, characterized in that: The terminal includes: a memory, a processor, and a data processing program based on Kafka and Elasticsearch stored in the memory and runnable on the processor. When the data processing program based on Kafka and Elasticsearch is executed by the processor, the steps of the data processing method based on Kafka and Elasticsearch as described in any one of claims 5 to 8 are implemented.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a data processing program based on Kafka and Elasticsearch. When the data processing program based on Kafka and Elasticsearch is executed by the processor, the steps of the data processing method based on Kafka and Elasticsearch according to any one of claims 5 to 8 are implemented.