Method, system and device for shunting Kafka data based on Flink and medium
By configuring KafkaSource in the Flink environment and converting it into a flow table form, and combining conditional judgment to realize Kafka data shunt, the problems of low processing efficiency, poor real-time and insufficient flexibility in traditional methods are solved, and efficient, real-time and flexible data shunt is achieved.
Patent Information
- Application Number
- CN202510071312.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2025-05-16
AI Technical Summary
The traditional data shunt method has low processing efficiency, is difficult to achieve real-time shunt, and is not flexible enough to support complex shunt rules, which cannot meet the needs of efficient, real-time and flexible shunt in big data processing.
Using Flink-based method, we can obtain the Flink execution environment, create a flow table environment, configure KafkaSource, convert it into a flow table form, and apply conditional judgment on the flow table to achieve efficient, real-time and flexible shunt of Kafka data.
It realizes efficient shunt of Kafka data, has high processing efficiency, strong real-time and good flexibility, can meet complex business needs, and is suitable for real-time business scenarios.
Smart Images

Figure CN120011171A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of big data processing and relates to a method, system, device and medium for distributing Kafka data based on Flink. Background Art
[0002] In the era of big data, the amount of data generated and transmitted is exploding. As a high-throughput, distributed message queue system, Kafka is widely used in data collection, transmission, and storage. However, in actual application scenarios, it is often necessary to further process and divert the data obtained from Kafka to meet the differentiated requirements of different business needs for data. For example, different types of log data can be classified and processed according to business modules, or they can be distributed to different subsequent processing processes according to the importance of the data.
[0003] Traditional data diversion methods have some limitations, such as low processing efficiency, difficulty in achieving real-time diversion, and insufficient flexibility in supporting complex diversion rules. Therefore, it is necessary to propose a new method and system that is efficient, flexible, and capable of diverting Kafka data in real time. Summary of the invention
[0004] The purpose of the present invention is to overcome the shortcomings of the above-mentioned prior art and provide a method, system, device and medium for diverting Kafka data based on Flink. The method, system, device and medium can realize the diversion of Kafka data and have the characteristics of high processing efficiency, strong real-time performance and good flexibility.
[0005] In order to achieve the above object, the present invention adopts the following technical scheme:
[0006] In one aspect, the present invention provides a method for shunting Kafka data based on Flink, comprising:
[0007] Get the Flink execution environment;
[0008] Create a flow table environment;
[0009] Configure KafkaSource;
[0010] Convert the configured KafkaSource into a stream table;
[0011] According to the diversion requirements, conditional judgment is applied to the flow table to achieve the diversion of Kafka data.
[0012] The method for shunting Kafka data based on Flink described in the present invention is further improved in that:
[0013] Furthermore, the process of obtaining the Flink execution environment is as follows:
[0014] Get the Flink execution environment through StreamExecutionEnvironment.get_execution_environment().
[0015] Furthermore, the process of creating the flow table environment is as follows:
[0016] Use StreamTableEnvironment.create(env) to create a stream table environment.
[0017] Furthermore, the process of configuring KafkaSource is as follows:
[0018] Set the Kafka server address, the topic to consume, the consumer group ID, and the offset from which to start consuming.
[0019] Furthermore, the process of converting the configured KafkaSource into a stream table is as follows:
[0020] Convert the configured KafkaSource into a stream table through the t_env.from_source(kafka_source,"kafka","kafka_table") operation.
[0021] In a second aspect of the present invention, the present invention provides a system for shunting Kafka data based on Flink, including:
[0022] Get module, used to get Flink's execution environment;
[0023] Create a module to create a flow table environment;
[0024] Configuration module, used to configure KafkaSource;
[0025] The conversion module is used to convert the configured KafkaSource into a stream table format;
[0026] The diversion module is used to apply conditional judgment on the flow table according to the diversion requirements to realize the diversion of Kafka data.
[0027] The system for distributing Kafka data based on Flink described in the present invention is further improved in that:
[0028] Furthermore, the process of obtaining the Flink execution environment is as follows:
[0029] Get the Flink execution environment through StreamExecutionEnvironment.get_execution_environment().
[0030] Furthermore, the process of creating the flow table environment is as follows:
[0031] Use StreamTableEnvironment.create(env) to create a stream table environment.
[0032] In a third aspect of the present invention, the present invention provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the method for shunting Kafka data based on Flink when executing the computer program.
[0033] In a fourth aspect of the present invention, the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the method for shunting Kafka data based on Flink are implemented.
[0034] The present invention has the following beneficial effects:
[0035] The method, system, device and medium for shunting Kafka data based on Flink described in the present invention can flexibly set shunting rules in a flow table environment during specific operations. At the same time, by utilizing Flink's distributed computing capabilities and stream processing characteristics, a large amount of Kafka data can be quickly processed to achieve efficient shunting operations with good real-time performance, thus meeting the needs of real-time business scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] The accompanying drawings constituting a part of the present invention are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the accompanying drawings:
[0037] Figure 1 The figure is a flow chart of the method of the present invention. DETAILED DESCRIPTION
[0038] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0039] In the description of the present invention, it should be understood that the terms “include” and “comprises” indicate the presence of described features, wholes, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or collections thereof.
[0040] It should also be understood that the terms used in the present specification are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in the present specification and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include plural forms.
[0041] It should be further understood that the term "and / or" used in the present specification and the appended claims refers to any combination of one or more of the associated listed items and all possible combinations, and includes these combinations. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in the present invention generally indicates that the associated objects are in an "or" relationship.
[0042] It should be understood that, although the terms first, second, third, etc. may be used to describe preset ranges, etc. in the embodiments of the present invention, these preset ranges should not be limited to these terms. These terms are only used to distinguish preset ranges from each other. For example, without departing from the scope of the embodiments of the present invention, the first preset range may also be referred to as the second preset range, and similarly, the second preset range may also be referred to as the first preset range.
[0043] The word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining" or "in response to detecting", depending on the context. Similarly, the phrases "if it is determined" or "if (stated condition or event) is detected" may be interpreted as "when it is determined" or "in response to determining" or "when detecting (stated condition or event)" or "in response to detecting (stated condition or event)", depending on the context.
[0044] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. The components of the embodiments of the present invention described and shown in the drawings here can usually be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0045] Various structural schematic diagrams of the embodiments disclosed in the present invention are shown in the accompanying drawings. These figures are not drawn to scale, and some details are magnified and some details may be omitted for the purpose of clear expression. The shapes of various regions and layers shown in the figures and the relative sizes and positional relationships therebetween are only exemplary, and may deviate in practice due to manufacturing tolerances or technical limitations, and those skilled in the art may additionally design regions / layers with different shapes, sizes, and relative positions according to actual needs.
[0046] As is known, flow table is a key data structure for storing and managing network data flow, which has different meanings and functions in different technical backgrounds and application scenarios. The following is a detailed explanation of flow table: Flow table can be regarded as a collection of events or an abstraction of the data forwarding function of network devices. In specific scenarios, flow table is used to store events that occur over time, and these events are continuously added to the table, so the collection is unbounded. In network devices, flow table is responsible for the search and forwarding of data packets, which is the basis for network devices to implement data forwarding functions. Event storage: In some application scenarios, such as log analysis and real-time monitoring, flow table is used to store events that occur over time. These events contain time attributes, such as ingestion time (the time when data is written to the flow engine), event time (the time defined in the business logic) and processing time (the time when the flow engine calculates and processes data). Flow table supports setting data expiration time, which defaults to 7 days, in order to manage the life cycle of data. Network devices: In network devices, such as OpenFlow switches, flow table is used to search and forward data packets. Each flow table entry defines a set of matching conditions and corresponding actions. When a data packet arrives, the network device will match it according to the rules in the flow table and perform corresponding actions, such as forwarding to a specified port, discarding, etc. OpenFlow flow table entries usually include a header field (for packet matching), a counter (for counting the number of matching packets), and an action (for indicating how to handle matching packets). OVS (Open vSwitch): The flow table of OVS is the key data structure for packet forwarding. It contains multiple flow table entries, each of which defines a set of matching conditions and corresponding actions. The flow table of OVS uses a caching mechanism to speed up the processing of data packets and reduce network latency. By configuring the flow table, OVS can achieve network isolation and communication between virtual machines, and supports a variety of network tunneling technologies, allowing virtual machines to communicate across physical networks.
[0047] Flink is an open source stream processing framework developed by the Apache Software Foundation. Its core is a distributed streaming data flow engine written in Java and Scala. The following is a detailed explanation of Flink: Flink is designed to process large-scale, high-throughput real-time data streams and batch data. It provides an efficient, reliable, and scalable way to process and analyze real-time data. Flink was originally designed to meet the needs of high-throughput, low-latency data stream processing. It supports unified processing of unbounded and bounded data, and regards batch processing as a special case of stream processing. Stream processing: Flink can process real-time data streams and provides low-latency data processing capabilities. Batch processing: In addition to stream processing, Flink can also process batch data. It can convert batch processing jobs into stream processing jobs and provides the same functions and optimizations as stream processing. Event time processing: Flink supports event time processing and can process data streams ordered by event time. Window operations: Flink provides a rich window operation that can group data streams, split them by time or quantity, and perform aggregation operations. State management: Flink can manage and maintain state information in stream processing to facilitate complex calculations and transformations. Event-driven processing: Flink supports an event-based processing model that can trigger and process specific events. Reliability guarantee: Flink provides a fault recovery mechanism to ensure the consistency and reliability of calculations. Data connection and integration: Flink can connect and integrate with various data sources and data storage, including message queues, databases, file systems, etc.
[0048] Embodiment 1
[0049] refer to Figure 1 The method for shunting Kafka data based on Flink of the present invention comprises the following steps:
[0050] 1) Get the Flink execution environment and set parameters;
[0051] The specific operations of step 1) are:
[0052] Get the Flink execution environment through StreamExecutionEnvironment.get_execution_environment(). During this process, set related parameters such as parallelism. For example, set the parallelism to 1 in the sample code.
[0053] 2) Create a flow table environment;
[0054] Use StreamTableEnvironment.create(env) to create a stream table environment to facilitate subsequent table operations on the data.
[0055] 3)Configure KafkaSource;
[0056] Specifically, configure KafkaSource and set the following Kafka-related parameters:
[0057] bootstrap_servers: Kafka server address.
[0058] topics: Topics to consume.
[0059] group_id: consumer group ID.
[0060] starting_offsets: The offset from which consumption starts.
[0061] Use the builder mode to build KafkaSource according to the parameters set above.
[0062] 4) Convert to flow table format;
[0063] Convert the configured KafkaSource into a stream table.
[0064] Specifically, the configured KafkaSource is converted into a stream table through the t_env.from_source(kafka_source,"kafka","kafka_table") operation, so that the data can be processed in the table environment.
[0065] 5) Perform data diversion;
[0066] According to the specific diversion requirements, conditional judgment is applied to the flow table to realize the diversion of Kafka data. For example, in the example, diversion is performed according to the value of the type field. The specific operations are as follows:
[0067] Use the where statement to filter out data that meets the conditions.
[0068] Use the add_sink directive to send data with different conditions to different targets, for example, sink1 and sink2 in the example.
[0069] The present invention has the following characteristics:
[0070] a) Efficiency: This paper uses Flink's distributed computing capabilities and stream processing characteristics to quickly process large amounts of Kafka data and achieve efficient diversion operations. Compared with traditional data diversion methods, it greatly improves data processing speed and diversion efficiency.
[0071] b) Real-time: Data can be obtained from Kafka in real time and diverted for processing, ensuring that the data can be distributed to the corresponding processing flow in a timely manner according to predetermined rules to meet the needs of real-time business scenarios.
[0072] c) Flexibility: By flexibly setting diversion rules in the flow table environment, for example, according to different field values, multiple condition combinations, etc., it can adapt to various complex business needs and accurately divert different types of Kafka data.
[0073] Embodiment 2
[0074] The system for distributing Kafka data based on Flink of the present invention includes:
[0075] Get module, used to get Flink's execution environment;
[0076] Create a module to create a flow table environment;
[0077] Configuration module, used to configure KafkaSource;
[0078] The conversion module is used to convert the configured KafkaSource into a stream table format;
[0079] The diversion module is used to apply conditional judgment on the flow table according to the diversion requirements to realize the diversion of Kafka data.
[0080] As an implementation manner of the present invention, the process of obtaining the execution environment of Flink is: obtaining the execution environment of Flink through StreamExecutionEnvironment.get_execution_environment().
[0081] As an implementation manner of the present invention, the process of creating a flow table environment is: using StreamTableEnvironment.create(env) to create a flow table environment.
[0082] As an implementation mode of the present invention, the process of configuring KafkaSource is:
[0083] Set the Kafka server address, the topic to consume, the consumer group ID, and the offset from which to start consuming.
[0084] As an implementation mode of the present invention, the process of converting the configured KafkaSource into a stream table is as follows:
[0085] Convert the configured KafkaSource into a stream table through the t_env.from_source(kafka_source,"kafka","kafka_table") operation.
[0086] The division of modules in the embodiments of the present application is schematic and is only a logical function division. There may be other division methods in actual implementation. In addition, each functional module in each embodiment of the present application may be integrated into a processor, or may exist physically separately, or two or more modules may be integrated into one module. The above-mentioned integrated modules may be implemented in the form of hardware or in the form of software functional modules.
[0087] Embodiment 3
[0088] A computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the method for shunting Kafka data based on Flink are implemented, for example, including: obtaining the execution environment of Flink; creating a flow table environment; configuring KafkaSource; converting the configured KafkaSource into a flow table form; applying conditional judgment on the flow table according to the shunting requirements to realize the shunting of Kafka data. Among them, the memory may include a memory, such as a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk memory, etc.; the processor, the network interface, and the memory are interconnected through an internal bus, and the internal bus may be an industrial standard architecture bus, a peripheral component interconnection standard bus, an extended industrial standard structure bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. The memory is used to store programs. Specifically, the program may include a program code, and the program code includes computer operation instructions. The memory may include a memory and a non-volatile memory, and provide instructions and data to the processor.
[0089] Embodiment 4
[0090] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the method for diverting Kafka data based on Flink are implemented, for example, including: obtaining the execution environment of Flink; creating a flow table environment; configuring KafkaSource; converting the configured KafkaSource into a flow table form; applying conditional judgment on the flow table according to the diversion requirements to implement the diversion of Kafka data. Specifically, the computer-readable storage medium includes, but is not limited to, for example, volatile memory and / or non-volatile memory. The volatile memory may include random access memory (RAM) and / or cache memory (cache), etc. The non-volatile memory may include a read-only memory (ROM), a hard disk, a flash memory, an optical disk, a magnetic disk, etc.
[0091] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that include computer-usable program code.
[0092] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0093] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0094] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0095] Those skilled in the art will readily appreciate other embodiments of the present invention after considering the specification and disclosure of the invention. This application is intended to cover any variations, uses or adaptations of the present invention that follow the general principles of the present invention and include common knowledge or customary techniques in the art that are not disclosed by the present invention. The specification and examples are to be considered exemplary only, and the true scope and spirit of the present invention are indicated by the following claims.
[0096] It should be understood that the present invention is not limited to the exact construction that has been described above and shown in the drawings and that various modifications and changes may be made without departing from the scope thereof. The scope of the present invention is limited only by the appended claims.
[0097] The above description is only a preferred embodiment of the present invention and does not limit the present invention in any way. Any simple modification, change and equivalent structural change made to the above embodiment based on the technical essence of the present invention still falls within the protection scope of the technical solution of the present invention.
Claims
1. A method for shunting Kafka data based on Flink, characterized in that: include: Get the Flink execution environment; Create a flow table environment; Configure KafkaSource; Convert the configured KafkaSource into a stream table; According to the diversion requirements, conditional judgment is applied to the flow table to achieve the diversion of Kafka data.
2. The method for shunting Kafka data based on Flink according to claim 1 is characterized in that: The process of obtaining the Flink execution environment is as follows: Get the Flink execution environment through StreamExecutionEnvironment.get_execution_environment().
3. The method for shunting Kafka data based on Flink according to claim 1 is characterized in that: The process of creating a flow table environment is as follows: Use StreamTableEnvironment.create(env) to create a stream table environment.
4. The method for shunting Kafka data based on Flink according to claim 1 is characterized in that: The process of configuring KafkaSource is as follows: Set the Kafka server address, the topic to consume, the consumer group ID, and the offset from which to start consuming.
5. The method for shunting Kafka data based on Flink according to claim 1, characterized in that: The process of converting the configured KafkaSource into a stream table is as follows: Convert the configured KafkaSource into a stream table through the t_env.from_source(kafka_source,"kafka","kafka_table") operation.
6. A system for shunting Kafka data based on Flink, characterized in that: include: Get module, used to get Flink's execution environment; Create a module to create a flow table environment; Configuration module, used to configure KafkaSource; The conversion module is used to convert the configured KafkaSource into a stream table format; The diversion module is used to apply conditional judgment on the flow table according to the diversion requirements to realize the diversion of Kafka data.
7. The system for distributing Kafka data based on Flink according to claim 6, characterized in that: The process of obtaining the Flink execution environment is as follows: Get the Flink execution environment through StreamExecutionEnvironment.get_execution_environment().
8. The system for distributing Kafka data based on Flink according to claim 6, characterized in that: The process of creating a flow table environment is as follows: Use StreamTableEnvironment.create(env) to create a stream table environment.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the method for shunting Kafka data based on Flink are implemented as described in any one of claims 1 to 5.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method for shunting Kafka data based on Flink are implemented as described in any one of claims 1 to 5.