A data processing and push method and system based on flink custom sink
A custom Flink sink with Redis-managed topic-table relationships addresses the inflexibility of Flink's sink services by enabling dynamic table management and reducing resource and maintenance costs through flexible data routing and configuration.
Patent Information
- Application Number
- CN202210173702.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-02-24
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2042-02-24
AI Technical Summary
The existing Flink sink services do not meet the requirements for flexible data routing based on message attributes like city code and time for storing data in HDFS directories, and do not allow multiple topics to be handled efficiently.
A custom sink is developed for Flink to handle data routing based on message attributes and store data dynamically in HDFS directories, using Redis to manage topic and table relationships, allowing dynamic addition and removal of jobs without affecting the Kafka cluster.
This approach reduces resource consumption and maintenance costs by enabling dynamic table management and flexible topic configuration, allowing multiple tables to share a Flink job and reducing the need for manual server modifications.
Smart Images

Figure CN114661491B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing technology, and more specifically, to a data processing and pushing method and system based on a flink custom sink. Background Art
[0002] In order to implement text push, the real-time received data is pushed to the specified ftp, sftp server and kafka cluster in real time according to the customer's specifications and format. And realize the flexible configuration of kafka topics to provide consumers with designated topics and consumer groups. If the groups and topics configured in the early stage do not meet the new user consumption requirements, you can modify the configuration on the page to meet the needs without affecting the kafka cluster. Based on Flink and other big data components, the real-time data collection scenarios are increasing. The sink service officially provided by Flink cannot meet our needs. At this time, it can be implemented through a custom sink.
[0003] The sink service officially provided by Flink cannot meet our needs. For example, one of the requirements is to determine the directory to be output to hdfssink based on the city code and time in the message. The official one can only provide one path for all messages in a topic, which cannot meet the actual scenario requirements. Summary of the invention
[0004] In view of the above problems, the present invention provides a data processing and push method based on flink custom sink, including:
[0005] Specify basic information in the flink startup command;
[0006] When the job list page is started, determine whether the topic corresponding to the current job has started the job based on the value of redis;
[0007] If the job is not started, send the flink start command and start the job, and add the relationship between the data topic and the table to redis;
[0008] If the job has been started, only the relationship between the topic and the table will be added to redis;
[0009] When the corresponding message in the process of starting the job passes through the calculation logic of the Flink start command, the storage location of the corresponding message will be determined according to the JSON content of the table corresponding to the job in Redis;
[0010] When deactivating on the job list page, if the table corresponding to the data topic in Redis only has the table corresponding to the current job, run the yarn application-kill command and close the corresponding job to delete the relationship between the data topic and the table in Redis.
[0011] Optionally, the basic information includes: basic job information, input adapter information, and output adapter information.
[0012] Optionally, input adapter information, including: Kafka-related information corresponding to Flink's input stream source.
[0013] Optional, output adapter information, including: information required to start three custom sinks of Flink.
[0014] The present invention also proposes a data processing and push system based on flink custom sink, including:
[0015] Information specification module, which specifies basic information in the flink startup command;
[0016] The job list startup unit, when the job list page is started, determines whether the topic corresponding to the current job has been started according to the value of redis; if the job has not been started, sends the flink startup command and starts the job, and adds the relationship between the topic and the table to redis; if the job has been started, only adds the relationship between the topic and the table to redis; when the corresponding message in the process of starting the job passes through the calculation logic of the flink startup command, the corresponding message storage location is determined according to the json content of the table corresponding to the job in redis;
[0017] The job list deactivation unit is deactivated on the job list page. If the table corresponding to the data receiving topic in redis only has the table corresponding to the current job, the yarn application-kill command is executed to close the corresponding job and delete the relationship between the data receiving topic and the table in redis.
[0018] Optionally, the basic information includes: basic job information, input adapter information, and output adapter information.
[0019] Optionally, input adapter information, including: Kafka-related information corresponding to Flink's input stream source.
[0020] Optional, output adapter information, including: information required to start three custom sinks of Flink.
[0021] The advantages of the present invention are as follows:
[0022] 1. The present invention reduces resource consumption, maintains the advantage of multiple tables accessing one Kafka topic in the current real-time stream processing platform, and allows multiple tables to share one Flink job, thereby reducing resource consumption;
[0023] 2. In the present invention, tables can be dynamically increased or decreased. When a job needs to be taken offline or online, you only need to click online or offline in the job list, which maintains the advantage of dynamic increase or decrease of tables; because flink only needs to know the cluster and topic information of kafka when it starts, the correspondence between topics and tables is stored in redis, and the tables in redis can be dynamically added or deleted when going online or offline.
[0024] 3. The present invention reduces operation and maintenance costs. In order to flexibly provide consumers with designated topics and consumer groups, that is, when the output Kafka related information of a table needs to be modified, it is only necessary to modify the Kafka output configuration in the job output configuration, and there is no need to modify the access program on the server, thereby reducing operation and maintenance costs. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 is a flow chart of the method of the present invention;
[0026] Figure 2 It is a structural diagram of the system of the present invention. DETAILED DESCRIPTION
[0027] Now, exemplary embodiments of the present invention are described with reference to the accompanying drawings. However, the present invention can be implemented in many different forms and is not limited to the embodiments described herein. These embodiments are provided to disclose the present invention in detail and completely and to fully convey the scope of the present invention to those skilled in the art. The terms used in the exemplary embodiments shown in the accompanying drawings are not intended to limit the present invention. In the accompanying drawings, the same units / elements are marked with the same reference numerals.
[0028] Unless otherwise specified, the terms (including technical terms) used herein have the commonly understood meanings to those skilled in the art. In addition, it is understood that the terms defined in commonly used dictionaries should be understood to have the same meanings as those in the context of the relevant fields, and should not be understood as idealized or overly formal meanings.
[0029] The present invention provides a data processing and push method based on flink custom sink, such as Figure 1 As shown, including:
[0030] Specify basic information in the flink startup command;
[0031] When the job list page is started, determine whether the topic corresponding to the current job has started the job based on the value of redis;
[0032] If the job is not started, send the flink start command and start the job, and add the relationship between the data topic and the table to redis;
[0033] If the job has been started, only the relationship between the topic and the table will be added to redis;
[0034] When the corresponding message in the process of starting the job passes through the calculation logic of the Flink start command, the storage location of the corresponding message will be determined according to the JSON content of the table corresponding to the job in Redis;
[0035] When deactivating on the job list page, if the table corresponding to the data topic in Redis only has the table corresponding to the current job, run the yarn application-kill command and close the corresponding job to delete the relationship between the data topic and the table in Redis.
[0036] The present invention will be further described below in conjunction with embodiments:
[0037] 1. Specify the basic job information, input adapter information, and output adapter information in the flink startup command
[0038] 2. The input adapter information contains the relevant information of Kafka corresponding to the input stream source of Flink
[0039] 3. The output adapter information contains the information needed to start the three custom sinks of Flink
[0040] 4. When you click Start on the job list page, determine whether the topic corresponding to the current job has been started based on the value of Redis. If the job has not been started, send the Flink command to start the job and add the relationship between the topic and the table to Redis.
[0041] If it has already been started, only the relationship between the topic and the table is added to redis, and the flink command is not sent to start the job. When the corresponding message passes through the flink calculation logic, the specific location where the message should be stored will be obtained based on the specific json content of the table corresponding to the job in redis.
[0042] When you click Deactivate in the job list, if the table corresponding to the data topic in Redis only has the table corresponding to the current job, run the yarn application-kill command to shut down the corresponding job and delete the relationship between the data topic and the table in Redis.
[0043] The storage structure design of Redis is shown in the following table:
[0044]
[0045] The present invention also proposes a data processing and push system 200 based on a flink custom sink, such as Figure 2As shown, including:
[0046] The information specifying module 201 specifies basic information in the flink startup command;
[0047] The job list startup unit 202, when the job list page is started, determines whether the topic corresponding to the current job has started the job according to the value of redis; if the job has not been started, sends the flink startup command and starts the job, and adds the relationship between the connection topic and the table to redis; if the job has been started, only adds the relationship between the connection topic and the table to redis; when the corresponding message in the process of starting the job passes through the calculation logic of the flink startup command, the corresponding message storage location is determined according to the json content of the table corresponding to the job in redis;
[0048] The job list deactivation unit 203 executes the yarn application-kill command and closes the corresponding job when the job list page is deactivated, if the table corresponding to the data subject in redis only has the table corresponding to the current job, and deletes the relationship between the data subject and the table in redis.
[0049] The basic information includes: basic job information, input adapter information, and output adapter information.
[0050] The input adapter information includes the relevant information of Kafka corresponding to the input stream source of Flink.
[0051] Among them, the output adapter information includes: the information required to start three custom sinks of Flink.
[0052] The advantages of the present invention are as follows:
[0053] 1. The present invention reduces resource consumption, maintains the advantage of multiple tables accessing one Kafka topic in the current real-time stream processing platform, and allows multiple tables to share one Flink job, thereby reducing resource consumption;
[0054] 2. In the present invention, tables can be dynamically increased or decreased. When a job needs to be taken offline or online, you only need to click online or offline in the job list, which maintains the advantage of dynamic increase or decrease of tables; because flink only needs to know the cluster and topic information of kafka when it starts, the correspondence between topics and tables is stored in redis, and the tables in redis can be dynamically added or deleted when going online or offline.
[0055] 3. The present invention reduces operation and maintenance costs. In order to flexibly provide consumers with designated topics and consumer groups, that is, when the output Kafka related information of a table needs to be modified, it is only necessary to modify the Kafka output configuration in the job output configuration, and there is no need to modify the access program on the server, thereby reducing operation and maintenance costs.
[0056] It will be appreciated by those skilled in the art that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes. The schemes in the embodiments of the present invention may be implemented in various computer languages, for example, object-oriented programming language Java and literal scripting language JavaScript, etc.
[0057] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0058] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0059] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0060] Although the preferred embodiments of the present invention have been described, those skilled in the art may make other changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention.
[0061] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalents, the present invention is also intended to include these modifications and variations.
Claims
1. A data processing and push method based on flink custom sink, the method comprising: Specify basic information in the flink startup command; When the job list page is started, determine whether the topic corresponding to the current job has started the job based on the value of redis; If the job is not started, send the flink start command and start the job, and add the relationship between the data topic and the table to redis; If the job has been started, only the relationship between the topic and the table will be added to redis; When the corresponding message in the process of starting the job passes through the calculation logic of the Flink start command, the storage location of the corresponding message will be determined according to the JSON content of the table corresponding to the job in Redis; When deactivating on the job list page, if the table corresponding to the data topic in Redis only has the table corresponding to the current job, run the yarn application-kill command and close the corresponding job to delete the relationship between the data topic and the table in Redis.
2. The method according to claim 1, wherein the basic information comprises: Basic job information, input adapter information, output adapter information.
3. The method according to claim 2, wherein the input adapter information comprises: Related information about Kafka corresponding to Flink's input stream source.
4. The method according to claim 2, wherein the information of the output adapter comprises: Information required to start three custom sinks of Flink.
5. A data processing and push system based on flink custom sink, the system comprising: Information specification module, specifies basic information in the flink startup command; The job list startup unit, when the job list page is started, determines whether the topic corresponding to the current job has been started according to the value of redis; if the job has not been started, sends the flink startup command and starts the job, and adds the relationship between the topic and the table to redis; if the job has been started, only adds the relationship between the topic and the table to redis; when the corresponding message in the process of starting the job passes through the calculation logic of the flink startup command, the corresponding message storage location is determined according to the json content of the table corresponding to the job in redis; The job list deactivation unit is deactivated on the job list page. If the table corresponding to the data receiving topic in redis only has the table corresponding to the current job, the yarn application-kill command is executed to close the corresponding job and delete the relationship between the data receiving topic and the table in redis.
6. The system according to claim 5, wherein the basic information comprises: Basic job information, input adapter information, output adapter information.
7. The system according to claim 6, wherein the input adapter information comprises: Related information about Kafka corresponding to Flink's input stream source.
8. The system according to claim 6, wherein the information of the output adapter comprises: Information required to start three custom sinks of Flink.
Citation Information
Patent Citations
Stream data processing method, system, apparatus, and computer-readable storage medium
CN109254982A
Data transmission method based on redis storage as message middleware
CN112882842A