Metadata change capturing method and system based on Flink CDC mode evolution
By using Kafka data source components and event memory in the metadata change capture method evolved in the Flink CDC mode, the metadata change event capture problem is solved when upstream business databases cannot enable the database change log capture feature, and efficient data integration tasks are achieved.
Patent Information
- Application Number
- CN202510139256.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-08
- Publication Date
- 2025-05-16
AI Technical Summary
When upstream business databases cannot enable database change log capture feature, it is difficult for the existing technology to effectively capture and handle metadata change events, resulting in high difficulty and low efficiency in data integration tasks.
A metadata change capture method based on the evolution of Flink CDC mode is designed, using the Kafka data source component, Kafka Connect JDBC record deserializer, memory event memory and Zookeeper event memory, read data and metadata through the JDBC general interface, and generate and send data change and metadata change events to downstream operators of Flink jobs.
This method reduces the difficulty of integrating business data into data warehouses, improves implementation efficiency, and can accurately capture and handle metadata change events when the upstream business database cannot enable the change log capture feature.
Smart Images

Figure CN120011110A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing technology, and in particular to a metadata change capture method and system based on Flink CDC mode evolution. Background Art
[0002] Flink CDC aims to provide users with a more comprehensive set of programming interfaces. It is a stream-based data integration tool. This tool allows users to elegantly define their ETL processes in the form of YAML configuration files, and helps users automatically generate customized Flink operators and submit Flink jobs.
[0003] As attached Figure 1 The Flink CDC shown in the figure is mainly composed of data source operators, transformation operators (Transformer), routing operators (Router), and data target operators. The data source operator reads the data change logs of upstream relational databases such as MySQL, Postgresql, KaiwuDB, and SQLServer to generate metadata change events and data change events, and optionally sends the time to the transformation and routing operators, and finally passes the event to the data target operator. The data target operator sinks the metadata and data into the data warehouse according to the event type; the Flink CDC pattern evolution feature ensures that when jobs are executed in parallel, the pattern saved in the pattern registration center is consistent with the business data, ensuring that downstream operators can be executed accurately and finally completing the data integration task.
[0004] In actual data integration tasks, the database change log capture feature cannot be enabled due to upstream business database versions, business system security, and other reasons. Therefore, the solution of reading data and metadata through the JDBC general interface becomes the best solution in this scenario. In order to support the solution of reading data and metadata using the JDBC general interface in the scenario where the upstream business database cannot enable the database change log capture feature in actual data integration tasks, it is necessary to design and implement a metadata change capture method that reads messages in the Kafka cluster, generates data change and metadata change events, and sends them to the downstream operators of the Flink job. Summary of the invention
[0005] The technical task of the present invention is to address the above shortcomings and provide a metadata change capture method and system based on the evolution of Flink CDC mode, which can reduce the implementation difficulty of integrating business data into the data warehouse and improve the implementation efficiency.
[0006] The technical solution adopted by the present invention to solve the technical problem is:
[0007] A metadata change capture method based on the evolution of Flink CDC mode. The implementation of this method includes Kafka data source components, Kafka Connect JDBC record deserializer, in-memory event storage, and Zookeeper event storage.
[0008] The Kafka data source factory creates a Kafka data source component based on the parameters and their values defined in the data source option parameter list; the Kafka data source component is responsible for interacting with the topic specified by the Kafka cluster, reading Kafka message records according to the partition start offset specified by the option parameter, maintaining consumer groups and offsets, and ensuring at least once or exactly once semantics; it is responsible for creating the Kafka Connect JDBC record deserialization component and passing the Kafka message record to it; it creates an in-memory event store or Zookeeper event store according to the event store type specified by the option parameter.
[0009] Further, the Kafka Connect JDBC record deserializer parses the topic name or the Header information in the Kafka message record to generate the tableId according to the TableId parsing method specified by the option parameter;
[0010] Deserialize received Kafka records into business data and metadata schema; pass the metadata schema to the event storage for further processing;
[0011] Generate data change (insert) events based on business data and pass them to downstream operators (conversion operators, routing operators or data target operators).
[0012] Furthermore, the TableId includes a database name, a mode name, and a data table name.
[0013] Furthermore, the memory event storage stores metadata schemas, and compares the newly received schema with the stored schema, generates metadata change events such as table creation events, column addition events, column deletion events, and column type change events according to the differences, and passes them to the SchemaOperator operator.
[0014] The in-memory event store stores the metadata schema in memory and is not shared among the nodes of a Flink job, so it is only used during testing.
[0015] Furthermore, after the Zookeeper event storage is started, the FlinkCDC job name specified in the creation parameter list is used as the root Znode temporary node ( / {flink_cdc_job_name}); the TableId (database name, mode name, data table name) is connected as the Znode temporary node name ( / {flink_cdc_job_name} / {table_id}), and its corresponding metadata schema is deserialized into binary data and saved as a value in the Znode; at the same time, a method of reading the binary data of the Znode node and deserializing it into the metadata schema is supported; the Znode temporary node name lock ( / {flink_cdc_job_name} / {table_id} / lock) is used as a table-level distributed lock.
[0016] Furthermore, the program flow of the table creation event is as follows:
[0017] First, determine whether the table schema has been saved in the " / {flink_cdc_job_name} / {table_id}" Znode temporary node of Zookeeper. If it already exists, record in memory that the table creation event has been sent; otherwise, obtain the table distributed lock and generate the table creation event and send it to the downstream SchemaOperator operator;
[0018] If the acquisition of the table distributed lock times out, that is, another Flink Task Slot acquires the distributed lock, re-determine whether the table schema has been persisted;
[0019] Finally, release the distributed lock of the table.
[0020] Furthermore, the procedure flow for other metadata change events is as follows:
[0021] First, if the schema of the table changes, obtain the distributed lock of the table, traverse the metadata to generate a schema change event, send the change event, and persist the new schema of the table;
[0022] If the acquisition of the table's distributed lock times out, that is, another Flink Task Slot acquires the distributed lock, re-read the Schema in Zookeeper and re-determine whether the Schema has changed;
[0023] Finally, release the distributed lock of the table;
[0024] The schema change events include table creation events, column addition events, column deletion events, and column type change events.
[0025] The present invention also claims a metadata change capture system based on the evolution of Flink CDC mode, including a Kafka data source component, a Kafka Connect JDBC record deserializer, an in-memory event storage, and a Zookeeper event storage;
[0026] The Kafka data source component is responsible for creating the Kafka Connect JDBC record deserialization component and passing the Kafka message record to it; creating an in-memory event store or Zookeeper event store according to the event store type specified by the option parameter;
[0027] The system specifically implements metadata change capture based on the evolution of Flink CDC mode through the above method.
[0028] The present invention also claims a metadata change capture device based on Flink CDC mode evolution, comprising: at least one memory and at least one processor;
[0029] The at least one memory is used to store a machine-readable program;
[0030] The at least one processor is used to call the machine-readable program to implement the above method.
[0031] The present invention also claims protection for a computer-readable medium having computer instructions stored thereon, which can implement the above method when executed by a processor.
[0032] Compared with the prior art, the metadata change capture method and system based on Flink CDC mode evolution of the present invention has the following beneficial effects:
[0033] The metadata change capture method designed and implemented by the present invention, in the scenario where the upstream business database cannot enable the database change log capture feature, generates data change and metadata change events by reading business data and metadata through the JDBC general interface and sends them to the corresponding operators downstream of the Flink CDC job, so as to finally complete the data integration task. The device can reduce the implementation difficulty of integrating business data into the data warehouse and improve the implementation efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 It is a schematic diagram of the Flink CDC Pipeline component provided by the background technology of the present invention;
[0035] Figure 2 This is a schematic diagram of Kafka Connect & Flink CDC components provided by an embodiment of the present invention;
[0036] Figure 3 It is a schematic diagram of a metadata change capture method based on Flink CDC mode evolution provided by an embodiment of the present invention;
[0037] Figure 4 is a flowchart of a table creation event program provided by an embodiment of the present invention;
[0038] Figure 5 This is a flowchart of other metadata change event procedures provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0039] The present invention will be further described below in conjunction with specific embodiments.
[0040] The embodiment of the present invention provides a metadata change capture method based on the evolution of Flink CDC mode. In the actual data integration task described in the background technology, the database change log capture feature cannot be enabled due to the upstream business database version, business system security and other reasons. The data and metadata are read through the JDBC universal interface. Figure 2 The kafka-connect-jdbc-source component in the Kafka Connect technology stack shown in the figure regularly assembles new business data and metadata into custom message records and publishes them to the Kafka cluster. The flink-cdc-pipeline-connector-kafka-source component is custom designed and implemented according to the Flink CDCPipeline interface specification. This device can read messages in the Kafka cluster, generate data change and metadata change events, and send them to the downstream operators of the Flink job, thus completing the data integration task.
[0041] like Figure 3 As shown, the specific implementation of this method includes a Kafka data source component, a Kafka Connect JDBC record deserializer, an in-memory event storage, and a Zookeeper event storage.
[0042] The Kafka data source factory creates a Kafka data source component based on the parameters and their values defined in the data source option parameter list; the Kafka data source component is responsible for interacting with the topic specified by the Kafka cluster, reading Kafka message records according to the partition start offset specified by the option parameter, maintaining consumer groups and offsets, and ensuring at least once or exactly once semantics; it is responsible for creating the Kafka Connect JDBC record deserialization component and passing the Kafka message record to it; it creates an in-memory event store or Zookeeper event store according to the event store type specified by the option parameter.
[0043] The Kafka Connect JDBC record deserializer is responsible for parsing the topic name or the Header information in the Kafka message record to generate tableId according to the TableId (database name, mode name, data table name) parsing method specified by the option parameter; deserializing the received Kafka record into business data and metadata schema; passing the metadata schema to the event storage for further processing; generating data change (insert) events based on business data and passing them to downstream operators (conversion operators, routing operators or data target operators).
[0044] The memory event storage is responsible for saving the metadata schema, comparing the newly received schema with the saved schema, generating metadata change events such as table creation events, column addition events, column deletion events, and column type change events according to the differences, and passing them to the SchemaOperator operator. The memory event storage stores the metadata schema in memory and will not be shared among the nodes of the Flink job, so it is only used during testing.
[0045] After the Zookeeper event storage is started, the Flink CDC job name specified in the creation parameter list is used as the root Znode temporary node ( / {flink_cdc_job_name}); the TableId (database name, mode name, data table name) is connected as the Znode temporary node name ( / {flink_cdc_job_name} / {table_id}), and its corresponding metadata schema is deserialized into binary data and saved as a value in the Znode; at the same time, it supports the method of reading the binary data of the Znode node and deserializing it into the metadata schema; the Znode temporary node name lock ( / {flink_cdc_job_name} / {table_id} / lock) is used as a table-level distributed lock.
[0046] like Figure 4 The following figure shows the program flow chart of the table creation event. First, it determines whether the table schema has been saved in the temporary Znode " / {flink_cdc_job_name} / {table_id}" of Zookeeper. If it already exists, it records in the memory that the table creation event has been sent. Otherwise, after acquiring the table distributed lock, it generates the table creation event and sends it to the downstream SchemaOperator operator. If the acquisition of the table distributed lock times out, that is, other Flink Task Slots obtain the distributed lock, it re-determines whether the table schema has been persisted. Finally, the table distributed lock is released.
[0047] like Figure 5 As shown in the figure, it is a program flow chart for other metadata change events. First, under the premise that the schema of the table changes, the distributed lock of the table is obtained, and then the metadata is traversed to generate schema change events (table creation events, column addition events, column deletion events, column type change events) and the new schema of the table is persisted after the change events are sent; if the acquisition of the distributed lock of the table times out, that is, other Flink Task Slots obtain the distributed lock, the schema in Zookeeper is re-read and it is re-determined whether the schema change occurs; finally, the distributed lock of the table is released.
[0048] This method implements the flink-cdc-pipeline-kafka-source component based on the interface specification of Flink CDC Pipeline, uses Zookeeper to store metadata information, and provides distributed lock capabilities for metadata change events, so that the Flink CDC job executed in parallel will only send the same metadata table creation event or change event to the downstream operator once. In the application scenario of the data service platform, the data in the upstream business system needs to be transferred to the data warehouse through ETL (extraction, transformation, loading). In actual applications, the database of the upstream business system cannot enable the change log capture feature due to objective reasons such as version and security. The metadata change capture device can generate metadata change events and send them to the downstream operators of the Flink CDC task by maintaining the metadata read by the JDBC general interface. This device can reduce the difficulty of integrating business data into the data warehouse and improve the implementation efficiency.
[0049] The embodiment of the present invention also provides a metadata change capture system based on the evolution of Flink CDC mode, including a Kafka data source component, a Kafka Connect JDBC record deserializer, an in-memory event storage, and a Zookeeper event storage;
[0050] The system implements metadata change capture based on Flink CDC mode evolution through the metadata change capture method based on Flink CDC mode evolution described in the above embodiment.
[0051] The Kafka data source factory creates a Kafka data source component based on the parameters and their values defined in the data source option parameter list; the Kafka data source component is responsible for interacting with the topic specified by the Kafka cluster, reading Kafka message records according to the partition start offset specified by the option parameter, maintaining consumer groups and offsets, and ensuring at least once or exactly once semantics; it is responsible for creating the Kafka Connect JDBC record deserialization component and passing the Kafka message record to it; it creates an in-memory event store or Zookeeper event store according to the event store type specified by the option parameter.
[0052] The Kafka Connect JDBC record deserializer is responsible for parsing the topic name or the Header information in the Kafka message record to generate tableId according to the TableId (database name, mode name, data table name) parsing method specified by the option parameter; deserializing the received Kafka record into business data and metadata schema; passing the metadata schema to the event storage for further processing; generating data change (insert) events based on business data and passing them to downstream operators (conversion operators, routing operators or data target operators).
[0053] The memory event storage is responsible for saving the metadata schema, comparing the newly received schema with the saved schema, generating metadata change events such as table creation events, column addition events, column deletion events, and column type change events according to the differences, and passing them to the SchemaOperator operator. The memory event storage stores the metadata schema in memory and will not be shared among the nodes of the Flink job, so it is only used during testing.
[0054] After the Zookeeper event storage is started, the Flink CDC job name specified in the creation parameter list is used as the root Znode temporary node ( / {flink_cdc_job_name}); the TableId (database name, mode name, data table name) is connected as the Znode temporary node name ( / {flink_cdc_job_name} / {table_id}), and its corresponding metadata schema is deserialized into binary data and saved as a value in the Znode; at the same time, it supports the method of reading the binary data of the Znode node and deserializing it into the metadata schema; the Znode temporary node name lock ( / {flink_cdc_job_name} / {table_id} / lock) is used as a table-level distributed lock.
[0055] The program flow of the table creation event is as follows:
[0056] First, determine whether the table schema has been saved in the " / {flink_cdc_job_name} / {table_id}" Znode temporary node of Zookeeper. If it already exists, record in memory that the table creation event has been sent; otherwise, obtain the table distributed lock and generate the table creation event and send the table creation event to the downstream SchemaOperator operator; if the acquisition of the table distributed lock times out, that is, other Flink Task Slots obtain the distributed lock, re-determine whether the table schema has been persisted; finally, release the table distributed lock.
[0057] The procedure flow for other metadata change events is as follows:
[0058] First, if the schema of the table changes, obtain the distributed lock of the table, traverse the metadata to generate schema change events (table creation events, column addition events, column deletion events, column type change events), send the change events and persist the new schema of the table; if the acquisition of the distributed lock of the table times out, that is, another Flink Task Slot obtains the distributed lock, re-read the schema in Zookeeper and re-determine whether the schema has changed; finally, release the distributed lock of the table.
[0059] The embodiment of the present invention further provides a metadata change capture device based on Flink CDC mode evolution, comprising: at least one memory and at least one processor;
[0060] The at least one memory is used to store a machine-readable program;
[0061] The at least one processor is used to call the machine-readable program to implement the metadata change capture method based on Flink CDC mode evolution described in the above embodiment.
[0062] The embodiment of the present invention further provides a computer-readable medium, on which computer instructions are stored, and when the computer instructions are executed by a processor, the metadata change capture method based on the evolution of the Flink CDC mode described in the above embodiment is implemented. Specifically, a system or device equipped with a storage medium can be provided, on which a software program code for implementing the functions of any of the above embodiments is stored, and a computer (or CPU or MPU) of the system or device is enabled to read and execute the program code stored in the storage medium.
[0063] In this case, the program code itself read from the storage medium can realize the function of any one of the above-mentioned embodiments, and thus the program code and the storage medium storing the program code constitute a part of the present invention.
[0064] The storage medium embodiments for providing the program code include a floppy disk, a hard disk, a magneto-optical disk, an optical disk (such as CD-ROM, CD-R, CD-RW, DVD-ROM, DVD-RAM, DVD-RW, DVD+RW), a magnetic tape, a non-volatile memory card, and a ROM. Alternatively, the program code can be downloaded from a server computer by a communication network.
[0065] In addition, it should be clear that the functions of any of the above embodiments can be implemented not only by executing the program code read by the computer, but also by enabling an operating system operating on the computer to complete part or all of the actual operations based on instructions from the program code.
[0066] In addition, it can be understood that the program code read from the storage medium is written to a memory provided in an expansion board inserted into the computer or written to a memory provided in an expansion unit connected to the computer, and then based on the instructions of the program code, a CPU installed on the expansion board or the expansion unit is enabled to perform part or all of the actual operations, thereby realizing the functions of any of the above-mentioned embodiments.
[0067] The present invention is shown and described in detail above through the accompanying drawings and preferred embodiments. However, the present invention is not limited to these disclosed embodiments. Based on the above multiple embodiments, those skilled in the art can know that the code review methods in the above different embodiments can be combined to obtain more embodiments of the present invention, and these embodiments are also within the protection scope of the present invention.
Claims
1. A metadata change capture method based on the evolution of Flink CDC mode, characterized in that: The implementation of this method includes Kafka data source components, Kafka Connect JDBC record deserializer, in-memory event storage, and Zookeeper event storage. The Kafka data source factory creates a Kafka data source component based on the parameters and their values defined in the data source option parameter list; the Kafka data source component is responsible for interacting with the topic specified by the Kafka cluster, reading Kafka message records according to the partition start offset specified by the option parameter, maintaining consumer groups and offsets, and ensuring at least once or exactly once semantics; it is responsible for creating the Kafka Connect JDBC record deserialization component and passing the Kafka message record to it; Creates an in-memory event store or a Zookeeper event store depending on the event store type specified by the options parameter.
2. According to a metadata change capture method based on Flink CDC mode evolution according to claim 1, it is characterized in that: The Kafka Connect JDBC record deserializer generates a tableId by parsing the topic name or the header information in the Kafka message record according to the TableId parsing method specified by the option parameter; Deserialize received Kafka records into business data and metadata schema; pass the metadata schema to the event storage for further processing; Generate data change events based on business data and pass them to downstream operators.
3. According to a metadata change capture method based on Flink CDC mode evolution according to claim 2, it is characterized in that: The TableId includes a database name, a mode name, and a data table name.
4. According to a metadata change capture method based on Flink CDC mode evolution according to claim 2, it is characterized in that: The memory event storage stores metadata schemas, compares newly received schemas with stored schemas, generates metadata change events based on differences, and passes them to the SchemaOperator operator.
5. A metadata change capture method based on Flink CDC mode evolution according to claim 2 or 3, characterized in that: After the Zookeeper event storage is started, the Flink CDC job name specified in the parameter list is created as the root Znode temporary node; the TableId is connected as the Znode temporary node name, and its corresponding metadata schema is deserialized into binary data and saved as a value in the Znode; at the same time, the method of reading the binary data of the Znode node and deserializing it into the metadata schema is supported; the Znode temporary node name lock is used as a table-level distributed lock.
6. According to a metadata change capture method based on Flink CDC mode evolution according to claim 5, it is characterized in that: The program flow of the table creation event is as follows: First, determine whether the table schema has been saved in the Znode temporary node of Zookeeper. If it already exists, record in memory that the table creation event has been sent; otherwise, obtain the table distributed lock and generate the table creation event and send it to the downstream SchemaOperator operator; If the acquisition of the table distributed lock times out, that is, another Flink Task Slot acquires the distributed lock, re-determine whether the table schema has been persisted; Finally, release the distributed lock of the table.
7. According to a metadata change capture method based on Flink CDC mode evolution according to claim 6, it is characterized in that: The procedure flow for other metadata change events is as follows: First, if the schema of the table changes, obtain the distributed lock of the table, traverse the metadata to generate a schema change event, send the change event, and persist the new schema of the table; If the acquisition of the table's distributed lock times out, that is, another Flink Task Slot acquires the distributed lock, re-read the Schema in Zookeeper and re-determine whether the Schema has changed; Finally, release the distributed lock of the table; The schema change events include table creation events, column addition events, column deletion events, and column type change events.
8. A metadata change capture system based on the evolution of Flink CDC mode, characterized in that: Includes Kafka data source components, Kafka Connect JDBC record deserializer, in-memory event storage, and Zookeeper event storage; The Kafka data source component is responsible for creating the Kafka Connect JDBC record deserializer component and passing the Kafka message record to it; Create an in-memory event store or a Zookeeper event store according to the event store type specified by the option parameter; The system specifically implements metadata change capture based on Flink CDC mode evolution through any one of the methods described in claims 1 to 7.
9. A metadata change capture device based on Flink CDC mode evolution, characterized in that: include: at least one memory and at least one processor; The at least one memory is used to store a machine-readable program; The at least one processor is used to call the machine-readable program to implement the method described in any one of claims 1 to 7.
10. A computer-readable medium, characterized in that The computer readable medium stores computer instructions, which, when executed by a processor, can implement the method according to any one of claims 1 to 7.