Data processing method, system and equipment based on data synchronization tool and medium
By synchronizing the source data and metadata table structure from the source database to the distributed event streaming platform and then to the data lake, the problem that Oracle GoldenGate cannot support cross-origin Schema Evolution data synchronization is solved, and efficient and stable data synchronization between the source database and the data lake is achieved.
Patent Information
- Application Number
- CN202411953142.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-27
- Publication Date
- 2025-06-03
AI Technical Summary
In the prior art, Oracle GoldenGate cannot support cross-origin Schema Evolution data synchronization into Paimon, and traditional methods rely on manual configuration mapping rules, which is time-consuming and error-prone, making it difficult to meet the real-time and efficient data synchronization needs.
The data synchronization tool is used to obtain the source data and metadata table structure from the source database, synchronize it to the distributed event flow platform, and then synchronize the source data and metadata table structure to the data lake through the distributed computing framework, solving the data synchronization problem under Schema Evolution.
It realizes efficient and stable data synchronization between the source database and the data lake, solves the problem that the existing technology cannot support cross-origin Schema Evolution data synchronization, and meets the needs of real-time and efficient data synchronization.
Smart Images

Figure CN120086281A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of data processing, and particularly to a data processing method, system, device and medium based on a data synchronization tool. Background Art
[0002] With the booming development of big data technology, data synchronization and data warehouse construction have become key links in enterprise informatization construction. Oracle GoldenGate (hereinafter referred to as OGG), as an efficient data synchronization tool, occupies an important position in the field of real-time data synchronization between databases by virtue of its powerful log capture, conversion and delivery capabilities. However, in practical applications, the Schema (table schema) differences and dynamic changes between the source database and the target database, namely Schema Evolution, pose severe challenges to data synchronization, and OGG cannot support cross-source Schema Evolution data synchronization into Paimon (advanced data processing framework).
[0003] Traditional methods often rely on manual configuration of mapping rules, which are not only time-consuming and laborious, but also error-prone and difficult to meet the requirements of real-time and efficient data synchronization. Summary of the Invention
[0004] The technical problem to be solved by the present disclosure is to overcome the defect that OGG in the prior art cannot support cross-source Schema Evolution data synchronization into Paimon, and provide a data processing method, system, device and medium based on a data synchronization tool.
[0005] The present disclosure solves the above technical problem through the following technical solutions:
[0006] The first aspect of the present disclosure provides a data processing method based on a data synchronization tool, and the data processing method includes:
[0007] Obtain source data and a metadata table structure from a source database through a data synchronization tool;
[0008] Synchronize the source data and the metadata table structure to a distributed event stream platform;
[0009] Obtain the source data and the metadata table structure from the distributed event stream platform through a distributed computing framework, and synchronize the source data and the metadata table structure to a data lake.
[0010] Preferably, a listening service is deployed in the data synchronization tool, and the step of obtaining the metadata table structure from the source database through the data synchronization tool includes:
[0011] Obtain source data in the source database through the data synchronization tool;
[0012] Generate a metadata structure file based on the data synchronization tool;
[0013] In response to detecting a change in the metadata structure file through the monitoring service, obtain the metadata table structure corresponding to the metadata structure file from the source database.
[0014] Preferably, the data processing method further includes:
[0015] In response to a change in the table structure in the source database, obtain table structure change information;
[0016] Generate a table structure mapping condition between the source database and the data lake according to the table structure change information;
[0017] Monitor the change of the metadata table structure according to the table structure mapping condition, and synchronize the metadata table structure to the distributed event stream platform;
[0018] Modify the data lake based on the metadata table structure in the source database so that the data lake meets the requirements of the table structure change;
[0019] and / or,
[0020] The data processing method further includes:
[0021] Display the synchronization progress of the source data and the metadata table structure.
[0022] Preferably, the step of synchronizing the metadata table structure to the data lake includes:
[0023] In response to the metadata table structure being a modified table, obtain the modified field names in the metadata table structure;
[0024] Add the modified field names to the data lake.
[0025] Preferably, the step of synchronizing the metadata table structure to the data lake further includes:
[0026] In response to the metadata table structure being a newly added table and there being a synchronized data table in the data lake, obtain the structure fields and types of the data lake synchronized data table;
[0027] Compare the structure fields and obtain the unioned structure fields;
[0028] Add the unioned structure fields to the data lake;
[0029] Or,
[0030] The step of synchronizing the metadata table structure to the data lake further includes:
[0031] In response to the metadata table structure being a newly added table and the synchronized data table not existing in the data lake, creating the data lake synchronized data table;
[0032] Adding the data lake synchronized data table to the data lake.
[0033] Preferably, the step of synchronizing the source data to the data lake includes:
[0034] Parsing the source data and obtaining metadata fields;
[0035] In response to the metadata fields being in the metadata table structure, adding the metadata fields to the data lake;
[0036] Or,
[0037] The step of synchronizing the source data to the data lake further includes:
[0038] In response to the metadata fields not being in the metadata table structure, caching the source data measurement flow output corresponding to the metadata fields;
[0039] In response to the metadata table structure completing the change, adding the source data with the measurement flow output cached to the data lake.
[0040] A second aspect of the present disclosure provides a data processing system based on a data synchronization tool, the data processing system includes:
[0041] A first acquisition module, configured to acquire source data and a metadata table structure from a source database through a data synchronization tool;
[0042] A first synchronization module, configured to synchronize the source data and the metadata table structure to a distributed event stream platform;
[0043] A second synchronization module, configured to acquire the source data and the metadata table structure from the distributed event stream platform through a distributed computing framework, and synchronize the source data and the metadata table structure to the data lake.
[0044] Preferably, a listening service is deployed in the data synchronization tool, and the first acquisition module includes:
[0045] A first acquisition unit, configured to acquire source data in the source database through the data synchronization tool;
[0046] A generation unit, configured to generate a metadata structure file based on the data synchronization tool;
[0047] A second acquisition unit, configured to acquire a metadata table structure corresponding to the metadata structure file from the source database in response to detecting a change in the metadata structure file through the monitoring service.
[0048] Preferably, the data processing system further includes:
[0049] A second acquisition module, configured to acquire table structure change information in response to a change in the table structure in the source database;
[0050] A generation module, configured to generate a table structure mapping condition between the source database and the data lake according to the table structure change information;
[0051] A monitoring module, configured to monitor changes in the metadata table structure according to the table structure mapping condition and synchronize the metadata table structure to the distributed event stream platform;
[0052] A modification module, configured to modify the data lake based on the metadata table structure in the source database so that the data lake meets the table structure change requirements;
[0053] And / or
[0054] The data processing system further includes:
[0055] A display module, configured to display the synchronization progress of the source data and the metadata table structure.
[0056] Preferably, the second synchronization module includes:
[0057] A third acquisition unit, configured to acquire the modified field names in the metadata table structure in response to the metadata table structure being a modified table;
[0058] A first addition unit, configured to add the modified field names to the data lake.
[0059] Preferably, the second synchronization module further includes:
[0060] A fourth acquisition unit, configured to acquire the structure fields and types of the data lake synchronization data table in response to the metadata table structure being a newly added table and there being a synchronization data table in the data lake;
[0061] A fifth acquisition unit, configured to compare the structure fields and acquire the unioned structure fields;
[0062] A second addition unit, configured to add the unioned structure fields to the data lake;
[0063] Or
[0064] The second synchronization module further includes:
[0065] A creation unit, configured to create the data lake synchronization data table in response to the metadata table structure being a newly added table and the synchronization data table not existing in the data lake;
[0066] A third addition unit, configured to add the data lake synchronization data table to the data lake.
[0067] Preferably, the second synchronization module further includes:
[0068] A sixth acquisition unit, configured to parse the source data and acquire metadata fields;
[0069] A fourth addition unit, configured to add the metadata fields to the data lake in response to the metadata fields existing in the metadata table structure;
[0070] Or,
[0071] The second synchronization module further includes:
[0072] A cache unit, configured to cache the source data flow measurement output corresponding to the metadata fields in response to the metadata fields not existing in the metadata table structure;
[0073] A new addition unit, configured to newly add the source data with the cached flow measurement output to the data lake in response to the completion of the change of the metadata table structure.
[0074] A third aspect of the present disclosure provides an electronic device, including a memory, a processor, and a computer program stored on the memory and configured to run on the processor, where when the processor executes the computer program, the data processing method based on the data synchronization tool described in the first aspect is implemented.
[0075] A fourth aspect of the present disclosure provides a computer-readable storage medium, on which a computer program is stored, where when the computer program is executed by a processor, the data processing method based on the data synchronization tool described in the first aspect is implemented.
[0076] A fifth aspect of the present disclosure provides a computer program product, including a computer program, where when the computer program is executed by a processor, the data processing method based on the data synchronization tool as described in the first aspect is implemented.
[0077] On the basis of conforming to common knowledge in the art, the above preferred conditions can be combined arbitrarily to obtain various preferred examples of the present disclosure.
[0078] The positive and progressive effects of the present disclosure are as follows:
[0079] The present disclosure synchronizes the source data and the metadata table structure obtained from the source database through a data synchronization tool to a distributed event stream platform; and then synchronizes the source data and the metadata table structure to a data lake through a distributed computing framework, solving the problem that the data synchronization tool cannot support cross-source Schema Evolution data synchronization into the lake Paimon, and realizing efficient and stable data synchronization between the source database and the data lake. Description of the Drawings
[0080] Figure 1 It is a flowchart of a data processing method based on a data synchronization tool provided in Embodiment 1 of the present disclosure;
[0081] Figure 2 It is a schematic diagram of modules of a data processing system based on a data synchronization tool provided in Embodiment 2 of the present disclosure;
[0082] Figure 3 It is a schematic diagram of the structure of an electronic device for implementing a data processing method based on a data synchronization tool in Embodiment 3 of the present disclosure. Detailed Embodiments
[0083] The present disclosure will be further described below by way of embodiments, but the present disclosure is not limited to the scope of the described embodiments.
[0084] In the embodiments of the present disclosure, prefix words such as "first" and "second" are only used to distinguish different described objects, and have no limiting effect on the position, order, priority, quantity or content of the described objects. The use of ordinal numbers and other prefix words for distinguishing described objects in the embodiments of the present disclosure does not constitute a limitation on the described objects. The description of the described objects refers to the description in the claims or the context of the embodiments, and should not constitute unnecessary limitations due to the use of such prefix words. In addition, in the description of this embodiment, unless otherwise specified, the meaning of "a plurality" is two or more.
[0085] In the embodiments of the present disclosure, the processing of collection, storage, use, processing, transmission, provision and disclosure of user personal information complies with the provisions of relevant laws and regulations and does not violate public order and good customs.
[0086] Embodiment 1
[0087] Figure 1 It is a flowchart of a data processing method based on a data synchronization tool provided in Embodiment 1 of the present disclosure. As Figure 1 shown, the data processing method includes:
[0088] S1. Obtain source data and a metadata table structure from a source database through a data synchronization tool;
[0089] In this embodiment, data in the Oracle (source database) is collected through OGG (data synchronization tool).
[0090] Specifically, during the real-time collection of Oracle data, the following OGG collection and transmission processes need to be configured:
[0091] Configure Oracle to enable archive logs, supplemental logs, and full-column supplemental logs, and enable OGG collection user permission configuration;
[0092] Enable DDL replication permission in the OGG collection extraction process to obtain full-column change records;
[0093] Configure DDL (Data Definition Language) replication permission in the OGG replication process, and complete metadata structure collection according to the full-column metadata change records of the table;
[0094] The OGG collection process does not support writing field types RAW, BLOB, UROWID, BFILE, and does not support updating field types CLOB, NCLOB, LONG;
[0095] S2. Synchronize the source data and the metadata table structure to the distributed event stream platform;
[0096] In this embodiment, the source data and the metadata table structure are synchronized to Kafka (distributed event stream platform);
[0097] In this embodiment, the metadata table structure in Kafka is divided into multiple topics according to the organization code;
[0098] It should be noted that there are cases where the table names of different organizations in the metadata table structure in Kafka are the same, but the metadata table structures are different, and they need to be stored in the same table in Paimon.
[0099] S3. Obtain the source data and the metadata table structure from the distributed event stream platform through the distributed computing framework, and synchronize the source data and the metadata table structure to the data lake.
[0100] In this embodiment, the source data and the metadata table structure are obtained from Kafka through Flink (distributed computing framework), and the source data and the metadata table structure are synchronized to the data lake Paimon.
[0101] In an optional embodiment, a listening service is deployed in the data synchronization tool, and step S1 includes:
[0102] Obtain the source data in the source database through the data synchronization tool;
[0103] Generate a metadata structure file based on the data synchronization tool;
[0104] In response to detecting a change in the metadata structure file through the monitoring service, obtain the metadata table structure corresponding to the metadata structure file from the source database.
[0105] In the specific implementation process, deploy the OGG source and target services to the server, establish a connection with the business Oracle database network, create a collection user, and set the permissions for the collection user. Deploy the Java monitoring service program on the OGG target server, start the monitoring service program to monitor changes in the metadata structure file, and create a Kafka metadata Topic; configure the OGG collection and replication process parameters, create a data transmission Topic, start the OGG collection of data, obtain the log data of the source database table, and analyze and extract the data from the log data; start the Flink application program to obtain the Paimon metadata, listen to the above metadata through the Topic to obtain the collected data and write it into Paimon; specifically, based on OGG, implement the real-time Schema Evolution synchronization of Oracle (Oracle version 11.2.0.4 and above) data and write it into the Paimon data lake. The main process is as follows: OGG collects Oracle data, deploys a Java monitoring program on the OGG target server to monitor the Oracle metadata and synchronize it to Kafka, and performs real-time Schema Evolution writing into Paimon through the Flink program. Detailed process: Based on OGG, configure the collection process to capture the source data in the Oracle database in real time, generate the corresponding metadata structure file according to the data synchronization tool, monitor the metadata structure file at the OGG target end through the Java monitoring service program, obtain the metadata table structure corresponding to the metadata structure file from Oracle through the monitoring service program according to the changed metadata structure file and send it to Kafka, the Flink program listens to the Kafka metadata distribution Topic, obtains the table metadata information, creates a target table in Paimon, and dynamically monitors the changes in the metadata table structure in real time and sends them to Kafka. It should be noted that currently, changes to the table support modifications to the field type and length. If there are modifications to the field name and table name at the source Oracle end, new fields and new tables are created in the Paimon target table, so as to complete the preliminary work of Schema Evolution for data synchronization. In the data synchronization process, the Flink program obtains the metadata in the target table of Paimon, starts the OGG synchronization component, performs Oracle data synchronization, sends the data to Kafka, obtains the data, judges the changes in the metadata table structure according to the metadata, and writes the data into the target table of Paimon.
[0106] In an optional embodiment, the data processing method further includes:
[0107] Upon detecting a change in the table structure of the source database, obtain the table structure change information;
[0108] Generate table structure mapping conditions between the source database and the data lake based on the table structure change information;
[0109] Monitor changes to the metadata table structure according to the table structure mapping conditions and synchronize the metadata table structure to the distributed event stream platform;
[0110] Modify the data lake based on the metadata table structure in the source database to make the data lake meet the table structure change requirements;
[0111] In this embodiment, during the data synchronization process, when the Schema of the source database changes, the present disclosure processes Schema Evolution using the following steps;
[0112] (1) Capture Schema changes: The OGG synchronization component captures the change information of the Schema by listening to the DDL (Data Definition Language) operations of the source database and generates a local file. The Java application listens to the local file to implement Schema change monitoring and obtains the corresponding metadata table structure through Oracle and distributes it to Kafka;
[0113] (2) Generate Schema mapping conditions: According to the captured Schema change information, automatically generate Schema mapping conditions between the source database and the data lake (Paimon), and create a synchronized data table for Paimon (for example, the target table);
[0114] (3) Apply Schema mapping conditions: During the data synchronization process, according to the generated Schema mapping conditions, monitor the changed metadata table structure and distribute it to Kafka, and modify the Paimon data lake through the metadata table structure to make it meet the Schema change requirements, thereby completing the change of the metadata table structure;
[0115] (4) Target Schema change conditions: When the source - side Oracle table name changes, it is a new table on the target - side Paimon. When the source - side Oracle table field name changes, it is a new field on the target - side Paimon. When the source - side Oracle field type and length change, it is to update the field type and length on the target - side Paimon.
[0116] In an alternative embodiment, the data processing method further includes:
[0117] Display the synchronization progress of the source data and the metadata table structure.
[0118] In this embodiment, the entire data synchronization process is monitored in real time, and the progress of data synchronization can also be displayed in real time; the received data volume and written data volume are statistically analyzed in real time, which facilitates users to troubleshoot faults and optimize performance.
[0119] In an alternative embodiment, the steps of synchronizing the metadata table structure to the data lake in step S3 include:
[0120] In response to the metadata table structure being a modified table, obtain the modified field names in the metadata table structure;
[0121] Add the modified field names to the data lake.
[0122] In an alternative embodiment, the steps of synchronizing the metadata table structure to the data lake in step S3 further include:
[0123] In response to the metadata table structure being a newly added table and there being a synchronized data table in the data lake, obtain the structure fields and types of the data lake synchronized data table;
[0124] Compare the structure fields and obtain the unioned structure fields;
[0125] Add the unioned structure fields to the data lake;
[0126] In an alternative embodiment, the steps of synchronizing the metadata table structure to the data lake in step S3 further include:
[0127] In response to the metadata table structure being a newly added table and there being no synchronized data table in the data lake, create a data lake synchronized data table;
[0128] Add the data lake synchronized data table to the data lake.
[0129] In an alternative embodiment, the steps of synchronizing the metadata table structure to the data lake in step S3 further include:
[0130] Parse the source data and obtain the metadata fields;
[0131] In response to the metadata fields being in the metadata table structure, add the metadata fields to the data lake;
[0132] In an alternative embodiment, the steps of synchronizing the metadata table structure to the data lake in step S3 further include:
[0133] In response to the metadata fields not being in the metadata table structure, cache the source data flow output corresponding to the metadata fields;
[0134] In this embodiment, if the metadata fields do not exist in the metadata table structure, the metadata fields are filtered and the corresponding source data is written at the same time.
[0135] In response to the completion of the change in the metadata table structure, the source data for caching the flow measurement output is newly added to the data lake.
[0136] In this embodiment, a trigger mechanism is set up to write the data into Paimon according to the primary key update.
[0137] In this embodiment, during the data synchronization process, the present disclosure writes the data into Paimon in the following manner:
[0138] a. Obtain the metadata in the synchronization data table corresponding to Paimon and write the data according to the metadata;
[0139] b. Data writing into Paimon currently supports INSERT, UPDATE, and UPSERT for primary key tables, does not support DELETE operations, and only supports INSERT operations for non-primary key tables;
[0140] c. Listen to Kafka through Flink to obtain data (for example, obtain source data and metadata table structure), distinguish between writing data and updating data based on the value of the "op_type" field in the data (I for writing data, U for updating data), and write the data through metadata comparison. If there is no change in the metadata table structure, the data is written into Paimon in the lake. If there is a change in the metadata table structure, the existing fields are written into the target table in Paimon, and the flow measurement output of the source data corresponding to the field is cached;
[0141] d. After the change in the metadata table structure is completed, start the flow measurement program through Flink to update the corresponding cached data into the target table to complete the final consistency of the data.
[0142] In this embodiment, Oracle (source database) is the starting point of data synchronization. The OGG synchronization component is responsible for capturing the changed data in the source database, performing conversion and delivery. The Java listening service program listens to the change in the metadata table structure at the OGG target end and issues the metadata table structure. Kafka is responsible for receiving and storing the synchronized data (for example, the synchronized source data and metadata table structure). Flink is responsible for Schema Evolution writing into Paimon, and the monitoring module is responsible for counting the data volume of the monitoring data flow link of the entire system.
[0143] It should be noted that the current overall process supports writing and updating data in the Paimon primary key table, does not support data deletion operations, and non-primary key tables only support data writing operations. When processing data writing involving real-time metadata table structure changes, there may be a situation where data arrives at Kafka before the metadata table structure. To solve this problem, the system first compares the arriving data with the metadata of the target table in Paimon, writes the data with existing fields in the target table to the target table, and outputs the data for flow measurement. After successfully obtaining the metadata and modifying the metadata table structure, the system will start a make-up task to update the flow measurement cache data to the target table.
[0144] In addition, during the entire process of data being fetched from Kafka and written to Paimon, the system also establishes a data flow monitoring mechanism to monitor metrics such as the amount of data received in Kafka and the amount of data written to Paimon. Users can view relevant task metrics through the Web interface of the Flink task, ensuring the smooth progress of the data synchronization and Schema Evolution process.
[0145] In this embodiment, Schema Evolution involves modifications to the database structure, such as adding or deleting tables, adding or deleting fields, and adjusting data types, etc. These changes require the data synchronization tool to be able to adapt flexibly, ensuring the accurate and consistent transmission of data between the source and target ends.
[0146] As an emerging streaming data lake storage technology, Apache Paimon, with its high throughput and low latency data ingestion capabilities, as well as rich streaming subscription and real-time query functions, has become an ideal choice for building modern data warehouses. However, combining OGG with Paimon to achieve efficient data synchronization under Schema Evolution.
[0147] Apache Kafka is an open-source distributed event streaming platform. With its high throughput and distributed architecture, it ensures the stable operation of the system under high concurrency. Persistent storage guarantees the security or reliability of messages. Partitioning and parallel processing enhance the message processing ability. It supports message publish / subscribe and can be flexibly applied to various data stream processing scenarios. Thus, it serves as the middle layer for OGG data to be stored in Paimon and can be deeply integrated with the parallelism mechanism of Flink, improving the data reading and writing efficiency. The reliability guarantee of Kafka and the checkpoint mechanism of Flink cooperate to ensure the consistency and reliability of data during the processing process.
[0148] In this embodiment, the source data and the metadata table structure obtained from the source database through the data synchronization tool are synchronized to the distributed event stream platform; then, through the distributed computing framework, the source data and the metadata table structure are synchronized to the data lake, which solves the problem that the data synchronization tool cannot support cross-source Schema Evolution data synchronization into the lake Paimon, and realizes efficient and stable data synchronization between the source database and the data lake. It has important practical significance and application value for improving the enterprise's data synchronization ability and accelerating the construction of the data warehouse, and meets the urgent need of the enterprise for efficient and reliable data synchronization technology.
[0149] Embodiment 2
[0150] Corresponding to the foregoing embodiment of the data processing method based on the data synchronization tool, the present disclosure also provides an embodiment of a data processing system based on the data synchronization tool.
[0151] Figure 2 It is a schematic diagram of the modules of a data processing system based on the data synchronization tool provided in Embodiment 2 of the present disclosure. As Figure 2 shown, the data processing system includes:
[0152] A first acquisition module 21, configured to obtain source data and a metadata table structure from a source database through a data synchronization tool;
[0153] In this embodiment, data in Oracle (source database) is collected through OGG (data synchronization tool);
[0154] Specifically, during the real-time data collection process of Oracle, the following OGG collection and transmission processes need to be configured:
[0155] Oracle is configured to enable archive logging, supplemental logging, and full-column supplemental logging, and enable OGG collection user permission configuration;
[0156] In the OGG collection extraction process configuration, enable DDL replication permission to obtain full-column change records;
[0157] In the OGG replication process, configure DDL (Data Definition Language) replication permission at the same time, and complete metadata structure collection according to the full-column metadata change records of the table;
[0158] The collection process OGG does not support writing field types RAW, BLOB, UROWID, BFILE, and does not support updating field types CLOB, NCLOB, LONG;
[0159] A first synchronization module 22, configured to synchronize the source data and the metadata table structure to the distributed event stream platform;
[0160] In this embodiment, the source data and the metadata table structure are synchronized to Kafka (a distributed event streaming platform).
[0161] In this embodiment, the metadata table structure in Kafka is partitioned into multiple topics according to the organizational unit code.
[0162] It should be noted that in Kafka, there are cases where the table names of different organizations are the same, but the metadata table structures are different, and they need to be stored in the same table in Paimon.
[0163] The second synchronization module 23 is configured to obtain the source data and the metadata table structure from the distributed event streaming platform through a distributed computing framework, and synchronize the source data and the metadata table structure to the data lake.
[0164] In this embodiment, the source data and the metadata table structure are obtained from Kafka through Flink (a distributed computing framework), and the source data and the metadata table structure are synchronized to the data lake Paimon.
[0165] In an alternative embodiment, a listening service is deployed in the data synchronization tool. The first acquisition module includes:
[0166] The first acquisition unit is configured to obtain the source data in the source database through the data synchronization tool.
[0167] The generation unit is configured to generate a metadata structure file based on the data synchronization tool.
[0168] The second acquisition unit is configured to obtain the metadata table structure corresponding to the metadata structure file from the source database in response to detecting a change in the metadata structure file through the listening service.
[0169] In the specific implementation process, deploy the OGG source and target services to the server, establish a connection with the business Oracle database network, create a collection user, and set the permissions for the collection user. Deploy a Java listening service program on the OGG target server, start the listening service program to monitor changes in the metadata structure file, and create a Kafka metadata Topic (theme); configure the OGG collection and replication process parameters, create a data transmission Topic, start the OGG data collection, obtain the log data of the source database table, and analyze and extract the data from the log data; start the Flink application program to obtain the Paimon metadata, listen to the above metadata through the Topic to obtain the collected data and write it into Paimon; specifically, based on OGG, realize the real-time Schema Evolution synchronization of Oracle (Oracle version 11.2.0.4 and above) data and write it into the Paimon data lake. The main process is as follows: OGG collects Oracle data, deploys a Java listening program on the OGG target server to monitor the Oracle metadata and synchronize it to Kafka, and performs real-time Schema Evolution writing into Paimon through the Flink program. Detailed process: Based on OGG, configure the collection process to capture the source data in the Oracle database in real time, generate the corresponding metadata structure file according to the data synchronization tool, monitor the OGG target-side metadata structure file through the Java listening service program, obtain the metadata table structure corresponding to the metadata structure file from Oracle according to the changed metadata structure file through the listening service program and send it to Kafka, the Flink program listens to the Kafka metadata distribution Topic, obtains the table metadata information, creates a target table in Paimon, and listens to the changes in the metadata table structure in real time and sends them to Kafka. It should be noted that currently, table changes support modifications to field types and lengths. For example, if there are modifications to the source-side Oracle field names and table names, new fields and new tables are created in the Paimon target-side table, so as to complete the data synchronization pre-work through Schema Evolution. In the data synchronization process, the Flink program obtains the metadata in the Paimon target table, starts the OGG synchronization component, performs Oracle data synchronization, sends the data to Kafka, obtains the metadata, judges the changes in the metadata table structure according to the metadata, and writes the metadata into the Paimon target table.
[0170] In an optional embodiment, the data processing system further includes:
[0171] A second acquisition module, configured to acquire table structure change information in response to a change in the table structure in the source database;
[0172] A generation module, configured to generate table structure mapping conditions between a source database and a data lake according to table structure change information;
[0173] A monitoring module, configured to monitor changes in the metadata table structure according to the table structure mapping conditions and synchronize the metadata table structure to a distributed event stream platform;
[0174] A modification module, configured to modify the data lake based on the metadata table structure in the source database so that the data lake meets the table structure change requirements;
[0175] In this embodiment, during the data synchronization process, when the Schema of the source database changes, the present disclosure processes Schema Evolution by the following steps;
[0176] (1) Capture Schema changes: The OGG synchronization component captures the change information of the Schema by listening to the DDL (Data Definition Language) operations of the source database and generates a local file. The Java application listens to the local file to implement Schema change monitoring and obtains the corresponding metadata table structure through Oracle and issues it to Kafka;
[0177] (2) Generate Schema mapping conditions: According to the captured Schema change information, automatically generate Schema mapping conditions between the source database and the data lake (Paimon), and create a synchronization data table of Paimon (for example, the target table);
[0178] (3) Apply Schema mapping conditions: During the data synchronization process, according to the generated Schema mapping conditions, monitor the changed metadata table structure and issue it to Kafka, and modify the Paimon data lake through the metadata table structure to make it meet the Schema change requirements, thereby completing the change of the metadata table structure;
[0179] (4) Target Schema change conditions: When the source - side Oracle table name changes, it is a new table at the target - side Paimon. When the source - side Oracle table field name changes, it is a new field at the target - side Paimon. When the source - side Oracle field type and length change, it is a new field type and length at the target - side Paimon.
[0180] In an optional embodiment, the data processing system further includes:
[0181] A display module, configured to display the synchronization progress of the source data and the metadata table structure.
[0182] In this embodiment, the entire data synchronization process is monitored in real time, and the progress of data synchronization can also be displayed in real time; the received data volume and written data volume are statistically analyzed in real time, which facilitates users to troubleshoot faults and optimize performance.
[0183] In an optional embodiment, the second synchronization module includes:
[0184] A third acquisition unit, configured to acquire the modified field names in the metadata table structure in response to the metadata table structure being a modified table;
[0185] A first addition unit, configured to add the modified field names to the data lake.
[0186] In an optional embodiment, the second synchronization module further includes:
[0187] A fourth acquisition unit, configured to acquire the structure fields and types of the data lake synchronization data table in response to the metadata table structure being a new table and the data lake having a synchronization data table;
[0188] A fifth acquisition unit, configured to compare the structure fields and acquire the structure fields after taking the union;
[0189] A second addition unit, configured to add the structure fields after taking the union to the data lake;
[0190] In an optional embodiment, the second synchronization module further includes:
[0191] A creation unit, configured to create a data lake synchronization data table in response to the metadata table structure being a new table and the data lake not having a synchronization data table;
[0192] A third addition unit, configured to add the data lake synchronization data table to the data lake.
[0193] In an optional embodiment, the second synchronization module further includes:
[0194] A sixth acquisition unit, configured to parse the source data and acquire the metadata fields;
[0195] A fourth addition unit, configured to add the metadata fields to the data lake in response to the metadata fields being in the metadata table structure;
[0196] In an optional embodiment, the second synchronization module further includes:
[0197] A cache unit, configured to cache the source data flow output corresponding to the metadata fields in response to the metadata fields not being in the metadata table structure;
[0198] In this embodiment, if the metadata fields do not exist in the metadata table structure, the metadata fields are filtered, and the corresponding source data is written at the same time.
[0199] A new unit is used to add the source data for caching the flow measurement output to the data lake in response to the completion of the change in the metadata table structure.
[0200] In this embodiment, a trigger mechanism is set up to write the data into Paimon according to the primary key update.
[0201] In this embodiment, during the data synchronization process, the present disclosure writes the data into Paimon in the following manner:
[0202] a. Obtain the metadata in the synchronization data table corresponding to Paimon and write the data according to the metadata;
[0203] b. The data writing into Paimon currently supports INSERT, UPDATE, and UPSERT for the primary key table, does not support the DELETE operation, and only supports the INSERT operation for the table without a primary key;
[0204] c. Listen to Kafka through Flink to obtain data (for example, obtain the source data and the metadata table structure), distinguish between writing data and updating data according to the value of the "op_type" field in the data (I represents writing data, U represents updating data), write the data through metadata comparison. If there is no change in the metadata table structure, the data is written into Paimon in the lake. If there is a change in the metadata table structure, write the existing fields into the target table of Paimon and cache the flow measurement output of the source data corresponding to the field;
[0205] d. After the change in the metadata table structure is completed, start the flow measurement program through Flink to update the corresponding cached data into the target table to complete the final consistency of the data.
[0206] In this embodiment, Oracle (source database) is the starting point of data synchronization. The OGG synchronization component is responsible for capturing the changed data in the source database, converting and delivering it. The Java listening service program listens to the change in the metadata table structure at the OGG target end and issues the metadata table structure. Kafka is responsible for receiving and storing the synchronized data (for example, the synchronized source data and the metadata table structure). Flink is responsible for Schema Evolution writing into Paimon, and the monitoring module is responsible for counting the data volume of the monitoring data flow link of the entire system.
[0207] It should be noted that currently the overall process supports writing and updating data in the Paimon primary key table, but does not support data deletion operations. For non-primary key tables, only data writing operations are supported. When dealing with data writing involving real-time metadata table structure changes, there may be a situation where data arrives at Kafka before the metadata table structure. To solve this problem, the system first compares the arriving data with the metadata of the target table in Paimon, writes the data with existing fields in the target table into the target table, and outputs the data for flow measurement. After successfully obtaining the metadata and modifying the metadata table structure, the system will start a make-up task to update the flow measurement cache data to the target table.
[0208] In addition, during the entire process of data being fetched from Kafka and written to Paimon, the system has also established a data flow monitoring mechanism to monitor metrics such as the amount of data received in Kafka and the amount of data written to Paimon. Users can view relevant task metrics through the web interface of the Flink task, ensuring the smooth progress of the data synchronization and Schema Evolution process.
[0209] In this embodiment, Schema Evolution involves modifications to the database structure, such as adding or deleting tables, adding or deleting fields, and adjusting data types, etc. These changes require the data synchronization tool to be able to adapt flexibly, ensuring accurate and consistent transmission of data between the source and target ends.
[0210] As an emerging streaming data lake storage technology, Apache Paimon, with its high throughput and low latency data ingestion capabilities, as well as rich streaming subscription and real-time query functions, has become an ideal choice for building modern data warehouses. However, combining OGG with Paimon to achieve efficient data synchronization under Schema Evolution.
[0211] Apache Kafka is an open-source distributed event streaming platform. With its high throughput and distributed architecture, it ensures the stable operation of the system under high concurrency. Persistent storage guarantees the security or reliability of messages. Partitioning and parallel processing improve the message processing ability. It supports message publishing / subscribing and can be flexibly applied to various data stream processing scenarios. It serves as the middle layer for OGG data to be stored in Paimon and can be deeply integrated with the parallelism mechanism of Flink, improving the data reading and writing efficiency. The reliability guarantee of Kafka and the checkpoint mechanism of Flink cooperate to ensure the consistency and reliability of data during the processing process.
[0212] In this embodiment, the source data and the metadata table structure obtained from the source database through the data synchronization tool are synchronized to the distributed event stream platform; then, through the distributed computing framework, the source data and the metadata table structure are synchronized to the data lake, solving the problem that the data synchronization tool cannot support cross-source Schema Evolution data synchronization into the lake Paimon, and realizing efficient and stable data synchronization between the source database and the data lake. It has important practical significance and application value for improving the enterprise's data synchronization ability and accelerating the construction of the data warehouse, and meets the urgent need of the enterprise for efficient and reliable data synchronization technology.
[0213] For the system embodiment, since it basically corresponds to the method embodiment, the relevant parts can be referred to the partial description of the method embodiment. The system embodiments described above are only illustrative. The units described as separate components may or may not be physically separated. The components as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of the present disclosure.
[0214] Embodiment 3
[0215] Figure 3 FIG. 10 is a schematic structural diagram of an electronic device shown in Embodiment 3 of the present disclosure. The electronic device includes a memory, a processor, and a computer program stored on the memory and configured to run on the processor. When the processor executes the computer program, it implements the data processing method based on the data synchronization tool described in any of the above embodiments. Figure 3 The electronic device 90 shown is only an example and should not impose any limitation on the functions and usage scope of the embodiments of the present disclosure.
[0216] As Figure 3 shown, the electronic device 90 may be presented in the form of a general computing device, for example, it may be a server device. The components of the electronic device 90 may include, but are not limited to: at least one of the above processors 91, at least one of the above memories 92, and a bus 93 connecting different system components (including the memory 92 and the processor 91).
[0217] The bus 93 includes a data bus, an address bus, and a control bus.
[0218] The memory 92 may include volatile memory, such as a random access memory (RAM) 921 and / or a cache memory 922, and may further include a read-only memory (ROM) 923.
[0219] The memory 92 may also include program utilities 925 (or utilities) having a set (at least one) of program modules 924, such program modules 924 including but not limited to: an operating system, one or more application programs, other program modules, and program data, and the implementation of a network environment may be included in each or some combination of these examples.
[0220] The processor 91 executes various functional applications and data processing by running computer programs stored in the memory 92, such as the data processing method based on the data synchronization tool provided in any of the above embodiments.
[0221] The electronic device 90 may also communicate with one or more external devices 94 (such as a keyboard, a pointing device, etc.). Such communication may be carried out through the input / output (I / O) interface 95. And, the electronic device 90 may also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through the network adapter 96. As Figure 3 shown, the network adapter 96 communicates with other modules of the electronic device 90 through the bus 93. It should be understood that although not shown in the figure, other hardware and / or software modules may be used in combination with the electronic device 90, including but not limited to: microcode, device drivers, redundant processors, external disk drive arrays, RAID (disk array) systems, tape drives, and data backup storage systems, etc.
[0222] It should be noted that although several units / modules or sub-units / modules of the electronic device are mentioned in the above detailed description, this division is merely exemplary and not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more of the above-described units / modules may be embodied in one unit / module. Conversely, the features and functions of one unit / module described above may be further divided and embodied by multiple units / modules.
[0223] Embodiment 4
[0224] Embodiment 4 of the present disclosure also provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the data processing method based on the data synchronization tool provided in any of the above embodiments.
[0225] Among them, the readable storage medium may more specifically include but not limited to: a portable disk, a hard disk, a random access memory, a read-only memory, an erasable programmable read-only memory, an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0226] Embodiment 5
[0227] Embodiment 5 of the present disclosure further provides a computer program product, including a computer program, which implements the data processing method based on a data synchronization tool described in any one of the above when executed by a processor.
[0228] Among them, the program code for executing the computer program product of the present disclosure can be written in any combination of one or more programming languages. The program code can be executed entirely on a user device, partially on a user device, executed as an independent software package, partially on a user device and partially on a remote device, or entirely on a remote device.
[0229] Among them, the program code for executing the computer program product of the present disclosure can be written in any combination of one or more programming languages. The program code can be executed entirely on a user device, partially on a user device, executed as an independent software package, partially on a user device and partially on a remote device, or entirely on a remote device.
[0230] Although the specific embodiments of the present disclosure have been described above, those skilled in the art should understand that this is only an example. The protection scope of the present disclosure is defined by the appended claims. Without departing from the principles and essence of the present disclosure, those skilled in the art can make various changes or modifications to these embodiments, but these changes and modifications all fall within the protection scope of the present disclosure.
Claims
1. A data processing method based on a data synchronization tool, characterized in that: The data processing method comprises: Use data synchronization tools to obtain source data and metadata table structures from the source database; Synchronizing the source data and the metadata table structure to a distributed event stream platform; The source data and the metadata table structure are obtained from a distributed event stream platform through a distributed computing framework, and the source data and the metadata table structure are synchronized to a data lake.
2. The data processing method based on the data synchronization tool according to claim 1, characterized in that: A monitoring service is deployed in the data synchronization tool. The step of obtaining the metadata table structure from the source database through the data synchronization tool includes: Acquire source data in the source database through the data synchronization tool; Generate a metadata structure file based on the data synchronization tool; In response to detecting, through the monitoring service, that the metadata structure file has been changed, obtaining a metadata table structure corresponding to the metadata structure file from the source database.
3. The data processing method based on the data synchronization tool according to claim 1, characterized in that: The data processing method further includes: In response to a change in the table structure in the source database, obtaining table structure change information; Generate a table structure mapping condition between the source database and the data lake according to the table structure change information; Monitor changes to the metadata table structure according to the table structure mapping conditions, and synchronize the metadata table structure to the distributed event stream platform; Modify the data lake based on the metadata table structure in the source database so that the data lake meets the table structure change requirement; and / or, The data processing method further includes: The synchronization progress of the source data and the metadata table structure is displayed.
4. The data processing method based on the data synchronization tool according to claim 1, characterized in that: The step of synchronizing the metadata table structure into the data lake includes: In response to the metadata table structure being a modification table, obtaining a modification field name in the metadata table structure; The modified field name is added to the data lake.
5. The data processing method based on the data synchronization tool according to claim 1, characterized in that: The step of synchronizing the metadata table structure into the data lake also includes: In response to the metadata table structure being a newly added table and the synchronized data table existing in the data lake, obtaining the structure fields and types of the synchronized data table in the data lake; Compare the structure fields and obtain the combined structure fields; Adding the unioned structure fields to the data lake; or, The step of synchronizing the metadata table structure into the data lake also includes: In response to the metadata table structure being a newly added table and the synchronization data table not existing in the data lake, creating the data lake synchronization data table; Add the data lake synchronization data table to the data lake.
6. The data processing method based on the data synchronization tool according to claim 1, characterized in that: The step of synchronizing the source data into the data lake includes: Parsing the source data and obtaining metadata fields; In response to the metadata field being in the metadata table structure, adding the metadata field to the data lake; or, The step of synchronizing the source data into the data lake further includes: In response to the metadata field not being in the metadata table structure, caching the source data stream output corresponding to the metadata field; In response to the metadata table structure being changed, source data cached by the flow measurement output is added to the data lake.
7. A data processing system based on a data synchronization tool, characterized in that: The data processing system comprises: A first acquisition module is used to acquire source data and metadata table structure from a source database through a data synchronization tool; A first synchronization module, used to synchronize the source data and the metadata table structure to a distributed event stream platform; The second synchronization module is used to obtain the source data and the metadata table structure from the distributed event stream platform through a distributed computing framework, and synchronize the source data and the metadata table structure to the data lake.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and used to run on the processor, characterized in that: When the processor executes the computer program, the data processing method based on the data synchronization tool according to any one of claims 1 to 6 is implemented.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the data processing method based on a data synchronization tool according to any one of claims 1 to 6 is implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the data processing method based on a data synchronization tool as described in any one of claims 1 to 6 is implemented.