A method and system for synchronously processing heterogeneous data

Through heterogeneous data synchronization processing method, the data change information of the Oracle database is synchronized to the target database, solving the problem of low correlation query efficiency between heterogeneous databases and achieving efficient data synchronization and query.

CN111984715BActive Publication Date: 2025-05-13SHENZHEN YINSHENG E-PAY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010838910.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-08-19
Publication Date
2025-05-13
Estimated Expiration
2040-08-19

AI Technical Summary

Technical Problem

In the prior art, data synchronization between heterogeneous databases has the problem of low correlation query efficiency, especially in big data scenarios, it is difficult to achieve efficient query between relational databases and non-relational databases.

Method used

Through a heterogeneous data synchronization processing method, the data change information of the source Oracle database's master and slave tables is stored in the queue file, and the transmission process is transmitted to the target system through TCP/IP, and finally processed and saved to the target database by the copy process to achieve heterogeneous synchronization of the data.

Benefits of technology

The data synchronization between the source database and the target database is realized, the data delay in sub-seconds is maintained, and the query efficiency is improved, solving the problem of low correlation query efficiency between heterogeneous databases.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN111984715B_ABST
    Figure CN111984715B_ABST
Patent Text Reader

Abstract

The embodiment of the present invention provides a heterogeneous data synchronization processing method, comprising the following steps: Step 1: storing the data change information corresponding to the main table of the source Oracle database of the first server and the data change information corresponding to the secondary table in a queue file; Step 2: using a transmission process to transmit the queue file of the first server to the target system corresponding to the second server through TCP / IP; Step 3: using a replication process to read the data change information from the queue file of the target system corresponding to the second server; Step 4: processing the data change information through a program specified by a replication process configuration file, and synchronously saving it to the target database of the second server. In the embodiment of the present invention, the data of the source database and the target database are synchronized and the data delay is maintained at the sub-second level, while improving the query efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of big data technology, and more specifically, to a method and system for synchronously processing heterogeneous data. Background Art

[0002] For business systems of different database types, in order to connect to the business, it is often necessary to synchronize the source data. In the big data scenario, the relational database (oracle) has a high I / O, and the relational database will also read the entire row of data from the storage device into the memory. Synchronize the data to another Nosql (mongodb) platform for maintenance, low-latency query, data retrieval and some other needs. Nosql database is horizontally scalable, and its storage is naturally distributed.

[0003] Usually, the query is not just a simple query of the data in a table, but also a related query. If a relational database has several large tables with hundreds of millions of data, the database performance will be greatly affected and the query efficiency will be low. When the synchronized database is a non-relational database, related queries cannot be performed.

[0004] SUMMARY OF THE INVENTION

[0005] In order to overcome the deficiencies of the prior art, the present invention provides a heterogeneous data synchronization processing method to solve the problem that the synchronized source database and target database data cannot be queried in association and the query efficiency is low.

[0006] The technical solution adopted by the present invention to solve the technical problem is: a method for synchronously processing heterogeneous data, comprising the following steps:

[0007] Step 1: Store the data change information corresponding to the source Oracle database master table and the data change information corresponding to the slave table of the first server into a queue file;

[0008] Step 2: using a transmission process to transmit the queue file of the first server to a target system corresponding to the second server via TCP / IP;

[0009] Step 3: Using a replication process, read data change information from the queue file of the target system corresponding to the second server;

[0010] Step 4: Process the data change information through the program specified by the replication process configuration file, and save it synchronously to the target database of the second server.

[0011] Preferably, before storing the data change information corresponding to the master table of the source-end Oracle database of the first server and the data change information corresponding to the slave table in the queue file, the step further includes:

[0012] Obtain a master table and a slave table from the source Oracle database of the first server.

[0013] Preferably, after obtaining the master table and the slave table from the source Oracle database of the first server, the step further includes:

[0014] Use the ogg tool to read the Online Redo Log or Archive Log from the source Oracle database of the first server;

[0015] The read Online Redo Log or Archive Log is parsed to obtain parsed data.

[0016] Specifically, the Online Redo Log or Archive Log is read from the source Oracle database of the first server through the ogg tool, and the steps include:

[0017] The Online RedoLog or Archive Log is read from the source Oracle database of the first server using the ogg tool and the extraction thread.

[0018] Preferably, after parsing the read Online Redo Log or Archive Log to obtain the parsed data, the step further includes:

[0019] Extracting the data change information from the parsed data;

[0020] The extracted data change information is converted into an intermediate format customized by GoldenGate.

[0021] A heterogeneous data synchronization processing system, the system comprising:

[0022] A storage unit, used to store data change information corresponding to a master table of a source-end Oracle database of the first server and data change information corresponding to a slave table into a queue file;

[0023] A transmission unit, used to transmit the queue file of the first server to a target system corresponding to the second server via TCP / IP using a transmission process;

[0024] a reading unit, configured to read data change information from the queue file of the target system corresponding to the second server by using a replication process;

[0025] A saving unit is used to process the data change information through a program specified by a replication process configuration file, and synchronously save the data to a target database of the second server.

[0026] Preferably, the system further comprises:

[0027] An acquisition unit is used to acquire a master table and a slave table from a source Oracle database of the first server.

[0028] Preferably, the system further comprises:

[0029] A reading unit, used to read Online RedoLog or Archive Log from the source Oracle database of the first server through an ogg tool;

[0030] The parsing unit is used to parse the read Online Redo Log or Archive Log to obtain parsed data.

[0031] Specifically, the reading unit includes:

[0032] The reading subunit is used to read the Online Redo Log or Archive Log from the source Oracle database of the first server by using the ogg tool and the extraction thread.

[0033] Preferably, the system further comprises:

[0034] An extraction unit, used for extracting the data change information from the parsed data;

[0035] The conversion unit is used to convert the extracted data change information into an intermediate format customized by GoldenGate.

[0036] The beneficial effect of the present invention is that the data change information of the main table in the source database and the data change information of the slave table are synchronously implemented in the corresponding wide table of the target database through the extraction process, the transmission process and the replication process, thereby realizing data synchronization between the source database and the target database and maintaining sub-second data delay, while improving query efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 The present invention is a flowchart of a heterogeneous data synchronization processing method.

[0038] Figure 2 It is a functional module diagram of a heterogeneous data synchronization processing method.

[0039] Figure 3 It is a diagram showing the principle of heterogeneous data synchronization processing method.

[0040] Figure 4 This is another principle implementation diagram of a heterogeneous data synchronization processing method.

[0041] Figure 5 This is another principle implementation diagram of a heterogeneous data synchronization processing method.

[0042] Figure 6 This is a diagram showing the effect of a method for synchronously processing heterogeneous data. DETAILED DESCRIPTION

[0043] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0044] The specific implementation of the present invention is described in detail below in conjunction with specific embodiments:

[0045] Embodiment 1:

[0046] Figure 1 The implementation process of a heterogeneous data synchronization processing method provided by the first embodiment of the present invention is shown. For the convenience of description, only the part related to the embodiment of the present invention is shown, which is described in detail as follows:

[0047] In step S101: data change information corresponding to the source end Oracle database master table of the first server and data change information corresponding to the slave table are stored in a queue file;

[0048] In an embodiment of the present application, a trail file is pre-constructed, and the source Oracle database of the first server contains hundreds of millions of merchant tables. Any two merchant tables with an associated relationship can be represented as a master table and a slave table. The data change information corresponding to the master table and the data change information corresponding to the slave table are stored in the trail file. The data change information is the field information after the fields in the master table and the slave table are changed, so that the master table and the slave table can save the changed fields in time.

[0049] Preferably, before storing the data change information corresponding to the master table of the source Oracle database of the first server and the data change information corresponding to the slave table in the queue file, the step further includes: obtaining the master table and obtaining the slave table from the source Oracle database of the first server.

[0050] Further preferably, after obtaining the master table and the slave table from the source Oracle database of the first server, the step further includes: reading the Online RedoLog or Archive Log from the source Oracle database of the first server through the ogg tool; parsing the read Online Redo Log or Archive Log to obtain the parsed data. Further preferably, after parsing the read Online Redo Log or Archive Log to obtain the parsed data, the step further includes: extracting the data change information from the parsed data; converting the extracted data change information into a GoldenGate customized intermediate format. Specifically, reading the Online Redo Log or Archive Log from the source Oracle database of the first server through the ogg tool, the step includes: reading the Online Redo Log or Archive Log from the source Oracle database of the first server through the ogg tool using an extraction thread.

[0051] In step S102: using a transmission process to transmit the queue file of the first server to a target system corresponding to the second server via TCP / IP;

[0052] In an embodiment of the present application, the transmission (Pump) process runs on the first server at the source end. When the first server at the source end uses the local queue file, the transmission process will send the queue file in the form of data blocks to the target system corresponding to the second server through the TCP / IP protocol. When the first server at the source end does not use the local queue file, the transmission process enters a dormant state.

[0053] In step S103: using a replication process to read data change information from the queue file of the target system corresponding to the second server;

[0054] In an embodiment of the present application, the replication process runs on the second server on the target side, which is the last stop of data transmission. The second server on the target side is responsible for reading the content in the queue (trail) file of the second server on the target side, parsing the queue file into DML or DDL statements, and then applying them to the target database.

[0055] In step S104: the data change information is processed by a program specified by the replication process configuration file, and is synchronously saved to a target database of the second server.

[0056] In the embodiment of the present application, the second server at the target end processes the data change information through the program specified by the replication process configuration file, such as DML operations - add, delete, and modify operations. The H table replication process can read the queue of the H table and synchronize the H table changes to the HU table. The U table replication process can read the queue of the U table and synchronize the U table changes to the HU table. Furthermore, the data change information of the main table in the source database of the first server and the data change information of the slave table are synchronously landed in the wide table corresponding to the target database of the second server, thereby achieving data synchronization between the source database and the target database and maintaining sub-second data latency, while improving query efficiency.

[0057] A person skilled in the art can understand that all or part of the steps in the above-mentioned embodiment method can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable storage medium, such as ROM / RAM, disk, CD-ROM, etc.

[0058] Embodiment 2:

[0059] Figure 2 The structure of a heterogeneous data synchronization processing system provided by the second embodiment of the present invention is shown. For the convenience of description, only the part related to the embodiment of the present invention is shown, which is described in detail as follows:

[0060] The storage unit 201 is used to store the data change information corresponding to the main table of the source Oracle database of the first server and the data change information corresponding to the secondary table into a queue file;

[0061] A transmitting unit 202, configured to transmit the queue file of the first server to a target system corresponding to the second server via TCP / IP using a transmission process;

[0062] A reading unit 203, configured to read data change information from the queue file of the target system corresponding to the second server by using a replication process;

[0063] The saving unit 204 is used to process the data change information through a program specified by the replication process configuration file, and synchronously save the data to the target database of the second server.

[0064] In an embodiment of the present invention, the data change information corresponding to the main table of the source Oracle database of the first server and the data change information corresponding to the sub-table are stored in a queue file, and the queue file of the first server is transmitted to the target system corresponding to the second server via TCP / IP using a transmission process, and the data change information is read from the queue file of the target system corresponding to the second server using a replication process, and the data change information is processed by a program specified by a replication process configuration file, and is synchronously saved to the target database of the second server. The data change information of the main table in the source database and the data change information of the sub-table are synchronously landed in the wide table corresponding to the target database through an extraction process, a transmission process, and a replication process, thereby achieving data synchronization between the source database and the target database and maintaining sub-second data delays, while improving query efficiency. The specific implementation methods of each unit can refer to the description of Example 1, which will not be repeated here.

[0065] Embodiment three:

[0066] Figure 3 A schematic diagram of a heterogeneous data synchronization processing method provided by Embodiment 3 of the present invention is shown. For ease of explanation, only the part related to the embodiment of the present invention is shown, including:

[0067] 1. Use the Extract Process to read the Online Redo Log or Archive Log from the source ORACLE database using the OGG tool, and then parse the Online Redo Log or Archive Log to convert the extracted change information into a GoldenGate-defined intermediate format and store it in a trail file.

[0068] Second, use the transmission process to transfer the trail file to another server target system via TCP / IP protocol.

[0069] 3. Finally, use the replication process to read the data change information from the trail file, process the data change information through the program specified by the replication process configuration file, and synchronize it to the target database MONGODB on the other end server.

[0070] It can be understood as "monitoring" the changes in ORACLE source data, converting these changed data into queue files (extraction process), forwarding the queue files to other servers (transmission process); and synchronously updating the received queue files to the target database through the "program" (copying process).

[0071] Embodiment 4:

[0072] Figure 4 Another principle implementation diagram of a heterogeneous data synchronization processing method provided by Embodiment 3 of the present invention is shown. For ease of explanation, only the part related to the embodiment of the present invention is shown, including:

[0073] The H table replication process can read the H table queue and synchronize the H table changes to the HU table:

[0074] When type is I, it is the insert queue: when inserting into the main table, determine whether the inserted MERC_ID has a value; if MERC_ID has no value, directly land the change data in the wide table; if MERC_ID has a value, reverse query the slave table field of the synchronized U table according to the MERC_ID value, and land it in the HU table together with the change data.

[0075] When type is U, it is an update queue: the main table is updated, and it is determined whether the update queue contains MERC_ID. The processing method is the same as that of insertion, which will not be repeated here.

[0076] When type is D, it is the delete queue: main table deletion, directly deleting the corresponding data from the wide table based on the primary key.

[0077] Embodiment five:

[0078] Figure 5 Another principle implementation diagram of a heterogeneous data synchronization processing method provided by Embodiment 3 of the present invention is shown. For ease of explanation, only the part related to the embodiment of the present invention is shown, including:

[0079] The U table replication process can read the U table queue and synchronize the U table changes to the HU table:

[0080] When type is I, it is an insert queue: insert from the table, check whether there is a record with MERC_ID of this value in the wide table according to the primary key MERC_ID value; if there is a record with MERC_ID of this value in the wide table, update the corresponding slave table field in the wide table according to the slave table data; if there is no record with MERC_ID of this value in the wide table, no operation is performed.

[0081] When type is U, it is an update queue: update from the table, determine whether the update queue contains the slave table field in the wide table; if the update queue contains the slave table field in the wide table, obtain the value of MERC_ID and update the slave table field of the record whose wide table MERC_ID is this value; if the update queue does not contain the slave table field in the wide table, do not operate.

[0082] When type is D, it is a delete queue: delete from the table, get the MERC_ID value, update the slave table field of the record whose MERC_ID in the wide table is this value, and set it to empty.

[0083] Embodiment six:

[0084] Figure 6 The effect diagram of implementing a heterogeneous data synchronization processing method provided by the third embodiment of the present invention is shown. For the convenience of explanation, only the part related to the embodiment of the present invention is shown, including:

[0085] In the source ORACLE database, obtain the associated H and U tables. Both H and U tables contain the MERC_ID field. The LOG_NO and AC_DT fields in the H table are primary keys, and other fields (including MERC_ID) are secondary keys. The MERC_ID field in the U table is the primary key, and other fields are secondary keys. Through the principle of heterogeneous synchronization, the H and U tables are synchronized and integrated into the target database MONGODB. The data change information of the primary table and the data change information of the secondary table in the source database are synchronized and landed in the corresponding wide table of the target database through the extraction process, transmission process, and replication process, thereby realizing data synchronization between the source database and the target database and maintaining sub-second data latency, while improving query efficiency.

[0086] Those skilled in the art will appreciate that the units and algorithm steps of the embodiments described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware or in a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution.

[0087] Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention. The above is only a specific implementation of the present invention, but the protection scope of the present invention is not limited thereto. Any person familiar with the technical field can easily think of changes or replacements within the technical scope disclosed by the present invention, which should be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be based on the protection scope of the claims.

Claims

1. A method for synchronously processing heterogeneous data, characterized in that: The following steps are involved: Step 1: Store the data change information corresponding to the source Oracle database master table and the data change information corresponding to the slave table of the first server into a queue file; Before storing the data change information corresponding to the main table of the source-end Oracle database of the first server and the data change information corresponding to the secondary table in the queue file, the step further includes: Obtain a master table and a slave table from a source Oracle database of the first server; After obtaining the master table and the slave table from the source Oracle database of the first server, the step further includes: Reading the Online Redo Log or Archive Log from the source Oracle database of the first server through the ogg tool; parsing the read Online Redo Log or Archive Log to obtain parsed data; Reading Online Redo Log or ArchiveLog from the source Oracle database of the first server through the ogg tool, the steps include: Use the ogg tool to use the extraction thread to read the Online Redo Log or Archive Log from the source Oracle database of the first server; After parsing the read Online Redo Log or Archive Log to obtain the parsed data, the step further includes: Extracting the data change information from the parsed data; Convert the extracted data change information into an intermediate format customized by GoldenGate; Step 2: using a transmission process to transmit the queue file of the first server to a target system corresponding to the second server via TCP / IP; Step 3: Using a replication process, read data change information from the queue file of the target system corresponding to the second server; Step 4: Process the data change information through the program specified by the replication process configuration file, and save it synchronously to the target database of the second server.

2. A heterogeneous data synchronization processing system, characterized in that: The system is used to implement a heterogeneous data synchronization processing method as claimed in claim 1, and the system includes: A storage unit, used to store data change information corresponding to a master table of a source-end Oracle database of the first server and data change information corresponding to a slave table into a queue file; A transmission unit, used to transmit the queue file of the first server to a target system corresponding to the second server via TCP / IP using a transmission process; a reading unit, configured to read data change information from the queue file of the target system corresponding to the second server by using a replication process; a saving unit, configured to process the data change information through a program specified by a replication process configuration file, and synchronously save the information to a target database of the second server; The system further comprises: An acquisition unit, configured to acquire a master table and a slave table from a source Oracle database of the first server; A reading unit, used to read the Online Redo Log or Archive Log from the source Oracle database of the first server through the ogg tool; A parsing unit, used for parsing the read Online Redo Log or Archive Log to obtain parsed data; An extraction unit, used for extracting the data change information from the parsed data; A conversion unit, used to convert the extracted data change information into an intermediate format customized by GoldenGate; The reading unit comprises: The reading subunit is used to read the Online Redo Log or Archive Log from the source Oracle database of the first server by using the ogg tool and the extraction thread.

Citation Information

Patent Citations

  • Quasi-real-time synchronizing method and device of data among databases

    CN107943979A

  • Big data fusion method and system for heterogeneous platform, electronic device and storage medium

    CN108763387A