Synchronization Method and Device for Incremental Data, Electronic Device, and Storage Medium
By acquiring and processing operation records of incremental data within the target business system, the ordering and synchronization of incremental data is achieved using the streaming processing engine and the second database, the problem of inaccurate incremental data synchronization in multiple data sources scenarios is solved, and incremental data synchronization with accurate order is achieved.
Patent Information
- Application Number
- CN202111229320.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-21
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2041-10-21
AI Technical Summary
In scenarios with multiple data sources, multiple tables, and large data volumes, the message queue cannot guarantee the global order of incremental data synchronization, resulting in deviations in data at the source and target ends.
By obtaining operation records corresponding to the incremental data generated when updated within the target service system, the operation records classification is sent to the topic partition of the distributed message queue by using the streaming processing engine, these operation records are extracted and consumed, written to the second database to generate the operation record sorting queue, and sorted according to the timestamp, and finally the sorted incremental data is sent to the target end.
The order accuracy of incremental data synchronization from the source end to the target end is achieved, and the problem of inaccurate incremental data synchronization is solved.
Smart Images

Figure CN114048217B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing, and in particular, to a method and apparatus for synchronizing incremental data, an electronic device, and a storage medium. Background Art
[0002] Currently, there are some scenarios where the data stored in the source database and the target database need to be synchronized. The common data synchronization method is the incremental data synchronization of the database. Incremental data, as the name implies, refers to newly added data. Among them, in the incremental synchronization of a relational database, it is generally required to transmit the operation records of the source database to the target end, and a message queue is used to achieve the peak shaving and valley filling of the operation record data to ensure the throughput.
[0003] However, in the scenario of multiple data sources, multiple tables, and large amounts of data, the message queue needs to enable the multi-partition feature. The speed at which consumers read messages from different partitions is inconsistent, which will result in inconsistent order of the database operation records transmitted to the target end. For example, if the operation records are in the order of first delete and then insert, and the order is reversed, the data inserted later will be lost, and there will be a deviation between the data at the source end and the target end. Therefore, the consumption of the message queue cannot guarantee the global order of messages, resulting in inaccurate incremental synchronization of the target end data.
[0004] Therefore, there is a problem in the related technology that the incremental data synchronization between the source end data and the target end data is inaccurate. Summary of the Invention
[0005] This application provides a method and apparatus for synchronizing incremental data, an electronic device, and a storage medium, so as to at least solve the problem that the incremental data synchronization between the source end data and the target end data is inaccurate in the related technology.
[0006] According to one aspect of the embodiments of the present application, a method for synchronizing incremental data is provided. The method includes:
[0007] When it is determined that the target sub-business in the target business system at the source end is updated, obtain the operation records corresponding to the incremental data generated at the time of the update, where the target business system includes multiple sub-businesses, and the operation records are used to represent multiple associated information corresponding to the generation of the incremental data, and the associated information includes a timestamp;
[0008] Send the operation records to each topic partition of the distributed message queue according to the classification of the associated information;
[0009] The streaming processing engine is used to extract and consume the operation records in the subject partition, and write the consumed operation records into a second database to obtain an operation record sorting queue. The second database stores the consumed operation records in a preset format, and the operation record sorting queue is a queue generated by sorting the consumed operation records according to the chronological order of the timestamps.
[0010] The streaming processing engine is used to read the operation record sorting queue stored in the second database, and send the incremental data corresponding to the operation record sorting queue to the target end, so that the target end completes the synchronization of the incremental data.
[0011] According to another aspect of the embodiments of the present application, there is also provided a device for synchronizing incremental data. The device includes:
[0012] A first acquisition unit, configured to acquire operation records corresponding to incremental data generated when a target sub-service in a target business system at a source end is updated. The target business system includes multiple sub-services, and the operation records are used to represent multiple associated information corresponding to the generation of the incremental data. The associated information includes a timestamp.
[0013] A first sending unit, configured to send the operation records to respective topic partitions of a distributed message queue according to the classification of the associated information.
[0014] A obtaining unit, configured to use a streaming processing engine to extract and consume the operation records in the subject partition, and write the consumed operation records into a second database to obtain an operation record sorting queue. The second database stores the consumed operation records in a preset format, and the operation record sorting queue is a queue generated by sorting the consumed operation records according to the chronological order of the timestamps.
[0015] A second sending unit, configured to use the streaming processing engine to read the operation record sorting queue stored in the second database, and send the incremental data corresponding to the operation record sorting queue to the target end, so that the target end completes the synchronization of the incremental data.
[0016] According to yet another aspect of the embodiments of the present application, there is also provided an electronic device, including a processor, a communication interface, a memory, and a communication bus. The processor, the communication interface, and the memory communicate with each other through the communication bus. The memory is used to store a computer program. The processor is configured to execute the steps of the method for synchronizing incremental data in any of the above embodiments by running the computer program stored on the memory.
[0017] According to another aspect of the embodiments of the present application, there is also provided a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps of the method for synchronizing incremental data in any of the above embodiments when running.
[0018] According to another aspect of the embodiments of the present application, there is also provided a computer program product or a computer program. The computer program product or the computer program includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the steps of the method for synchronizing incremental data in any of the above embodiments.
[0019] In the embodiments of the present application, an incremental data synchronization method is adopted. When it is determined that a target sub-service in a target business system at the source end is updated, operation records corresponding to the incremental data generated at the time of update are obtained. The target business system includes multiple sub-services, and the operation records are used to represent multiple associated information corresponding to the generation of the incremental data. The associated information includes a timestamp. The operation records are respectively sent to each topic partition of a distributed message queue according to the classification of the associated information. A streaming processing engine extracts and consumes the operation records in the topic partition, and writes the consumed operation records into a second database to obtain an operation record sorting queue. The second database stores the consumed operation records in a preset format, and the operation record sorting queue is a queue generated by sorting the consumed operation records according to the sequence of timestamps. The streaming processing engine reads the operation record sorting queue stored in the second database and sends the incremental data corresponding to the operation record sorting queue to the target end, so that the target end completes the synchronization of the incremental data. Since the embodiments of the present application utilize the operation records corresponding to the incremental data generated when the target sub-service is updated, record these operation records in the second database, use the second database as a transfer storage medium for incremental synchronization operation records, and automatically sort these operation records based on the sequence of timestamps corresponding to the operation records by the second database. In this way, when the processing engine reads the sorted operation records and sends the corresponding incremental data to the target end, the sequence of messages is ensured, and the order accuracy of incremental data synchronization from the source end to the target end is achieved, thereby solving the problem of inaccurate incremental data synchronization between the source end and the target end in the related art. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] The accompanying drawings herein are incorporated into and constitute a part of this specification, showing embodiments consistent with the present application and, together with the specification, are used to explain the principles of the present application.
[0021] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0022] Figure 1 It is a schematic flowchart of an optional method for synchronizing incremental data according to an embodiment of the present application;
[0023] Figure 2 It is a schematic overall flowchart of an optional method for synchronizing incremental data according to an embodiment of the present application;
[0024] Figure 3 It is a structural block diagram of an optional device for synchronizing incremental data according to an embodiment of the present application;
[0025] Figure 4 It is a structural block diagram of an optional electronic device according to an embodiment of the present application. Detailed implementation manners
[0026] In order to enable those skilled in the art to better understand the solutions of the present application, the following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0027] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned accompanying drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of the present application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily need to be limited to those clearly listed steps or units, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0028] In order to solve the problem of inaccurate incremental data synchronization from the source to the target, related technology 1: synchronization based on MySQL auto-increment columns, each synchronization caches the starting number of the current auto-increment field and the amount of field increment, and then extracts the result data of the SQL query and imports it into the target in batches. However, this method is too dependent on the data in the table and cannot be applied to scenarios without auto-increment columns, such as scenarios where the user's cleansed and encrypted ID is used as the unique primary key, or scenarios where the business does not add auto-increment columns when configuring the form.
[0029] Related technology 2: briefly suspend the business, export the source data to the target, and after the target catches up with the source data, adopt the dual-write mode and periodically verify the data on both sides. However, the dual-write mode involved in this method is only applicable to simple business scenarios. For core and complex businesses, it is obviously impossible to suspend the business and wait for catching up.
[0030] Related technology 3: Parse MySQL Binlog logs, restore the change records of the target database table, and repeat the change operation in the target table. However, the corresponding streaming parsing of Binlog logs and repeating the change operation in Binlog in the target table can only write to one Kafka partition to ensure the order consistency when using Kafka as a message queue. Consumers can only read from one partition, which greatly weakens Kafka's ability to distribute load. Due to the different business volumes on the source side, it is easy to have uneven partition distribution, message queue blocking, and consumer overload.
[0031] Based on the defects of the above-mentioned related technologies, the present application embodiment provides a method for synchronizing incremental data, such as Figure 1 As shown, the method is applied to a background server, and the method includes:
[0032] Step S101, when it is determined that the target sub-business in the target business system of the source end is updated, obtain the operation record corresponding to the incremental data generated when it is updated, wherein the target business system contains multiple sub-businesses, and the operation record is used to represent multiple related information corresponding to the generation of incremental data, and the related information includes a timestamp.
[0033] Optionally, in an embodiment of the present application, the source end for data synchronization is first determined, and each source end is composed of a target business system, which in turn contains multiple sub-businesses. When the background server obtains that the target sub-business of the source end has been updated, it indicates that incremental data will be generated at this time. At this time, the operation record corresponding to the incremental data generated when the target sub-business is updated is obtained, and the operation record contains multiple associated information related to the incremental data, such as timestamp, database operation type, etc.
[0034] Step S102: Send the operation records to each topic partition of the distributed message queue according to the classification of the associated information.
[0035] Optionally, synchronously send these operation records to a distributed message queue, such as Kafka. It should be noted that when the distributed message queue selected in the embodiment of the present application is Kafka, one of the usage scenarios of Kafka is to decouple the producer and the consumer. The producer and the consumer no longer directly contact each other. It can perform topic partitioning. That is, in the Kafka file storage, there are multiple different partitions under the same topic. At this time, the operation records can be sent to each topic partition of Kafka respectively based on the business scenario classification corresponding to the associated information in the operation records.
[0036] Step S103: Use a streaming processing engine to extract and consume the operation records in the topic partition, and write the consumed operation records into the second database to obtain an operation record sorting queue. Among them, the second database stores the consumed operation records according to a preset format, and the operation record sorting queue is a queue generated by sorting the consumed operation records according to the chronological order of the timestamps.
[0037] Optionally, use a streaming processing engine, such as Flink, Spark Streaming, etc., to extract and consume the operation records in each topic partition, and then write the consumed operation records into the second database, such as the Hbase database. Then, based on the characteristics of the database, the sorting of the operation records is realized to obtain an operation record sorting queue. Among them, the Hbase database is a distributed, column-oriented open-source database. It is distributed and is good at processing big data. The embodiment of the present application uses Flink as the streaming processing engine.
[0038] It can be understood that the streaming processing engine and the Kafka queue are usually applied in the scenario of message production and message consumption between the producer and the consumer. The operation record sorting queue is also a queue generated by sorting according to the chronological order of the timestamps recorded in the operation records. In addition, the Hbase database has its own table structure. At this time, when recording the operation records in the Hbase database, they need to be stored according to the preset format of the table structure.
[0039] Step S104: Use a streaming processing engine to read the operation record sorting queue stored in the second database, and send the incremental data corresponding to the operation record sorting queue to the target end, so that the target end completes the synchronization of the incremental data.
[0040] Optionally, the sorted operation records have been stored in the second database (i.e., the Hbase database). The background server of the embodiment of the present application will start the streaming processing engine to read the operation record sorting queue in the second database again, consume the operation records in the operation record sorting queue one by one, and send the read incremental data to the database at the target end for synchronous writing, so that the target end finally completes the synchronization of the incremental data. At the same time, after Flink consumes each operation record, it can clear the corresponding data of this operation record stored in Hbase.
[0041] In the embodiment of the present application, an incremental data synchronization method is adopted. When it is determined that the target sub-service in the target business system at the source end is updated, the operation records corresponding to the incremental data generated at the time of the update are obtained. Among them, the target business system contains multiple sub-services, and the operation records are used to represent multiple associated information corresponding to the generation of the incremental data, and the associated information includes a timestamp; the operation records are sent to each topic partition of the distributed message queue; the streaming processing engine is used to extract and consume the operation records in the topic partition, and write the consumed operation records into the second database to obtain an operation record sorting queue. Among them, the second database stores the consumed operation records in a preset format, and the operation record sorting queue is a queue generated by sorting the consumed operation records according to the order of timestamps; the streaming processing engine reads the operation record sorting queue stored in the second database, and sends the incremental data corresponding to the operation record sorting queue to the target end, so that the target end completes the synchronization of the incremental data. Since the embodiment of the present application uses the operation records corresponding to the incremental data generated when the target sub-service is updated, records these operation records in the second database, uses the second database as a transfer storage medium for incremental synchronization operation records, and automatically sorts these operation records based on the order of timestamps corresponding to the operation records by the second database. In this way, when the processing engine reads the sorted operation records and sends the corresponding incremental data to the target end, the order of messages is guaranteed, and the order accuracy of incremental data synchronization from the source end to the target end is realized, thereby solving the problem of inaccurate incremental data synchronization between the source end and the target end in the related art.
[0042] As an alternative embodiment, obtaining the operation records corresponding to the incremental data generated at the time of the update includes:
[0043] Obtain the first database in the target business system, where the first database stores the incremental data after the update of the target sub-service;
[0044] Perform instantiation configuration on the incremental data stored in the first database to obtain readable data;
[0045] Obtain the operation records corresponding to the readable data.
[0046] Optionally, the database is used to record some operations performed by the business system on the data stored in the database. In this case, the embodiment of the present application needs to obtain the first database in the source-side target business system. The first database can be a MySQL database, and the first database stores the incremental data after the target sub-business is updated and the original data before the update.
[0047] At this time, in order to read the required incremental data from the MySQL database, it is necessary to perform instantiation configuration on the incremental data to obtain readable data, and then find some corresponding operation records based on the readable data.
[0048] In the embodiment of the present application, since all data updates are obtained or executed in the database, in order to obtain readable data, it is necessary to instantiate the database. Only then can the incremental data be read from the memory, and at this time, the incremental data is readable and operable. Then, the corresponding search for operation records is performed to obtain accurate incremental data.
[0049] As an alternative embodiment, obtaining the operation records corresponding to the readable data includes:
[0050] Converting the readable data into a binary log file, where the binary log file is used to record the updated data and the data to be updated stored in the first database;
[0051] Parsing the binary log file according to a parsing tool to obtain the operation records executed when the target sub-business is updated.
[0052] Optionally, the operations of the target sub-business in the target business system on the database will be saved as a binary log file, such as a Binlog log file. The binary log file records the updated data and the data to be updated stored in the first database. That is, as long as there is data update in the first database, it can be saved as a binary log file.
[0053] Then, parse the Binlog log file through a database operation record parsing tool (such as the canal open-source parsing tool of xxxx) to generate the operation records for the database tables.
[0054] The following are examples of operation records:
[0055] / / UPDATE
[0056] {"data":[{"id":"29","order_id":"291029","amount":"281727.39","create_time":"2020-09-25 03:17:59"}],"database":"testbase","es":1600975079000,"id":4,"isDdl":false,"mysqlType":{"id":"BIGINT","order_id":"VARCHAR(64)","amount":"DECIMAL(10,2)","create_time":"DATETIME"},"old":[{"amount":"666.0"}],"pkNames":["id"],"sql":"","sqlType":{"id":-5,"order_id":12,"amount":3,"create_time":93},"table":"order","ts":1600975179000,"type":"UPDATE"}。
[0057] Among them, data: The modified data is included here, that is, the key and value of the table structure fields; id: Identification code; order_id: Identification code order; amount: Total number; create_time: Creation time; database: The database name corresponding to the Binlog (i.e., binary log); es: The execution timestamp of the Binlog, which can be in milliseconds; id in database: The affiliated identification code of the Binlog; isDdl: Whether it is a ddl statement; mysqlType: Table structure; old: The value of the field before modification; pkNames: Primary key field name; sql, sqltype: The original sql instruction text and classification parsed and restored; table: Table name; ts: The timestamp during parsing, which can be in milliseconds; type: Database operation type. It should be noted that the meaning of id (ID) in each embodiment of this application is the identification code.
[0058] In the embodiments of this application, by converting the readable data into a binary log file and parsing the binary log file according to a parsing tool, the operation records executed when the target sub-business is updated are obtained, so as to achieve the purpose of quickly obtaining various operations performed on the database by the target sub-business in the target business system.
[0059] As an alternative embodiment, a streaming processing engine is used to extract and consume the operation records in the topic partition, and write the consumed operation records into the second database, and the obtained operation record sorting queue includes:
[0060] The streaming processing engine consumes the operation records in the topic partition based on preset key information to obtain the consumed operation records;
[0061] Obtain the preset format for storing data in the second database;
[0062] Use the timestamp as the version number and write the consumed operation records into the preset unit corresponding to the preset format according to the preset format;
[0063] Sort the consumed operation records in ascending order of the version number to obtain an operation record sorting queue.
[0064] Optionally, when the embodiments of the present application use Flink to batch consume the operation records in each topic partition, some preset key information can be set in advance, and based on these preset key information, some key information is recorded during the consumption process.
[0065] Specifically, during the processing and calculation of Flink, based on the preset key information, the operation records that match the preset key information in the operation records are recorded. For example, the preset key information is: database name, table name, Binlog ID, timestamp, etc. Among them, recording the database name, table name, and Binlog ID realizes deduplication to prevent the situation of duplicate sending of upstream system data due to network reasons. Recording the timestamp is for the realization of the sorting function of the second database. In addition, the entire operation record can also be directly recorded. In this way, a status information is recorded and status persistence is performed to ensure fault tolerance. When the Flink task fails and hangs, it can be automatically restarted from the status in persistence without losing data.
[0066] Based on the database table structure of the second database, obtain the preset format for storing data in the second database. Since the embodiments of the present application can select Hbase as the second database;
[0067] At the same time, the principle of Hbase is automatic sorting based on the version number. Therefore, in the embodiments of the present application, Hbase uses the timestamp as the version number, automatically sorts the operation records consumed by Flink according to the size order of the version number, and writes the operation records consumed by Flink into the preset unit corresponding to the preset format of the second database according to the preset format. The binary code stream can be set as the primary key id in mysql, the operation record is used as the value in Hbase, and the parsed operation timestamp is used as the version differentiation method. An example of the table structure format supported by Hbase is as follows: (rowkey,handle_msg,datasource,table,handle_type,handle_id,version).
[0068] Among them, rowkey: primary key id; handle_msg: operation record; datasource: data source; table: table; handle_type: operation type; handle_id: operation id; version: version.
[0069] An example of writing data is as follows:
[0070] (1,{"id":"1","order_id":"10086","amount":"10087.0","create_time":"2020-03-02 05:12:49"},"test","order","UPDATE","4",1583143974870),
[0071] At this time, "1583143974870" is the version number. According to the size order of the version numbers, the consumed operation records are sorted to obtain an operation record sorting queue.
[0072] It should be noted that in the embodiments of the present application, the timestamp can be obtained by calculating the Unix timestamp corresponding to the current time. For example, at 18:00, the calculated Unix timestamp is: 1630663219. Among them, the numerical size of the current time is in a proportional relationship with the numerical size of the Unix timestamp. The larger the numerical value of the current time, the larger the corresponding timestamp value.
[0073] In the embodiments of the present application, the data consumed by Flink all carry operation record timestamps. This timestamp is used as the version number of a single unit record in Hbase. For out-of-order data caused by network latency or inconsistent consumption rates in multiple partitions, Hbase can automatically sort based on the multi-version nature of the timestamp.
[0074] As an alternative embodiment, before sending the incremental data corresponding to the operation record sorting queue to the target end, the method further includes:
[0075] Converting the format of the incremental data corresponding to the operation record sorting queue to obtain the incremental data after format conversion, where the incremental data after format conversion meets the data format supported by the target end;
[0076] Sending the incremental data after format conversion to the target end, where the number of target ends is at least one.
[0077] Optionally, since the number of target ends corresponding to the source end is at least one, in order to adapt to the database support formats of different target ends, the embodiments of the present application convert the format of the incremental data corresponding to the operation record sorting queue to obtain the incremental data after format conversion, and then send the incremental data after format conversion to the target end, so that the storage of the incremental data can be completed in the database of the target end.
[0078] For example, if the database at the source end is mysql and the databases at the target ends are Hive, ElasticSearch, etc., at this time, according to the data write operation instruction, the format of the incremental data in the mysql database needs to be changed to adapt to the data formats supported by the databases at the target ends such as the Hive data warehouse and Elastic Search (search service engine, ES).
[0079] In the embodiments of the present application, the ability to synchronize across different database types can be provided to realize the synchronization of data between different databases.
[0080] As an optional embodiment, using a streaming processing engine to read the operation record sorting queue stored in the second database includes:
[0081] The streaming processing engine consumes the read operation record sorting queue according to a preset consumption scheme, where the preset consumption scheme is used to instruct the streaming processing engine to perform a consumption operation on each operation record in the operation record sorting queue exactly once.
[0082] Optionally, to ensure the consumption efficiency of the streaming consumption engine, the embodiments of the present application use a preset consumption scheme, such as the exactly-once consumption state, to ensure that the operation records are consumed and only consumed once. At the same time, all operation records can also be written to the Flink side through the checkpiont checkpoint mechanism to ensure data consistency. Here, the consumption object of the preset consumption scheme is each operation record in the operation record sorting queue stored in the second database.
[0083] As an optional embodiment, after the streaming processing engine consumes the read operation record sorting queue according to the preset consumption scheme, the method further includes:
[0084] Obtain the identification codes of each operation record in the operation record sorting queue;
[0085] Perform a check mechanism operation on the identification codes, where the check mechanism includes: an identification code error check mechanism and an identification code duplicate check mechanism;
[0086] Delete the operation records indicated as incorrect by the check mechanism from the operation record sorting queue.
[0087] Optionally, when the embodiments of the present application consume operation records using a streaming processing engine, the identification codes of each operation record in the operation record sorting queue are obtained, where the identification code is used to uniquely represent the operation record, and each operation record corresponds to an identification code. At this time, duplicate checking and error checking operations are performed on these identification codes, and the operation records with the same identification code number or incorrect identification code number are deleted from the operation record sorting queue.
[0088] At the same time, the front and back operation states of these processed operation records are both persistently stored.
[0089] In the embodiments of the present application, by performing an operation of the inspection mechanism on the identification codes of each operation record, the accuracy and uniqueness of the operation record can be guaranteed.
[0090] As an alternative embodiment, after using the streaming processing engine to read the operation record sorting queue stored in the second database memory, the method further includes:
[0091] In the case of using the streaming processing engine to consume operation records, determine the first version number corresponding to the currently consumed operation record;
[0092] Obtain all the version numbers stored in the preset unit;
[0093] Compare the numerical sizes of all the version numbers with the first version number, and delete the operation records corresponding to the second version number that is less than the first version number, where the second version number is any other version number except the first version number among all the version numbers.
[0094] Optionally, when Flink consumes, it records the first version number corresponding to the currently consumed operation record, that is, the current latest version number, and then compares the numerical sizes of all the version numbers stored in the preset unit with this first version number, and deletes the operation records corresponding to the version numbers less than this first version number. It can be understood that the version number is the timestamp, and the timestamp is converted from the current moment. Therefore, if there is a second version number whose value is less than the first version number among all the version numbers, it proves that the operation record has been consumed, and the data of these operation records is equivalent to being "expired", so the "expired" operation records need to be deleted. Among them, the second version number is any other version number except the first version number among all the version numbers.
[0095] In the embodiments of the present application, the consumed operation records are deleted through the version number to achieve the effect of data redundancy removal.
[0096] As an alternative embodiment, as Figure 2 shown Figure 2It is a schematic diagram of the overall process of an optional incremental data synchronization method according to an embodiment of the present application, and its execution steps are as follows:
[0097] (1) Obtain the Binlog log file of the mysql data source.
[0098] Specifically, after the mysql instance is reasonably configured, the operations of the business system on the database are saved as binary Binlog files.
[0099] (2) Parse the Binlog information in the Binlog log file and put it into the kafka message queue.
[0100] Specifically, the Binlog log of mysql can be parsed through a database operation record parsing tool (such as the canal open-source parsing tool) to obtain each Insert / Update / Delete instruction executed, along with the data before and after the modification, aggregated into a json format to form an operation record for the database table.
[0101] (3) Flink extracts the messages in the kafka message queue and writes them into Hbase.
[0102] Specifically, the parsed operation records are synchronously sent to the topic in kafka. Multiple partitions can be configured for the topic according to the business scenario, and through distributed writing and reading, high throughput of the operation record messages of the business database tables with frequent changes can be achieved.
[0103] During the processing and calculation process of Flink, the key information in the previously consumed operation records is recorded (for example, the database name, table name, and Binlog ID are recorded. Since the data of the upstream system may be sent repeatedly due to network reasons, deduplication is achieved through the database name, table name, and Binlog ID), or the entire handle msg is directly recorded (that is, the state is recorded), and the state is persisted to ensure fault tolerance. When the Flink task fails and hangs, it can be automatically restarted from the state in the persistence without losing data. Flink writes the timestamp and other key information of the parsed operation records into a single cell in HBase. The rowkey can be set as the primary key id in mysql, the operation record is used as the value in Hbase, and the parsed operation timestamp is used as the version number of a single cell record in Hbase (the number of versions is set to max). For out-of-order data caused by network latency or inconsistent consumption rates under multiple partitions, Hbase can automatically sort based on the multi-version nature of the timestamp.
[0104] (4) Flink reads data from Hbase and writes it to the target end. At the same time, before writing to the target end, it is necessary to delete the past version numbers, that is, delete the expired operation records, and input different target ends, such as the Hive data warehouse, the ES search service engine, and the mysql instance.
[0105] Specifically, the Flink program reads multi-version data in Hbase. Through Flink's support for the canal json format, the sorted json data stored in Hbase can be parsed into Flink sql in sequence. And Flink connects to Elasticsearch or other Hbase instances, and cooperates with Flink's linker to use Flink sql to synchronously write to the target-end database, ultimately achieving incremental data synchronization. Among them, when writing, the original data can be converted into the writing formats of different data sources (such as a doc in ES and a put in Hbase), and at the same time, the offset is saved in the state information.
[0106] In this application, the offset refers to the ID number of the Binlog record currently consumed by Flink. Here, it is also a stateful read. For records with the same ID number, duplicate removal operations will be taken. The data read by Flink is stored in the state before and after the conversion operation and is persisted.
[0107] In this application, by virtue of Flink's exactly-once state consistency, it is ensured that the message is consumed only once, and the data is checkpointed to the Flink end to ensure fault tolerance. After Flink consumes the data, the data in Hbase within this time window can be cleared.
[0108] In addition, the incremental data synchronization methods provided in the above various embodiments can be applied to various business scenarios:
[0109] 1. In a simple data center backup and disaster recovery scenario for off-site disaster tolerance, by establishing backup systems in different locations, the disaster tolerance ability of the business to resist various possible security factors can be improved. This requires real-time synchronization of the main business source data to the off-site backup data center. When a disaster occurs, the off-site backup data can be used to achieve rapid takeover, ensuring data security and business continuity.
[0110] 2. In multiple scenarios of multiple off-site data centers, to ensure that multiple data centers can provide services to the business simultaneously and implement capabilities such as traffic allocation, it is required that the target data center has the same data as the source data center, including database data, distributed cache data, distributed queue data, and various middleware data, etc. This application is mainly applied to the backup scenario of the database. The backup of databases in multiple centers involves the backup of shared data and exclusive data. The shared data is backed up unidirectionally. Through a competitive service, the data is uniformly written into one data center, and then synchronized to other data centers through the data synchronization mechanism described in this application. The exclusive data requires cross-backup, and this application can also be applied to multiple data sources.
[0111] 3. To provide strong data support for the business data analysis department, enterprises often establish large data warehouse systems, gather data from different data sources into a unified warehouse, and through data cleaning, extraction, and analysis processes, produce analytical reports to provide data support for various decisions of the enterprise, guide process improvement, monitor business costs, customer acquisition rates, daily active users, page views (PV), unique visitors (UV), and other important parameter changes.
[0112] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that this application is not limited by the described action sequence, because according to this application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0113] According to another aspect of the embodiments of the present application, there is also provided an incremental data synchronization device for implementing the above-mentioned incremental data synchronization method. Figure 3 is a structural block diagram of an optional incremental data synchronization device according to an embodiment of the present application, as Figure 3 shown. The device may include:
[0114] A first acquisition unit 301, configured to acquire an operation record corresponding to the incremental data generated when it is determined that a target sub-business in the target business system at the source end is updated. Among them, the target business system contains multiple sub-businesses, and the operation record is used to represent multiple associated information corresponding to the generation of the incremental data, and the associated information includes a timestamp;
[0115] A first sending unit 302, connected to the first acquisition unit 301, and configured to send the operation records to each topic partition of the distributed message queue according to the classification of the associated information;
[0116] A obtaining unit 303, connected to the first sending unit 302, is configured to extract operation records in a topic partition for consumption by using a streaming processing engine, and write the consumed operation records into a second database to obtain an operation record sorting queue, where the second database stores the consumed operation records in a preset format, and the operation record sorting queue is a queue generated by sorting the consumed operation records according to the chronological order of timestamps;
[0117] A second sending unit 304, connected to the obtaining unit 303, is configured to read the operation record sorting queue stored in the second database by using a streaming processing engine, and send incremental data corresponding to the operation record sorting queue to a target end, so that the target end completes the synchronization of the incremental data.
[0118] It should be noted that the first obtaining unit 301 in this embodiment may be used to execute the above step S101, the first sending unit 302 in this embodiment may be used to execute the above step S102, the obtaining unit 303 in this embodiment may be used to execute the above step S103, and the second sending unit 304 in this embodiment may be used to execute the above step S104.
[0119] Through the above modules, by using operation records corresponding to incremental data generated when a target sub-service is updated, these operation records are recorded in a second database, the second database is used as an intermediate storage medium for incremental synchronization operation records, and these operation records are automatically sorted by the second database based on the chronological order of timestamps corresponding to the operation records. In this way, when the processing engine reads the sorted operation records and sends the corresponding incremental data to the target end, the order of messages is ensured, the order accuracy of incremental data synchronization from the source end to the target end is achieved, and thus the problem of inaccurate incremental data synchronization between the source end and the target end in the related art is solved.
[0120] As an optional embodiment, the first obtaining unit includes:
[0121] A first obtaining module, configured to obtain a first database in a target service system, where the first database stores incremental data after a target sub-service is updated;
[0122] A first obtaining module, configured to perform instantiation configuration on the incremental data stored in the first database to obtain readable data;
[0123] A second obtaining module, configured to obtain operation records corresponding to the readable data.
[0124] As an optional embodiment, the second obtaining module includes:
[0125] A conversion subunit, configured to convert readable data into a binary log file, where the binary log file is used to record updated data and to-be-updated data stored in a first database;
[0126] An obtaining subunit, configured to parse the binary log file according to a parsing tool to obtain an operation record executed when a target sub-business is updated.
[0127] As an optional embodiment, the obtaining unit includes:
[0128] A second obtaining module, configured to enable a streaming processing engine to consume operation records in a topic partition based on preset key information to obtain consumed operation records;
[0129] A third obtaining module, configured to obtain a preset format for storing data in a second database;
[0130] A writing module, configured to use a time stamp as a version number and write the consumed operation records into a preset unit corresponding to the preset format according to the preset format;
[0131] A sorting module, configured to sort the consumed operation records in ascending order of the version number to obtain an operation record sorting queue.
[0132] As an optional embodiment, the apparatus further includes:
[0133] A conversion unit, configured to convert the incremental data corresponding to the operation record sorting queue into a converted format before sending the incremental data corresponding to the operation record sorting queue to a target end, where the incremental data in the converted format meets the data format supported by the target end;
[0134] A third sending unit, configured to send the incremental data in the converted format to the target end, where the number of target ends is at least one.
[0135] As an optional embodiment, the second sending unit further includes:
[0136] A consuming module, configured to enable a streaming processing engine to consume the read operation record sorting queue according to a preset consumption scheme, where the preset consumption scheme is used to instruct the streaming processing engine to perform a consumption operation on each operation record in the operation record sorting queue exactly once.
[0137] As an optional embodiment, the apparatus further includes:
[0138] A second obtaining unit, configured to obtain an identification code of each operation record in the operation record sorting queue after the streaming processing engine consumes the read operation record sorting queue according to the preset consumption scheme;
[0139] An operation unit for performing an inspection mechanism operation on an identification code, where the inspection mechanism includes: an identification code error inspection mechanism and an identification code duplicate inspection mechanism;
[0140] A first deletion unit for deleting the operation records indicated as errors by the inspection mechanism from the operation record sorting queue.
[0141] As an optional embodiment, the device further includes:
[0142] A determination unit for determining a first version number corresponding to the currently consumed operation record when consuming the operation record after reading the operation record sorting queue stored in the second database by using a streaming processing engine;
[0143] A third acquisition unit for acquiring all version numbers stored in a preset unit;
[0144] A second deletion unit for comparing the numerical sizes of all version numbers with the first version number and deleting the operation records corresponding to the second version number that is less than the first version number, where the second version number is any other version number except the first version number among all version numbers.
[0145] It should be noted here that the examples and application scenarios implemented by the above modules and corresponding steps are the same, but are not limited to the content disclosed in the above embodiments.
[0146] According to another aspect of the embodiments of the present application, there is also provided an electronic device for implementing the above incremental data synchronization method, and the electronic device may be a server, a terminal, or a combination thereof.
[0147] Figure 4 is a structural block diagram of an optional electronic device according to the embodiments of the present application, as Figure 4 shown, including a processor 401, a communication interface 402, a memory 403, and a communication bus 404. Among them, the processor 401, the communication interface 402, and the memory 403 complete mutual communication through the communication bus 404, where,
[0148] The memory 403 is used to store a computer program;
[0149] When the processor 401 executes the computer program stored on the memory 403, the following steps are implemented:
[0150] When it is determined that a target sub-service in the target business system at the source end is updated, obtain the operation records corresponding to the incremental data generated during the update. Among them, the target business system includes multiple sub-services, and the operation records are used to represent multiple associated information corresponding to the generation of the incremental data, and the associated information includes a timestamp;
[0151] Send the operation records to each topic partition of the distributed message queue respectively according to the classification of the associated information;
[0152] Use a streaming processing engine to extract and consume the operation records in the topic partition, and write the consumed operation records into the second database to obtain an operation record sorting queue. Among them, the second database stores the consumed operation records in a preset format, and the operation record sorting queue is a queue generated by sorting the consumed operation records according to the chronological order of timestamps;
[0153] Use a streaming processing engine to read the operation record sorting queue stored in the second database, and send the incremental data corresponding to the operation record sorting queue to the target end, so that the target end completes the synchronization of the incremental data.
[0154] Optionally, in this embodiment, the above communication bus may be a PCI (Peripheral Component Interconnect) bus, an EISA (Extended Industry Standard Architecture) bus, etc. The communication bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity of representation, Figure 4 only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus.
[0155] The communication interface is used for communication between the above electronic device and other devices.
[0156] The memory may include RAM, and may also include non-volatile memory, for example, at least one disk memory. Optionally, the memory may also be at least one storage device located far from the aforementioned processor.
[0157] As an example, as Figure 4 shown, the above memory 403 may but is not limited to include the first acquisition unit 301, the first sending unit 302, the obtaining unit 303, and the second sending unit 304 in the above incremental data synchronization device. In addition, it may also include but is not limited to other module units in the above incremental data synchronization device, which will not be elaborated in this example.
[0158] The above-mentioned processor may be a general-purpose processor, including but not limited to: CPU (Central Processing Unit, central processing unit), NP (Network Processor, network processor), etc.; it may also be a DSP (Digital Signal Processing, digital signal processor), ASIC (Application Specific Integrated Circuit, application-specific integrated circuit), FPGA (Field-Programmable Gate Array, field-programmable gate array) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0159] In addition, the above-mentioned electronic device further includes: a display for displaying the synchronization result of the incremental data.
[0160] Optionally, the specific examples in this embodiment may refer to the examples described in the above embodiments, and will not be elaborated herein.
[0161] Those of ordinary skill in the art can understand that Figure 4 The structure shown is only schematic. The device for implementing the method for synchronizing incremental data may be a terminal device, which may be a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a handheld computer, and a Mobile Internet Device (MID), a PAD and other terminal devices. Figure 4 It does not limit the structure of the above-mentioned electronic device. For example, the terminal device may further include more or fewer components (such as a network interface, a display device, etc.) than those shown in Figure 4 or have a different configuration from that shown in Figure 4 shown.
[0162] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing the relevant hardware of the terminal device through a program, and the program can be stored in a computer-readable storage medium. The storage medium may include: a flash drive, a ROM, a RAM, a magnetic disk or an optical disc, etc.
[0163] According to another aspect of the embodiments of the present application, a storage medium is further provided. Optionally, in this embodiment, the above-mentioned storage medium may be used to execute the program code of the method for synchronizing incremental data.
[0164] Optionally, in this embodiment, the above-mentioned storage medium may be located on at least one of the multiple network devices in the network shown in the above embodiments.
[0165] Optionally, in this embodiment, the storage medium is configured to store program code for performing the following steps:
[0166] When it is determined that the target sub-service in the target business system at the source end is updated, obtain the operation record corresponding to the incremental data generated during the update. Here, the target business system contains multiple sub-services, and the operation record is used to represent multiple associated information corresponding to the generation of the incremental data. The associated information includes a timestamp;
[0167] Send the operation records to each topic partition of the distributed message queue according to the classification of the associated information;
[0168] Use a streaming processing engine to extract and consume the operation records in the topic partition, and write the consumed operation records into a second database to obtain an operation record sorting queue. Here, the second database stores the consumed operation records in a preset format, and the operation record sorting queue is a queue generated by sorting the consumed operation records according to the chronological order of the timestamps;
[0169] Use a streaming processing engine to read the operation record sorting queue stored in the second database, and send the incremental data corresponding to the operation record sorting queue to the target end, so that the target end completes the synchronization of the incremental data.
[0170] Optionally, the specific examples in this embodiment may refer to the examples described in the above embodiments, and will not be elaborated herein.
[0171] Optionally, in this embodiment, the above storage medium may include, but is not limited to: various media such as USB flash drives, ROMs, RAMs, mobile hard disks, magnetic disks, or optical discs that can store program code.
[0172] According to another aspect of the embodiments of the present application, there is also provided a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium; a processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the steps of the method for synchronizing incremental data in any one of the above embodiments.
[0173] The serial numbers of the embodiments of the present application above are only for description and do not represent the advantages and disadvantages of the embodiments.
[0174] If the integrated units in the above embodiments are implemented in the form of software functional units and sold or used as independent products, they can be stored in the above computer-readable storage media. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing one or more computer devices (which can be personal computers, servers, or network devices, etc.) to execute all or part of the steps of the method for synchronizing incremental data in various embodiments of this application.
[0175] In the above embodiments of this application, the descriptions of the various embodiments each have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0176] In several embodiments provided by this application, it should be understood that the disclosed client can be implemented in other ways. Among them, the device embodiments described above are only illustrative. For example, the division of units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of units or modules can be in electrical or other forms.
[0177] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution provided in this embodiment.
[0178] In addition, the functional units in various embodiments of this application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.
[0179] The above are only the preferred embodiments of this application. It should be noted that for those of ordinary skill in the art, without departing from the principle of this application, several improvements and refinements can still be made, and these improvements and refinements should also be regarded as the protection scope of this application.
Claims
1. A method for synchronizing incremental data, characterized in that, the method includes: When it is determined that the target sub-business in the target business system at the source end is updated, obtain the operation record corresponding to the incremental data generated during the update. Among them, the target business system contains multiple sub-businesses, and the operation record is used to represent multiple associated information corresponding to the generation of the incremental data, and the associated information includes a timestamp; Send the operation records to each topic partition of the distributed message queue according to the classification of the associated information; Use the streaming processing engine to extract and consume the operation records in the topic partition, and write the consumed operation records into the second database to obtain an operation record sorting queue, including: the streaming processing engine consumes the operation records in the topic partition based on preset key information to obtain the consumed operation records; obtain the preset format for storing data in the second database; use the timestamp as the version number, and write the consumed operation records into the preset unit corresponding to the preset format according to the preset format; sort the consumed operation records in the order of the size of the version number to obtain the operation record sorting queue, where the second database stores the consumed operation records according to the preset format, and the operation record sorting queue is a queue generated by sorting the consumed operation records according to the sequence of timestamps; Use the streaming processing engine to read the operation record sorting queue stored in the second database, and send the incremental data corresponding to the operation record sorting queue to the target end, so that the target end completes the synchronization of the incremental data.
2. The method according to claim 1, characterized in that, the obtaining of the operation record corresponding to the incremental data generated during the update includes: Obtain the first database in the target business system, where the first database stores the incremental data after the update of the target sub-business; Perform instantiation configuration on the incremental data stored in the first database to obtain readable data; Obtain the operation record corresponding to the readable data.
3. The method according to claim 2, characterized in that, the obtaining of the operation record corresponding to the readable data includes: Convert the readable data into a binary log file, where the binary log file is used to record the updated data and the data to be updated stored in the first database; Parse the binary log file according to the parsing tool to obtain the operation record executed when the target sub-business is updated.
4. The method according to claim 1, characterized in that, Before sending the incremental data corresponding to the operation record sorting queue to the target end, the method further includes: Convert the format of the incremental data corresponding to the operation record sorting queue to obtain the incremental data after the conversion of the format, where the incremental data after the conversion of the format meets the data format supported by the target end; Send the increment data after the conversion format to the target end, where the number of the target ends is at least one.
5. The method according to claim 1, wherein, the step of using the stream processing engine to read the operation record sorting queue stored in the second database includes: The stream processing engine consumes the read operation record sorting queue according to a preset consumption scheme, where the preset consumption scheme is used to instruct the stream processing engine to perform a consumption operation on each of the operation records in the operation record sorting queue exactly once.
6. The method according to claim 5, wherein, after the stream processing engine consumes the read operation record sorting queue according to the preset consumption scheme, the method further includes: Obtain the identification codes of each of the operation records in the operation record sorting queue; Perform a check mechanism operation on the identification codes, where the check mechanism includes: an identification code error check mechanism and an identification code duplicate check mechanism; Delete the operation records indicated as errors by the check mechanism from the operation record sorting queue.
7. The method according to claim 1, wherein, after using the stream processing engine to read the operation record sorting queue stored in the second database, the method further includes: When using the stream processing engine to consume the operation record, determine the first version number corresponding to the currently consumed operation record; Obtain all the version numbers stored in the preset unit; Compare the numerical sizes of all the version numbers with the first version number, and delete the operation records corresponding to the second version numbers that are smaller than the first version number, where the second version number is any other version number except the first version number among all the version numbers.
8. An increment data synchronization device, wherein, the device includes: A first acquisition unit, configured to acquire operation records corresponding to increment data generated when a target sub-service in a target business system of a source end is updated, where the target business system includes multiple sub-services, and the operation records are used to represent multiple association information corresponding to the generation of the increment data, and the association information includes a timestamp; A first sending unit, configured to send the operation records to each topic partition of the distributed message queue; A obtaining unit is configured to consume the operation records in the topic partition by using a streaming processing engine, and write the consumed operation records into a second database to obtain an operation record sorting queue: The streaming processing engine consumes the operation records in the topic partition based on preset key information to obtain the consumed operation records; obtain a preset format for storing data in the second database; use the time stamp as a version number, and write the consumed operation records into a preset unit corresponding to the preset format according to the preset format; sort the consumed operation records in ascending order of the version number to obtain the operation record sorting queue, where the second database stores the consumed operation records in a preset format, and the operation record sorting queue is a queue generated by sorting the consumed operation records according to the sequence of the time stamps; A second sending unit is configured to read the operation record sorting queue stored in the second database by using the streaming processing engine, and send the incremental data corresponding to the operation record sorting queue to a target end, so that the target end completes the synchronization of the incremental data.
9. An electronic device, comprising a processor, a communication interface, a memory, and a communication bus, wherein, the processor, the communication interface, and the memory communicate with each other through the communication bus, and is characterized in that, the memory is configured to store a computer program; the processor is configured to execute the steps of the incremental data synchronization method according to any one of claims 1 to 7 by running the computer program stored on the memory.
10. A computer-readable storage medium, characterized in that, the storage medium stores a computer program, wherein the computer program is configured to execute the steps of the incremental data synchronization method according to any one of claims 1 to 7 when running.
11. A computer program product, comprising a computer program / instructions, characterized in that, the computer program / instructions, when executed by a processor, implement the incremental data synchronization method according to any one of claims 1 - 7.
Citation Information
Patent Citations
Data stream type synchronization system, device and method
CN112818022A