Data synchronization method, machine readable storage medium and computer equipment
By using global transaction identifiers (GTIDs) in MySQL database cluster to find the start location of data synchronization and filter duplicate transactions, the data synchronization consistency problem after master-secure switching is solved, and efficient and accurate data continuity synchronization is achieved.
Patent Information
- Application Number
- CN202411982761.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-05-06
AI Technical Summary
When data synchronization is performed by MySQL database cluster, breakpoints cannot be accurately parsed after the main and backup switch, making it difficult to ensure data synchronization consistency.
By obtaining the global transaction identifier (GTID), the start location of data synchronization is found in the logical log of the new host, and the repeated parsed transactions are filtered out, thereby achieving continuous data synchronization.
Ensure data synchronization is carried out from the appropriate nodes, avoid repeated processing of past data and missing key data, optimize data synchronization efficiency, and enhance data accuracy and consistency.
Smart Images

Figure CN119938786A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data synchronization, and in particular to a data synchronization method, a machine-readable storage medium and a computer device. Background Art
[0002] When using data synchronization tools for real-time data synchronization, there are three stages: the first stage is to initialize the loading of existing data to obtain the basic point for data synchronization; the second stage is to synchronize incremental data based on the synchronization basic point established by the initial data loading; the third stage is to regularly compare and verify the source data and target data of data synchronization to confirm that no data is lost during the data synchronization process. The second and third stages will be in parallel for a long time.
[0003] In the second phase of incremental data synchronization, it is a common real-time replication technology to obtain incremental data by analyzing database logs to achieve real-time data synchronization. This technology parses the source database online log or archive log to obtain data additions, deletions and changes, and then converts these changes into a specific message format inside the synchronization software in units of transactions, and sends them to the target data synchronization software through the data synchronization software's private transmission protocol. Finally, the target synchronization software restores the acquired transaction logs into SQL statements supported by the target database and executes them on the target database to achieve real-time data synchronization and maintain data consistency between the source and target databases.
[0004] When using the data synchronization tool to synchronize the MySQL database: on the one hand, the MySQL master-slave synchronization mechanism is adopted, the tool is registered as a standby machine, and the binlog log is parsed to obtain incremental data. The corresponding log and offset are also recorded so that the parsing can be continued according to the breakpoint after restart to ensure data consistency; on the other hand, under complex business, a single MySQL server has performance and reliability shortcomings, and cluster deployment is required to improve performance. However, when the cluster encounters a host failure and the master-slave switch occurs, because the binlog of each node is recorded separately, the data synchronization tool cannot find the accurate parsing breakpoint on the new host using the breakpoint information of the old host, resulting in the failure of breakpoint parsing and the inability to ensure data synchronization consistency. Summary of the invention
[0005] An object of the first aspect of the present invention is to overcome at least one technical defect in the prior art and to provide a data synchronization method, a machine-readable storage medium and a computer device.
[0006] A further object of the first aspect of the present invention is to achieve data connection synchronization after the primary and standby switching to ensure data consistency.
[0007] Another further object of the first aspect of the present invention is to improve the flexibility and reliability of data synchronization in different scenarios and ensure the stable operation of the database cluster.
[0008] In particular, according to a first aspect of the present invention, the present invention provides a data synchronization method, comprising:
[0009] When a host fails, a backup server is automatically selected and upgraded to the new host.
[0010] Obtaining a global transaction identifier indicating a data breakpoint, wherein the global transaction identifier includes a unique identifier of a server and a self-incrementing transaction number and is globally unique in a database cluster environment;
[0011] In the logic log of the new host, searching for the starting position of data synchronization according to the global transaction identifier indicating the data breakpoint;
[0012] Continue parsing the logical log from the starting position, and filter out transactions that are parsed repeatedly.
[0013] Optionally, if the user specifies a global transaction identifier in the online instruction of the data synchronization tool, the step of obtaining the global transaction identifier indicating the data breakpoint includes:
[0014] Using the global transaction identifier specified by the user as the global transaction identifier indicating the data breakpoint;
[0015] If the user does not specify a global transaction identifier in the online command of the data synchronization tool, the steps for obtaining the global transaction identifier indicating the data breakpoint include:
[0016] The global transaction identifier of the last message object in the disk is used as the global transaction identifier indicating the data breakpoint.
[0017] Optionally, the step of searching for a starting position of data synchronization according to the global transaction identifier indicating the data breakpoint includes:
[0018] Obtain the file names of all logical logs in the new host;
[0019] Determine whether the logical log corresponding to the global transaction identifier indicating the data breakpoint has been parsed;
[0020] If it is not parsed, the last logical log is used as the starting point for data synchronization.
[0021] Optionally, after the step of determining whether the logical log corresponding to the global transaction identifier indicating the data breakpoint has been parsed, the method further includes:
[0022] If it has been parsed, searching for logical logs corresponding to the global transaction identifier indicating the data breakpoint in sequence from back to front;
[0023] If found, the previous logical log of the logical log corresponding to the global transaction identifier indicating the data breakpoint is used as the starting position of data synchronization;
[0024] If not found, the last logical log is used as the starting position for data synchronization.
[0025] Optionally, the step of searching for logical logs corresponding to the global transaction identifier indicating the data breakpoint in sequence from back to front includes:
[0026] Obtaining a preceding global transaction identifier set in each logical log from back to front, wherein the preceding global transaction identifier set is a set consisting of global transaction identifiers corresponding to all transactions occurring before the current transaction;
[0027] Determine whether there is a single preceding global transaction identifier set that includes both the unique identifier and the transaction code of the global transaction identifier indicating the data breakpoint;
[0028] If so, determining that a logical log corresponding to the global transaction identifier indicating the data breakpoint has been found;
[0029] If not, it is determined that the logical log corresponding to the global transaction identifier indicating the data breakpoint is not found.
[0030] Optionally, the step of filtering out duplicate parsed transactions includes:
[0031] Parse an event;
[0032] If the currently parsed event is a global transaction identifier event, obtain the global transaction identifier of the current transaction;
[0033] Determine whether the global transaction identifier indicating the data breakpoint is empty;
[0034] If it is empty, parse the next event;
[0035] Optionally, after the step of determining whether the global transaction identifier indicating the data breakpoint is empty, the method further includes:
[0036] If it is not empty, determine whether the global transaction identifier of the current transaction is equal to the global transaction identifier indicating the data breakpoint;
[0037] If they are equal, then the global transaction identifier indicating the data breakpoint is set to null;
[0038] If not equal, parse the next event.
[0039] Optionally, after parsing the first event, also include:
[0040] If the event currently being parsed is a non-global transaction identifier event, continue parsing other events in the current transaction;
[0041] Determine whether the global transaction identifier indicating the data breakpoint is empty;
[0042] If it is empty, the transaction is encapsulated and the message object is returned;
[0043] If not empty, parse the next event.
[0044] According to a second aspect of the present invention, the present invention provides a machine-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements any of the above-mentioned data synchronization methods.
[0045] According to a third aspect of the present invention, the present invention provides a computer device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, and the processor implements any one of the above-mentioned data synchronization methods when executing the computer program.
[0046] The data synchronization method of the present invention can automatically switch the standby machine to a new host when a sudden failure occurs in the host, quickly restore services, and ensure the smooth operation of the business. After the service is restored, the global transaction identifier indicating the data breakpoint is obtained, and then the starting position of the data synchronization is found in the logical log of the new host using the global transaction identifier indicating the data breakpoint obtained, and finally the logical log is continued to be parsed from the starting position, and the transactions that are repeatedly parsed are filtered out. In this way, it can be ensured that data synchronization is continued from the appropriate node, which not only avoids the ineffective labor of repeatedly processing past data, but also eliminates the risk of missing key data. This not only optimizes the efficiency of data synchronization and reduces unnecessary consumption of system resources, but also further enhances the accuracy of the data, providing solid data support for business applications that rely on data consistency.
[0047] Furthermore, in the data synchronization method of the present invention, when the user specifies a global transaction identifier in the online instruction of the data synchronization tool, the identifier specified by the user is directly used as the global transaction identifier indicating the data breakpoint. This gives the user a high degree of autonomy and control, allowing it to accurately select where to start the continuation of data synchronization based on a deep understanding of the business process, data status, and past failure conditions. If the user does not specify a global transaction identifier, the global transaction identifier of the last message object on the disk is used as the global transaction identifier indicating the data breakpoint. Since the last recorded message object on the disk is closest to the latest data status before the failure occurs, using this as a breakpoint can minimize the risk of data loss caused by the failure and ensure that data synchronization covers the latest data changes as completely as possible.
[0048] Based on the following detailed description of specific embodiments of the present invention in conjunction with the accompanying drawings, those skilled in the art will become more aware of the above and other objects, advantages and features of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Various other advantages and benefits will become apparent to those of ordinary skill in the art by reading the detailed description of the preferred embodiments below. The accompanying drawings are only for the purpose of illustrating the preferred embodiments and are not to be considered as limiting the present invention. Moreover, the same reference symbols are used throughout the accompanying drawings to represent the same components. In the accompanying drawings:
[0050] Figure 1 is a schematic flow chart of a data synchronization method according to an embodiment of the present invention;
[0051] Figure 2 is a schematic flow chart of searching for a starting position of data synchronization according to an embodiment of the present invention;
[0052] Figure 3 This is the parsing status diagram when the parsing speed of the data synchronization tool is faster than that of the standby machine;
[0053] Figure 4 This is a parsing status diagram when the parsing speed of the data synchronization tool is slower than that of the standby machine.
[0054] Figure 5 is a schematic flow chart of searching logical logs in reverse order according to an embodiment of the present invention;
[0055] Figure 6 is a schematic flow chart of filtering out repeated parsed transactions according to an embodiment of the present invention;
[0056] Figure 7 is a schematic flow chart of filtering out repeated parsed transactions according to another embodiment of the present invention;
[0057] Figure 8 is a schematic diagram of a machine-readable storage medium according to an embodiment of the present invention;
[0058] Fig. 9 is a schematic diagram of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0059] It should be understood by those skilled in the art that the embodiments described below are only some embodiments of the present invention, rather than all embodiments of the present invention, and these embodiments are intended to explain the technical principles of the present invention, rather than to limit the protection scope of the present invention. Based on the embodiments provided by the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should still fall within the protection scope of the present invention.
[0060] It should be noted that the logic and / or steps represented in the flowchart or described in other ways herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, device or equipment (such as a computer-based system, a system including a processor, or other system that can fetch instructions from an instruction execution system, device or equipment and execute instructions), or used in combination with these instruction execution systems, devices or equipment.
[0061] The embodiment of the present invention first provides a data synchronization method, which aims to solve the problem that during the data synchronization process of a MySQL database cluster by a data synchronization tool, when the cluster undergoes a master-slave switch, breakpoint analysis operations cannot be performed on the data, thereby causing data synchronization to be blocked and consistency to be difficult to ensure.
[0062] It is worth mentioning that MySQL database has introduced the global transaction identifier GTID (Global Transaction Identifier) since version 5.6 and higher. GTID consists of server_id and transaction_id, where server_id is the database unique identifier generated when the MySQL database is first started, and transaction_id is an auto-increment count starting from 1, which increases in order with the birth of each new transaction. This identifier is automatically generated and assigned by the MySQL database server, is globally unique, and is used to track transactions generated in a cluster environment. Based on this, the data synchronization tool can leverage this to parse the MySQL database GTID and use it as a breakpoint to locate and parse data in the MySQL cluster.
[0063] In the embodiment of the present invention, in order to further adapt to the diverse application scenarios and meet the needs of users for fine-tuning data synchronization, it is allowed to provide external configuration parameters to control whether to enable GTID resolution. After enabling GTID resolution, the Event that records a data change transaction in the binlog log of the MySQL database will consist of the following parts: GTIDLogEvent, which records the global consistency ID of the transaction; QueryLogEvent, which usually records a BEGIN, indicating the start of a transaction; TableMapLogEvent, which records the metadata of the table involved in the change data; RowsLogEvent, which records the specific change data including insert, update and delete; XidLogEvent, which indicates the end of a transaction.
[0064] If data synchronization is performed in a MySQL cluster environment and breakpoint analysis is required in the event of a master-slave switch, users can choose to enable GTID analysis.
[0065] When the data parsing module parses the database change transaction in the binlog log, it will first parse the transaction's GTIDLogEvent, obtain the transaction's GTID value, and store it in the memory in the form of server_id:transaction_id. Subsequently, it will parse other events to obtain the change transaction data. Then, it will parse XidLogEvent. Finally, it will encapsulate the transaction into a specific message object inside the synchronization software and store the stored GTID in the metadata attribute of the object in the format of gtid=GTID, and persist the message object to the disk in a unified data format.
[0066] Figure 1 is a schematic flow chart of a data synchronization method according to an embodiment of the present invention. Figure 1 As shown, the data synchronization method at least includes the following steps S102 to S108.
[0067] Step S102: When a main machine fails, a backup machine is automatically selected and upgraded to a new main machine.
[0068] In a cluster environment, the host is responsible for the core data processing and service provision tasks. Once a host fails, the system will automatically select a suitable standby from a number of standby machines based on preset rules. The screening process may consider factors such as the standby machine's performance status, current load, and compatibility with other components to ensure that the selected standby machine can undertake subsequent work. Subsequently, the selected standby machine will be quickly upgraded to a new host, allowing it to take over the role of the original host and continue to provide services to the outside world. At this point, the data synchronization tool will identify the master-slave switch, obtain the database connection method of the new host, and launch the service.
[0069] It is worth noting that in a database cluster environment, multiple hosts may be deployed to ensure service continuity and data reliability. During the master-slave switch process, if the current host fails, the system will automatically select a device as the new host to take over. This device may be a machine originally set as a backup, or it may be another host in the cluster that plays a different role.
[0070] Step S104, obtaining a global transaction identifier indicating a data breakpoint, wherein the global transaction identifier includes a unique identifier of the server and a self-incrementing transaction number, and is globally unique in a database cluster environment.
[0071] The unique identification of the server is like the "identity card" of the database server, which accurately distinguishes different servers in the cluster environment. The self-incrementing transaction number gives each transaction a unique serial number, which increases in sequence as transactions are generated. The combination of the two, the GTID, can ensure that each transaction is uniquely identified in the entire database cluster, no matter how complex the data flow is or how frequent the interaction between nodes is.
[0072] Step S106: searching the starting position of data synchronization in the logical log of the new host according to the global transaction identifier indicating the data breakpoint.
[0073] The logical log of the new host records a series of transaction operation information that occurred on the host, which is an important basis for data synchronization. After obtaining the GTID indicating the data breakpoint, the system starts the search process. Once the log record containing the GTID is found, it can be determined as the starting position of data synchronization.
[0074] Step S108, continue parsing the logical log from the starting position, and filter out transactions that are parsed repeatedly.
[0075] After the starting position of data synchronization is successfully determined according to the previous steps, the system immediately starts the in-depth analysis process of the new host logical log. During the analysis process, if the system identifies the current transaction as a duplicate transaction, it will immediately skip the transaction and no additional data processing will be performed on the transaction.
[0076] The data synchronization method of this embodiment can ensure that data synchronization is continued from the appropriate node, avoiding the ineffective labor of repeatedly processing past data and eliminating the risk of missing key data. This not only optimizes data synchronization efficiency and reduces unnecessary consumption of system resources, but also further enhances data accuracy and provides solid data support for business applications that rely on data consistency.
[0077] In an optional embodiment, if the user specifies a global transaction identifier in the online instruction of the data synchronization tool, the step of obtaining the global transaction identifier indicating the data breakpoint may be: using the global transaction identifier specified by the user as the global transaction identifier indicating the data breakpoint.
[0078] This gives users a high degree of autonomy and control, allowing them to accurately choose where to start data synchronization based on a deep understanding of business processes, data status, and past failure scenarios.
[0079] In another optional embodiment, if the user does not specify a global transaction identifier in the online instruction of the data synchronization tool, the step of obtaining the global transaction identifier indicating the data breakpoint may be: using the global transaction identifier of the last message object in the disk as the global transaction identifier indicating the data breakpoint.
[0080] It should be noted that when the database cluster switches between the master and the slave, the data synchronization tool will identify the master-slave switch, automatically obtain the database connection method of the new host and launch the service. When the synchronization service is launched, the GITD breakpoint information will be obtained. If the user specifies the connection method in the online command and the online service tool identifies the master, the breakpoint information uses the GTID specified by the user. If the user does not specify, the synchronization service will obtain the GTID recorded in the last message object stored on the disk as the breakpoint information.
[0081] Since the last recorded message object on the disk is closest to the latest data status before the failure, using this as a breakpoint can minimize the risk of data loss due to failure and ensure that data synchronization covers the latest data changes as completely as possible.
[0082] The above two acquisition methods can cooperate with each other, which can not only meet the needs of users to actively intervene and finely control the starting point of data synchronization, but also have the adaptive ability to deal with the situation when the user does not give clear instructions.
[0083] Figure 2 is a schematic flow chart of searching for a starting position of data synchronization according to an embodiment of the present invention. Figure 2 As shown, searching for the starting position of data synchronization according to the global transaction identifier indicating the data breakpoint includes the following steps S202 to S212.
[0084] Step S202, obtaining the file names of all logical logs in the new host.
[0085] Step S204, determining whether the logical log corresponding to the global transaction identifier indicating the data breakpoint has been parsed, if it has been parsed, executing step S206, if it has not been parsed, executing step S212.
[0086] Step S206, searching for logical logs corresponding to the global transaction identifier indicating the data breakpoint in sequence from back to front.
[0087] Step S218: If found, the previous logical log of the logical log corresponding to the global transaction identifier indicating the data breakpoint is used as the starting position of data synchronization.
[0088] Step S210: If not found, the last logical log is used as the starting position for data synchronization.
[0089] Step S212: directly use the last logical log as the starting position for data synchronization.
[0090] Since the data synchronization speeds of the standby machine (new host) and the data synchronization tool are different, there are two data analysis states during the breakpoint analysis operation:
[0091] In case 1, if Figure 3 As shown in the figure, the parsing speed of the data synchronization tool is faster than that of the standby machine. At this time, the current breakpoint GTID of the data synchronization tool cannot be found in the standby machine. At this time, the data synchronization tool will start parsing from the last binlog log of the standby machine and filter out the transactions that have been parsed in the data synchronization tool.
[0092] In case 2, if Figure 4 As shown in the figure, the parsing speed of the data synchronization tool is slower than that of the standby machine. At this time, the current breakpoint GTID of the data synchronization tool has been parsed by the standby machine. At this time, the data synchronization tool needs to continue parsing from the found breakpoint and filter out some transactions that are repeatedly parsed in the binlog log.
[0093] It is understandable that in data synchronization scenarios, especially in complex and ever-changing database cluster environments, the amount of data is huge and transactions are frequently updated. Back-to-front search fully utilizes the time characteristics of log records. Newly generated transactions are always appended to the end of the log. The closer to the end, the closer it is to the current data state.
[0094] The data synchronization method of this embodiment adopts this reverse search method when searching for the starting position of data synchronization, which can quickly locate the potential matching position closest to the indicated data breakpoint. Compared with searching sequentially from front to back, it greatly reduces the invalid search range and greatly improves the search efficiency, allowing the data synchronization process to lock the starting position more quickly, saving valuable time resources and ensuring that the system can keep up with the update rhythm of the source data in a timely manner.
[0095] In addition, the database may encounter various unexpected situations during operation, such as host failure, temporary power outage, etc., which may cause data synchronization interruption. When resuming synchronization, searching from the back to the front helps to accurately connect the breakpoints. Prioritize the latest generated log part to effectively avoid repeated processing of a large amount of synchronized data due to forward search from the beginning, ensure data consistency and accuracy, and reduce the risk of data inconsistency caused by repeated operations.
[0096] Figure 5 FIG. 1 is a schematic flow chart of searching logical logs in reverse order according to an embodiment of the present invention. Figure 5 As shown, searching for the logical logs corresponding to the global transaction identifier indicating the data breakpoint in sequence from back to front includes the following steps S502 to S508.
[0097] Step S502: Obtain the preceding global transaction identifier set in each logical log in sequence from the back to the front.
[0098] Step S504, determine whether there is a single preceding global transaction identifier set that includes both the unique identifier of the global transaction identifier indicating the data breakpoint and the transaction code. If so, execute step S506; if not, execute step S508.
[0099] Step S506: determine that a logical log corresponding to the global transaction identifier indicating the data breakpoint has been found.
[0100] Step S508: determining that no logical log corresponding to the global transaction identifier indicating the data breakpoint is found.
[0101] In this embodiment, the preceding global transaction identifier set is a set consisting of global transaction identifiers corresponding to all transactions occurring prior to the current transaction.
[0102] The preceding global transaction identifier set can be represented by pre-gtids, which is part of the log records (such as binlog logs) in the database system. In the MySQL binary log (binlog) records, the relevant information of each transaction is stored in sequence. Pre-gtids is the GTID set of other transactions before the transaction stored in the log. This information can help understand the order and dependency between transactions.
[0103] During the operation of the database, when a transaction is executed and recorded in a log (such as binlog), the system will also record some pre-information related to the transaction, including pre-gtids. These pre-GTID information are automatically generated and recorded by the database logging mechanism during the transaction execution. For example, in the scenario of master-slave replication, when the master records a transaction, it will record the GTID set of previously executed transactions as pre-gtids. This information will be written into the binlog log along with the transaction, and then passed to the slave for data synchronization and transaction order judgment.
[0104] The database system will record pre-gtids according to the actual execution order of transactions. Assuming that transactions T1, T2, and T3 are executed in sequence, when recording the log information of T3, the GTIDs of T1 and T2 will be recorded as pre-gtids of T3 to indicate the transactions that have been executed before T3. In this way, in subsequent data processing, synchronization, or fault recovery scenarios, the order and association between transactions can be understood by viewing pre-gtids.
[0105] Specifically, in this embodiment, after obtaining the pre-gtids of a binlog log, first determine whether the pre-gtids contains the sid in the breakpoint GTID. If it does not contain sid, obtain the pre-gtids of the previous binlog log. If it contains sid, further determine whether the pre-gtids contains the tid in the breakpoint GTID. If it contains tid, it means that the logical log corresponding to the breakpoint GTID is found, and the parsing can be started from the previous logical log of the logical log. If it does not contain tid, further determine whether the tid in the breakpoint GTID is greater than the largest tid in the pre-gtids. If it is greater, the parsing can be started from the last logical log. If it is less, it is necessary to continue to obtain the pre-gtids of the previous binlog log.
[0106] After traversing all binlog logs in reverse order, if the logical log corresponding to the breakpoint GTID cannot be found, the data synchronization tool will also start parsing from the last logical log.
[0107] Figure 6 is a schematic flow chart of filtering out repeated parsed transactions according to an embodiment of the present invention. Figure 6 As shown, filtering out transactions that are repeatedly parsed may include the following steps S602 to S614.
[0108] Step S602, parsing an event.
[0109] Step S604: if the event currently being parsed is a global transaction identifier event, obtain the global transaction identifier of the current transaction.
[0110] Step S606, determine whether the global transaction identifier of the current transaction is empty, if it is empty, execute step S608, if not empty, execute step S610.
[0111] Step S608, parsing the next event.
[0112] Step S610, determining whether the global transaction identifier of the current transaction is equal to the global transaction identifier indicating the data breakpoint, if so, executing step S612, if not, executing step S614.
[0113] Step S612: Set the global transaction identifier indicating the data breakpoint to null.
[0114] Step S614, parse the next event.
[0115] During the parsing process, when encountering a global transaction identifier event, accurately obtain its identifier to provide a key basis for subsequent judgments and ensure that no key nodes that may be repeated are missed. When the global transaction identifier of the current transaction is empty, directly parse the next event to avoid meaningless redundant operations and save system resources. When the global transaction identifier indicating the data breakpoint is not empty, further compare the current transaction with the global transaction identifier indicating the data breakpoint to see if they are equal. If they are equal, it means that a duplicate parsed transaction is found. The global transaction identifier indicating the data breakpoint is set to empty, and the repeated part can be accurately skipped in the future to avoid the risk of data inconsistency caused by repeated work and ensure data accuracy; if they are not equal, continue to parse the next event.
[0116] Figure 7 is a schematic flow chart of filtering out repeated parsed transactions according to another embodiment of the present invention. Figure 7 As shown, filtering out transactions that are repeatedly parsed may include the following steps S702 to S710.
[0117] Step S702, parsing an event.
[0118] Step S704: If the event currently being parsed is a non-global transaction identifier event, continue parsing other events in the current transaction.
[0119] Step S706, determining whether the global transaction identifier indicating the data breakpoint is empty, if it is empty, executing step S708, if not empty, executing step S710.
[0120] Step S708, encapsulate the transaction and return the message object;
[0121] Step S710, parsing the next event.
[0122] During the parsing process, when encountering a non-global transaction identifier event, continue to parse other events in the current transaction. Here, "other events" refer to events corresponding to operations such as data insertion, update, and deletion in the current transaction other than the global transaction identifier event, aiming to fully process the non-global transaction identifier part of the current transaction. After the parsing is completed, further determine whether the global transaction identifier indicating the data breakpoint is empty. If it is empty, it means that there is no unprocessed transaction before the current transaction. Encapsulate the transaction and return the message object to complete the current transaction processing. If it is not empty, parse the next event.
[0123] Through the above two methods, when faced with the continuous influx of massive data and complex and ever-changing database scenarios, we will not be deadlocked by fixed rules, but will always maintain a flexible response attitude, continuously screen and identify data, ensure that the data synchronization process continues to move forward in an orderly manner, and seamlessly connect the data processing needs at different stages.
[0124] The embodiment of the present invention further provides a machine-readable storage medium 20 and a computer device 30 . Figure 8 is a schematic diagram of a machine-readable storage medium 20 according to an embodiment of the present invention, Fig. 9 is a schematic diagram of a computer device 30 according to one embodiment of the present invention.
[0125] The machine-readable storage medium 20 stores the computer program 21, which implements the steps of the data synchronization method of any of the above embodiments when the computer program 21 is executed by the processor 32. The computer device 30 may include a memory 31, a processor 32, and the computer program 21 stored in the memory 31 and running on the processor 32.
[0126] The computer program 21 for performing the operation of the present invention may be an assembly instruction, an instruction set architecture (ISA) instruction, a machine instruction, a machine-related instruction, a microcode, a firmware instruction, a state setting data, a configuration data of an integrated circuit, or a source code or an object code written in any combination of one or more programming languages and process programming languages. The computer program 21 may be executed entirely on the user's computer, partially on the user's computer, as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter case, the remote computer may be connected to the user's computer via any type of network (including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., using an Internet service provider via the Internet). In some embodiments, in order to perform various aspects of the present invention, an electronic circuit including, for example, a programmable logic circuit, a field programmable gate array (FPGA) or a programmable logic array (PLA) may execute computer-readable program instructions by utilizing the state information of the computer-readable program instructions to personalize the electronic circuit.
[0127] For the purposes of the description of this embodiment, the machine-readable storage medium 20 is a tangible device capable of retaining and storing a computer program 21, which may be any device that can contain, store, communicate, propagate or transmit the program 21 for use with or in conjunction with an instruction execution system, device or apparatus. More specific examples (a non-exhaustive list) of the machine-readable storage medium 20 include the following: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disk read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanical encoding device, and any suitable combination of the foregoing.
[0128] The computer device 30 may be, for example, a server, a desktop computer, a notebook computer, a tablet computer, or a smart phone. In some examples, the computer device 30 may be a cloud computing node. The computer device 30 may be described in the general context of computer system executable instructions (such as program modules) executed by a computer system. Typically, a program module may include routines, programs, target programs, components, logic, data structures, etc. that perform specific tasks or implement specific abstract data types. The computer device 30 may be implemented in a distributed cloud computing environment where remote processing devices linked via a communication network perform tasks. In a distributed cloud computing environment, program modules may be located on a local or remote computing system storage medium including a storage device.
[0129] The computer device 30 may include a processor 32 adapted to execute stored instructions, and a memory 31 providing temporary storage space for the operation of the instructions during operation. The processor 32 may be a single-core processor, a multi-core processor, a computing cluster, or any number of other configurations. The memory 31 may include a random access memory (RAM), a read-only memory, a flash memory, or any other suitable storage system.
[0130] The computer device 30 may also include a network adapter / interface and an input / output (I / O) interface. The I / O interface allows data to be input and output with external devices that may be connected to the computer device. The network adapter / interface may provide communication between the computer device and a network, which is typically shown as a communication network.
[0131] At this point, those skilled in the art should recognize that, although multiple exemplary embodiments of the present invention have been shown and described in detail herein, many other variations or modifications that conform to the principles of the present invention can still be directly determined or derived based on the content disclosed in the present invention without departing from the spirit and scope of the present invention. Therefore, the scope of the present invention should be understood and recognized as covering all these other variations or modifications.
Claims
1. A data synchronization method, comprising: When a host fails, a backup server is automatically selected and upgraded to the new host. Obtaining a global transaction identifier indicating a data breakpoint, wherein the global transaction identifier includes a unique identifier of a server and a self-incrementing transaction number and is globally unique in a database cluster environment; In the logic log of the new host, searching for the starting position of data synchronization according to the global transaction identifier indicating the data breakpoint; Continue parsing the logical log from the starting position, and filter out transactions that are parsed repeatedly.
2. The data synchronization method according to claim 1, wherein: If the user specifies a global transaction identifier in the online instruction of the data synchronization tool, the steps of obtaining the global transaction identifier indicating the data breakpoint include: Using the global transaction identifier specified by the user as the global transaction identifier indicating the data breakpoint; If the user does not specify a global transaction identifier in the online command of the data synchronization tool, the steps for obtaining the global transaction identifier indicating the data breakpoint include: The global transaction identifier of the last message object in the disk is used as the global transaction identifier indicating the data breakpoint.
3. The data synchronization method according to claim 1, wherein: The step of searching for the starting position of data synchronization according to the global transaction identifier indicating the data breakpoint comprises: Obtain the file names of all logical logs in the new host; Determine whether the logical log corresponding to the global transaction identifier indicating the data breakpoint has been parsed; If it is not parsed, the last logical log is used as the starting point for data synchronization.
4. The data synchronization method according to claim 3, wherein: After the step of determining whether the logical log corresponding to the global transaction identifier indicating the data breakpoint has been parsed, the method further includes: If it has been parsed, searching for logical logs corresponding to the global transaction identifier indicating the data breakpoint in sequence from back to front; If found, the previous logical log of the logical log corresponding to the global transaction identifier indicating the data breakpoint is used as the starting position of data synchronization; If not found, the last logical log is used as the starting position for data synchronization.
5. The data synchronization method according to claim 4, wherein: The step of searching the logical logs corresponding to the global transaction identifier indicating the data breakpoint in sequence from back to front comprises: Obtaining a preceding global transaction identifier set in each logical log from back to front, wherein the preceding global transaction identifier set is a set consisting of global transaction identifiers corresponding to all transactions occurring before the current transaction; Determine whether there is a single preceding global transaction identifier set that includes both the unique identifier and the transaction code of the global transaction identifier indicating the data breakpoint; If so, determining that a logical log corresponding to the global transaction identifier indicating the data breakpoint has been found; If not, it is determined that the logical log corresponding to the global transaction identifier indicating the data breakpoint is not found.
6. The data synchronization method according to claim 1, wherein: The steps to filter out duplicate parsed transactions include: Parse an event; If the currently parsed event is a global transaction identifier event, obtain the global transaction identifier of the current transaction; Determine whether the global transaction identifier indicating the data breakpoint is empty; If empty, parse the next event.
7. The data synchronization method according to claim 6, wherein: After the step of determining whether the global transaction identifier indicating the data breakpoint is empty, the method further includes: If it is not empty, determine whether the global transaction identifier of the current transaction is equal to the global transaction identifier indicating the data breakpoint; If they are equal, setting the global transaction identifier indicating the data breakpoint to empty; If not equal, parse the next event.
8. The data synchronization method according to claim 6, wherein: After the first event parsing step, also include: If the event currently being parsed is a non-global transaction identifier event, continue parsing other events in the current transaction; Determine whether the global transaction identifier indicating the data breakpoint is empty; If it is empty, the transaction is encapsulated and the message object is returned; If not empty, parse the next event.
9. A machine-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the data synchronization method according to any one of claims 1 to 8.
10. A computer device, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the processor implements the data synchronization method according to any one of claims 1 to 8 when executing the computer program.