Data Synchronization Method and Device
By using the target storage system of the MySQL database as a virtual master node, establishing a connection with the slave node and synchronizing pre-write log data, the problem of the inability to reproduce the historical binlog files after recycling during the incremental synchronization of the MySQL database is solved, and it is ensured that the slave node can be set to read-only mode to ensure data consistency.
Patent Information
- Application Number
- CN202411274115.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-11
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2044-09-11
AI Technical Summary
During the incremental synchronization process of MySQL database, the historical binlog file cannot be replayed lately or replayed from specific historical sites after it is recycled. When synchronized by third-party middleware, the slave node cannot be set to read-only mode, which has the risk of data inconsistency.
Use the target storage system as the virtual master node, establish a connection between the slave node and the virtual master node, determine the synchronized pre-write log data in the slave node, and synchronize the data between the virtual master node and the slave node based on this data, simulating the native master-slave replication mechanism of the database.
The slave node can be set to read-only mode, while ensuring data consistency between the slave node and the target storage system, avoiding the risk of data inconsistency between the master and slave nodes.
Smart Images

Figure CN119128015B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of computer technologies, specifically to the fields of databases and cloud services technologies, and in particular to a data synchronization method, apparatus, computer-readable medium, electronic device, and computer program product. Background Art
[0002] There are mainly two solutions for incremental synchronization (replication) of MySQL databases: one is the native master-slave replication of MySQL; the other is to introduce a third-party publish-subscribe middleware (such as: Canal, Maxwell, etc.).
[0003] Since the storage space on a single MySQL node is limited, an expiration mechanism (i.e., "binlog rotate") is required to periodically delete expired binlog (binary log) files. Thus, once the historical binlog files are recycled, it is impossible to perform delayed replay based on the historical binlog files or replay from a specific historical position.
[0004] By introducing a publish-subscribe middleware, the binlog is dumped into a third-party storage system. As long as the storage capacity of the third-party storage system can be continuously horizontally expanded, it is even possible to cache all binlog files generated in the history of a MySQL instance. However, when the adapter program at the subscription end connects to the slave node of the MySQL database, only an ordinary client-server connection can be established. In this way, the slave node cannot be set to read-only mode, and there is a risk of data inconsistency between the master and slave nodes. Summary of the Invention
[0005] The embodiments of the present application propose a data synchronization method, apparatus, computer-readable medium, electronic device, and computer program product.
[0006] In a first aspect, the embodiments of the present application provide a data synchronization method, including: using a target storage system as a virtual master node, establishing a connection between a slave node and the virtual master node; determining the pre-written log data that has been synchronized in the slave node; and performing data synchronization between the virtual master node and the slave node according to the pre-written log data that has been synchronized.
[0007] In some examples, establishing a connection between the slave node and the virtual master node as described above includes: in response to the slave node calling the original connection function, calling the new connection function through the function hook corresponding to the original connection function; in response to the preset network parameters being consistent with the master node address of the virtual master node, establishing a virtual connection between the slave node and the virtual master node, where the original connection function is used to establish a connection between the slave node and the actual master node, and during the startup process of the application program of the slave node, the dynamic link library including the new connection function is loaded.
[0008] In some examples, establishing a connection between the slave node and the virtual master node as described above further includes: creating a client adapted to the virtual master node in the slave node according to the node type of the virtual master node; establishing a real connection between the slave node and the virtual master node through the client.
[0009] In some examples, determining the synchronized pre-written log data in the slave node as described above includes: in response to the slave node calling the original sending function, calling the new sending function through the function hook corresponding to the original sending function, and sending a dump request to the virtual master node, where the original sending function is used to send a dump request to the actual master node corresponding to the slave node, and during the startup process of the application program of the slave node, the dynamic link library including the new sending function is loaded; parsing the dump request to determine the synchronized pre-written log data of the slave node relative to the virtual master node.
[0010] In some examples, performing data synchronization between the virtual master node and the slave node according to the synchronized pre-written log data as described above includes: in response to the slave node calling the original polling function, calling the new polling function through the function hook corresponding to the original polling function to determine whether the virtual master node includes unsynchronized pre-written log data other than the synchronized pre-written log data, where the original polling function is used to determine whether unsynchronized pre-written log data sent by the actual master node is received in the receive buffer of the network connection between the slave node and the actual master node, and during the startup process of the application program of the slave node, the dynamic link library including the new polling function is loaded; in response to determining that there is unsynchronized pre-written log data and the slave node calling the original receiving function, calling the new receiving function through the function hook corresponding to the original receiving function, and synchronizing the unsynchronized pre-written log data to the slave node, where the original receiving function is used to receive data from the actual master node, and during the startup process of the application program of the slave node, the dynamic link library including the new receiving function is loaded.
[0011] In some examples, determining whether the virtual master node includes unsynchronized pre-written log data other than the synchronized pre-written log data includes: in response to determining that the virtual master node supports a polling operation, determining whether the virtual master node includes unsynchronized pre-written log data through the polling interface of the virtual master node; in response to determining that the virtual master node does not support a polling operation, determining whether the virtual master node includes unsynchronized pre-written log data according to the log file indicated by the latest generated metadata in the virtual master node.
[0012] In some examples, synchronizing the unsynchronized pre-written log data to the slave node includes: storing the unsynchronized pre-written log data in a receiving buffer; dumping the unsynchronized pre-written log data in the receiving buffer as the relay log of the slave node.
[0013] In some examples, using the target storage system as the virtual master node includes: using the distributed storage system and / or message queue in the third-party storage system as the virtual master node.
[0014] In a second aspect, an embodiment of the present application provides a data synchronization device, including: a connection unit configured to use the target storage system as a virtual master node and establish a connection between the slave node and the virtual master node; a determination unit configured to determine the synchronized pre-written log data in the slave node; a synchronization unit configured to perform data synchronization between the virtual master node and the slave node according to the synchronized pre-written log data.
[0015] In some examples, the connection unit is further configured to: in response to the slave node calling the original connection function, call the new connection function through the function hook corresponding to the original connection function, and in response to the preset network parameters being consistent with the master node address of the virtual master node, establish a virtual connection between the slave node and the virtual master node, where the original connection function is used to establish a connection between the slave node and the actual master node, and during the startup process of the application program of the slave node, the dynamic link library including the new connection function is loaded.
[0016] In some examples, the connection unit is further configured to: create a client adapted to the virtual master node in the slave node according to the node type of the virtual master node; establish a real connection between the slave node and the virtual master node through the client.
[0017] In some examples, the determination unit is further configured to: in response to the slave node calling the original sending function, call the new sending function through the function hook corresponding to the original sending function, and send a dump request to the virtual master node, where the original sending function is used to send a dump request to the actual master node corresponding to the slave node, and during the startup process of the application program of the slave node, the dynamic link library including the new sending function is loaded; parse the dump request to determine the synchronized pre-written log data of the slave node relative to the virtual master node.
[0018] In some examples, the above synchronization unit is further configured to: in response to a slave node calling the original polling function, call the new polling function through the function hook corresponding to the original polling function to determine whether the virtual master node includes unsynchronized pre-written log data other than the synchronized pre-written log data, where the original polling function is used to determine whether unsynchronized pre-written log data sent by the actual master node is received in the receive buffer of the network connection between the slave node and the actual master node. During the startup process of the application program of the slave node, the dynamic link library including the new polling function is loaded; in response to determining that there is unsynchronized pre-written log data and the slave node calls the original receiving function, call the new receiving function through the function hook corresponding to the original receiving function to synchronize the unsynchronized pre-written log data to the slave node, where the original receiving function is used to receive data from the actual master node. During the startup process of the application program of the slave node, the dynamic link library including the new receiving function is loaded.
[0019] In some examples, the above synchronization unit is further configured to: in response to determining that the virtual master node supports the polling operation, determine whether the virtual master node includes unsynchronized pre-written log data through the polling interface of the virtual master node; in response to determining that the virtual master node does not support the polling operation, determine whether the virtual master node includes unsynchronized pre-written log data according to the log file indicated by the latest generated metadata in the virtual master node.
[0020] In some examples, the above synchronization unit is further configured to: store the unsynchronized pre-written log data in the receive buffer; and dump the unsynchronized pre-written log data in the receive buffer as the relay log of the slave node.
[0021] In some examples, the above connection unit is further configured to: use the distributed storage system and / or message queue in the third-party storage system as the virtual master node.
[0022] In a third aspect, an embodiment of the present application provides a computer-readable medium, on which a computer program is stored, where when the program is executed by a processor, the method described in any implementation manner of the first aspect is implemented.
[0023] In a fourth aspect, an embodiment of the present application provides an electronic device, including: one or more processors; a storage device, on which one or more programs are stored, and when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any implementation manner of the first aspect.
[0024] In a fifth aspect, a computer program product is provided, including: a computer program, and when the computer program is executed by a processor, the method described in any implementation manner of the first aspect is implemented.
[0025] The data synchronization method and device provided by the embodiments of the present application establish a connection between a slave node and a virtual master node by using a target storage system as the virtual master node; determine the pre-written log data that has been synchronized in the slave node; and perform data synchronization between the virtual master node and the slave node according to the pre-written log data that has been synchronized, thereby simulating the native master-slave replication mechanism of a database and realizing the data synchronization process from the target storage system as the virtual master node to the slave node, enabling the slave node to be set to read-only mode, and ensuring the data consistency between the slave node and the target storage system on the basis of ensuring the data storage effect of the target storage system. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Other features, objects, and advantages of the present application will become more apparent by reading the following detailed description of non-limiting embodiments with reference to the accompanying drawings:
[0027] Figure 1 is an exemplary system architecture diagram to which an embodiment of the present application can be applied;
[0028] Figure 2 is a flowchart of an embodiment of the data synchronization method according to the present application;
[0029] Figure 3 is a schematic diagram of the master-slave replication mechanism in the native master-slave mode of a database;
[0030] Figure 4 is a schematic diagram of the master-slave synchronization process of a database based on the group commit mechanism;
[0031] Figure 5 is a schematic diagram of the data synchronization principle of a third-party real-time data publishing tool;
[0032] Figure 6 is a schematic diagram of an application scenario of the data synchronization method according to this embodiment;
[0033] Figure 7 is a schematic diagram of the passing process between the target storage system and the slave node;
[0034] Figure 8 is a schematic diagram of the replacement principle under the Linux system;
[0035] Figure 9 is a schematic diagram of the functional modules of the data synchronization system;
[0036] Figure 10 is an interaction flowchart between the target storage system and the slave node;
[0037] Figure 11 is a schematic diagram of the process of the connect hook function;
[0038] Figure 12 It is a schematic flowchart of the send hook function;
[0039] Figure 13 It is a schematic flowchart of the poll hook function;
[0040] Figure 14 It is a schematic flowchart of the recv hook function
[0041] Figure 15 It is a flowchart of another embodiment of the data synchronization method according to the present application;
[0042] Figure 16 It is a structural diagram of an embodiment of the data synchronization device according to the present application;
[0043] Figure 17 It is a schematic structural diagram of a computer system suitable for implementing the embodiments of the present application. Detailed implementation manners
[0044] The present application will be further described in detail below with reference to the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the related invention, rather than limiting the invention. Additionally, it should be noted that for the sake of convenience of description, only parts related to the relevant invention are shown in the drawings.
[0045] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other. The present application will be described in detail below with reference to the drawings and embodiments.
[0046] It should be noted that in the technical solutions of the present disclosure, in aspects such as the collection, gathering, updating, analysis, processing, use, transmission, and storage of user personal information, they all comply with the provisions of relevant laws and regulations, are used for legal purposes, and do not violate public order and good customs. Necessary measures are taken for user personal information to prevent illegal access to user personal information data, and to safeguard user personal information security, network security, and national security.
[0047] Figure 1 An exemplary architecture 100 to which the data synchronization method and device of the present application can be applied is shown.
[0048] As Figure 1 shown, the system architecture 100 may include terminal devices 101, 102, 103, a network 104, and a server 105. The terminal devices 101, 102, 103 are communicatively connected to form a topology network, and the network 104 is used as a medium to provide a communication link between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.
[0049] The terminal devices 101, 102, and 103 can interact with the server 105 through the network 104 to receive or send data, etc. The terminal devices 101, 102, and 103 can be hardware devices or software that support network connections for data interaction and data processing. When the terminal devices 101, 102, and 103 are hardware, they can be various electronic devices that support network connections, information acquisition, interaction, display, processing, etc., including but not limited to smartphones, in-vehicle computers, tablet computers, e-book readers, laptop computers, and desktop computers, etc. When the terminal devices 101, 102, and 103 are software, they can be installed in the above-listed electronic devices. It can be implemented as, for example, multiple software or software modules for providing distributed services, or can also be implemented as a single software or software module. No specific limitation is made here.
[0050] The server 105 can be a server that provides various services. For example, based on the data update operations sent by the user through the terminal devices 101, 102, and 103 to the database, the main node of the database synchronizes the pre-written log data to the target storage system with a larger data capacity for backup, and uses the target storage system as the virtual main node to implement the data synchronization process between the virtual main node and the slave node corresponding to the virtual main node. As an example, the server 105 can be a cloud server.
[0051] It should be noted that the server can be hardware or software. When the server is hardware, it can be implemented as a distributed server cluster composed of multiple servers, or can also be implemented as a single server. When the server is software, it can be implemented as multiple software or software modules (such as software or software modules for providing distributed services), or can also be implemented as a single software or software module. No specific limitation is made here.
[0052] It should also be noted that the data synchronization method provided by the embodiments of the present application is generally executed by the server, but it does not exclude the possibility of being executed by the terminal device, or being executed by the server and the terminal device in cooperation with each other. Correspondingly, each part (such as each unit) included in the data synchronization device can be all set in the server, or can be all set in the terminal device, or can also be respectively set in the server and the terminal device.
[0053] It should be understood that Figure 1 the numbers of the terminal devices, networks, and servers in [[ ]] are only illustrative. According to the implementation requirements, there can be any number of terminal devices, networks, and servers. When the electronic device on which the data synchronization method runs does not need to perform data transmission with other electronic devices, the system architecture can only include the electronic device (such as a server or a terminal device) on which the data synchronization method runs.
[0054] Continue to refer to Figure 2 , which shows the flow 200 of an embodiment of the data synchronization method, including the following steps:
[0055] Step 201, use the target storage system as the virtual master node and establish a connection between the slave node and the virtual master node.
[0056] In this embodiment, the execution entity of the data synchronization method (such as Figure 1 the terminal device or server in) can use the target storage system as the virtual master node and establish a connection between the slave node and the virtual master node.
[0057] Taking the MySQL database system as an example, there are mainly two types of solutions for incremental synchronization (replication): one is the native master-slave replication of MySQL; the other is to introduce a third-party publish-subscribe middleware (such as: Canal, Maxwell, etc.). Generally speaking, the subscription end of MySQL (i.e., the slave library) is called "slave (slave node)" or "replica (subscription end)".
[0058] Taking the MySQL database system as an example, continue to refer to Figure 3 , which shows the schematic diagram of the master-slave replication mechanism in the native master-slave mode of the database system.
[0059] The MySQL database system can store data update operations in the form of binlog (such as: mysql-bin.000001, mysql-bin.000002, mysql-bin.000003, and the suffix of the log file increases in the order of log generation time). By parsing and replaying the binlog data in sequence, the replication of the data source (i.e., the MySQL instance that generates binlog) can be achieved. In a MySQL data synchronization solution, the system used for binlog dumping is called the "Publish / Subscribe" system, or the CDC (Change data capture) system. First, the binlog generated by the MySQL instance (i.e., the data producer) can be collected and temporarily stored in the above dumping system; then, these log files can be pushed to different data consumers, or subscribed by different consumers. Through the above dumping system, the MySQL instance can synchronize data to multiple heterogeneous data consumers (such as: relational databases, such as PostgreSQL, Oracle; document databases, such as ElasticSearch; KV databases, such as TiKV).
[0060] The master node of MySQL converts data updates into the binlog format and transfers the binlog data to the slave node through the data replication protocol between the master and slave nodes (i.e., the binlogdump command). In the slave node, the IO (input / output) thread is responsible for network communication with the master node and dumps the content of the binlog data to a local disk file in the form of "Relay log". The SQL (Structured Query Language) thread in the slave node reads and parses the relay log and replays the log content to achieve data synchronization with the master node.
[0061] In the native master-slave replication mode, the slave node can be set to read-only mode. The read-only mode of the slave node and the data synchronization process from the master node to the slave node responsible by the SQL thread are two independent processes. The read-only setting of the slave node mainly affects external data writing operations to the slave node, while the data synchronization process from the master node to the slave node is automatically completed through the replication mechanism of the MySQL system and is not affected by the read-only state of the slave node.
[0062] According to different data reliabilities, the synchronization process is divided into asynchronous replication, semi-synchronous replication, and synchronous replication. Among them, in the asynchronous replication mode, after the master node sends the binlog data to the slave node, it can complete the commit of the write transaction without waiting for the relay log of the slave node to be stored on disk; in the semi-synchronous replication mode, after the relay log in the slave node is stored on disk, it sends an ack (acknowledgement) packet to the master node, and the master node can complete the commit of the write transaction only after receiving the ack. In this way, on the one hand, it can ensure the consistency of write transactions between the master and slave, and on the other hand, it does not require the slave node to confirm the commit after determining consistency with the master node (this mechanism is called "synchronous replication"), which is a solution that takes into account both performance and consistency.
[0063] Continue to refer to Figure 4 , which shows a schematic diagram of the master-slave synchronization process of the database based on the group commit mechanism.
[0064] The binlog group commit mechanism of the MySQL database can be regarded as a way to optimize write performance, that is, a group of transactions to be committed are submitted in a batch manner, rather than each transaction performing the commit action separately. The specific implementation method is that within a certain time interval, the transactions that enter the committable state store their respective updated contents in the binlog cache (binary log file cache); then the transaction with the smallest transaction number in this group of transactions is selected as the commit leader, responsible for driving the commit actions of all transactions in the binlog cache. On the one hand, the data in the binlog cache will be appended to the end of the binlog file on the master node locally, and it will be persisted (flushed) to the disk through the "disk write" operation (i.e., fdatasync); on the other hand, this part of the binlog will be sent to the slave node through the master-slave replication protocol of MySQL. When the binlogdump reply (ack packet) from the slave node is received, the commit of this group of transactions can be confirmed, and all waiting transactions in this commit group will be awakened. Among them, the reason for using fdatasync as the confirmation operation for log persistence is that in order to improve the file write performance, the database system does not directly persist the data to the disk in each write operation, but first writes the data into the system cache (called page cache in the Linux system), and then batches the binlog in the system cache and persists it to the disk through the fdatasync operation.
[0065] The above group commit mechanism is actually a process similar to the "two-phase commit". Process ② (the master node persists the binlog data to the local file) and process ③ (the master node sends the binlogdump data packet to the slave node after receiving the binlogdump request from the slave node) are asynchronous and concurrent. When both process ② and process ④ (the slave node sends the confirmation packet corresponding to the data packet in process ③ to the master node) are completed, then process ⑤ is executed, waking up all members of the group commit and returning the commit success to the client. Therefore, the overall time consumption of a group commit is the time consumption of the process with the longest execution time in process ② and process ③—>④, that is:
[0066] Ttotal=max(T2,T3+T4)
[0067] Among them, Ttotal represents the overall time consumption of a group commit transaction, and T2, T3, and T4 represent the time consumptions of process ②, ③, and ④ respectively. Generally, the overall time consumption of process ③ and ④ is significantly greater than that of process ②. Thus, the above formula can be approximated as:
[0068] Ttotal≈T3+T4
[0069] When there are multiple slave nodes, a binlogdump thread will be created between the master node and each slave node, and the T3 + T4 time depends on the slave node with the slowest ack response speed among all slave nodes. Generally speaking, the more slave nodes there are, the worse the master-slave replication performance will be.
[0070] Continue to refer to Figure 5 , which shows the data synchronization schematic diagram of a third-party real-time data publishing tool.
[0071] Third-party publish-subscribe middleware generally simulates the master-slave replication protocol of the MySQL database system (that is, virtualizes a replica), and dumps the collected binlog data to a message queue or distributed storage. Taking Canal as an example, the subscription end needs to run a client program to pull the binlog data dumped by Canal from the message queue, parse it, and then replay it to the MySQL slave instance.
[0072] Canal server simulates the behavior of a MySQL slave, pulls binlog data from the master node, and dumps it to Kafka (a mainstream open-source message queue). At the subscription end, a client program needs to be run, which is responsible for subscribing to binlogs from Kafka, parsing the binlog content, and replaying the update operations to the target database. To replay to the target database, developers need to provide an adapter program in the client that is adapted to the target database; that is, a program that can connect to the target database and perform corresponding update operations in that database. Canal is developed based on MySQL and comes with a client program adapted to the MySQL database system.
[0073] For the pre-written log data in the target storage system, if the method of introducing a third-party publish-subscribe middleware is used to synchronize it to the slave node, not only does it need to introduce new components (Canal client and adapter program) into the system, increasing the operation and maintenance burden; but also it makes the slave node unable to run in the "read-only" mode, with the risk of master-slave inconsistency.
[0074] In this embodiment, the target storage system can be a system with data storage functions. By taking the target storage system as the virtual master node in the database, the purpose is to simulate the native master-slave replication protocol of the database system, so that the slave node can still be set to the read-only mode, ensuring data consistency between the target storage system and the slave node. The database involved can be a database related to the synchronization process of pre-written log data, such as databases like MySQL, MariaDB, Percona Server for MySQL, and PostgreSQL.
[0075] Write-Ahead Logging (WAL) is a logging mechanism widely used in database systems. Its core concept is to record data update operations in the log file before actually updating the data file. For a database, the data update content is written both to the data file (where there will be a large number of scattered writes) and to the log file (the log file is an append-only file with good write performance). When committing a write transaction, only the log file needs to be persisted (this is the "write-ahead"), and then the data file can be asynchronously persisted at an opportune time. When committing a write transaction, there is no need to perform a persistence operation on the data file because the data file can be repaired by replaying the log file. In this embodiment, only the pre-written log data needs to be synchronized, and there is no need to synchronize the data file because the data update operations can be replayed based on the pre-written log data to obtain the data corresponding to the data file. For example, the pre-written log data is binlog data.
[0076] A database generally includes multiple nodes, such as a master node (master database) and multiple slave nodes (slave databases). The target storage system in this embodiment can be used as a slave node of the master node to obtain pre-written log data from the master node for data synchronization. For the target storage system, it can be used as a virtual master node to establish a virtual connection with the slave node corresponding to the target storage system. Among them, the slave node corresponding to the target storage system is generally different from the slave node corresponding to the master node.
[0077] In this embodiment, during the startup process, the slave node can take the target storage system as the virtual master node to simulate the handshake connection process with the actual master node and establish a connection between the two.
[0078] The target storage system can be any storage system with data storage capabilities. The target storage system can be a distributed storage system and / or a message queue in a third-party storage system. A distributed storage system is a storage method that disperses data storage across multiple physical nodes and enables unified management and access through network connections, featuring advantages such as high availability, high performance, and high scalability. Examples include HDFS (Hadoop Distributed File System), Ceph, XDFS, etc. A message queue (Message Queue) is a cross-process communication mechanism used to asynchronously transfer data between different applications or different components of the same application. It allows communication between producers and consumers in a decoupled manner, meaning that the producer does not need to know the specific implementation of the consumer and does not need to wait for the consumer to process the data before sending the next piece of data. Examples include RabbitMQ, Kafka, ActiveMQ, etc.
[0079] In this implementation manner, the above-mentioned execution entity can use the distributed storage system and / or message queue in the third-party storage system as the virtual master node.
[0080] Step 202, determine the synchronized pre-written log data in the slave node.
[0081] In this embodiment, the above-mentioned execution entity determines the synchronized pre-written log data in the slave node.
[0082] As an example, the above-mentioned execution entity can determine the synchronized pre-written log data in the slave node based on the Figure 3 、 4 shown data processing process. Specifically, the slave node sends a binlog dump request to the target storage system through a virtual connection. The request includes the data position of the synchronized pre-written log data in the slave node. The simulated sending function injected into the database parses the binlog dump request to determine the synchronized pre-written log data in the slave node of the database.
[0083] Step 203, perform data synchronization between the virtual master node and the slave node according to the synchronized pre-written log data.
[0084] In this implementation manner, the above-mentioned execution entity can perform data synchronization between the virtual master node and the slave node according to the synchronized pre-written log data.
[0085] As an example, the above-mentioned execution entity can be based on the above Figure 3 、 4Perform data synchronization between the virtual master node and the slave node according to the processes ③ and ④ shown. Specifically, after the simulated sending function injected into the database determines that the pre-written log data has been synchronized, it determines the unsynchronized pre-written log data other than the synchronized pre-written log data, and downloads the unsynchronized pre-written log data during the virtual receiving process, encapsulates it into a master-slave replication protocol packet, returns it to the caller (IO thread) of the receiving function, and determines that the synchronization operation is completed after receiving the ack response from the slave node.
[0086] Continue to refer to Figure 6 , Figure 6 is a schematic diagram 600 of the application scenario of the data synchronization method according to this embodiment. In Figure 6 In the application scenario, the user 601 issues a data update operation to the master node 603 of the database server through the terminal device 602. The master node 603 uses the group commit mechanism to determine the pre-written log data to be synchronized, and synchronizes it to the target storage system 604 with a larger data storage capacity for backup. For the slave node (subscription end) 606 of the target storage system 604, the target storage system is used as the virtual master node, and a connection is established between the slave node and the virtual master node; the synchronized pre-written log data in the slave node is determined; data synchronization is performed between the virtual master node and the slave node according to the synchronized pre-written log data.
[0087] The method provided by the above embodiment of the present application, by using the target storage system as the virtual master node, establishing a connection between the slave node and the virtual master node; determining the synchronized pre-written log data in the slave node; performing data synchronization between the virtual master node and the slave node according to the synchronized pre-written log data, thus simulating the native master-slave replication mechanism of the database, without introducing new components between the target storage system and the slave node, can realize the data synchronization process of using the target storage system as the virtual master node to the slave node, and is perfectly integrated with the native master-slave replication mechanism of the database, without introducing additional high-availability mechanisms, and the overall availability of the system is higher; the slave node can be set to read-only mode, ensuring data consistency between the slave node and the target storage system.
[0088] Continue to refer to Figure 7 , which shows a schematic diagram of the interaction process between the target storage system and the slave node.
[0089] In some implementation manners of this embodiment, within the process of the slave node (i.e., the subscriber), a virtual master node is simulated by modifying system library functions related to network I / O, and the data downloaded from the target storage system is encapsulated into data packets of the MySQLbinlogdump command and passed to the binlog I / O thread of the slave node. Among them, the system library functions are, for example, the connect (connection) function, the send (send function) function, the poll (polling) function, the recv (receiving) function, etc.
[0090] For the modification of system library functions, dynamic library injection and function hook technologies are adopted, which can ensure that there is no need to intrude into the source code of the MySQL database system. Similarly, open-source memory allocators jemalloc and tcmalloc also use similar technologies to implement hooks for functions such as malloc / free.
[0091] Continue to refer to Figure 8 , which shows a schematic diagram of the replacement principle under the Linux system. In the Linux system, functions declared with the "weak symbol" attribute in an application can be overridden by the same-named "strong symbol"; in addition, during the application startup phase, a dynamic link library to_inject.so can be pre-loaded into the program through the "LD_PRELOAD" mechanism. The C runtime library (glibc) in the Linux system generally declares library functions in the form of weak symbols; if a strong symbol with the same name is implemented in a dynamic link library and this dynamic library is loaded through "LD_PRELOAD" when the application starts, the native weak symbol function of glibc can be replaced by the strong symbol function implemented in the dynamic library.
[0092] Using the above technologies, a dynamic link library is implemented and loaded at the startup of MySQL through the LD_PRELOAD method, so as to achieve the override of the Linux native network I / O related library functions (such as connect, poll, send, recv, etc.). Assuming that the name of this dynamic link library is "my_subscriber.so", then starting MySQL with the following command can achieve dynamic library injection and system function replacement:
[0093] [Linux] LD_PRELOAD= / your_library_dir / my_subscriber.so mysqld --defaults-file=my.cnf &
[0094] Continue to refer to Figure 9, showing a schematic diagram of the functional modules of the data synchronization system.
[0095] The data synchronization system includes the following functional modules:
[0096] Network IO related library function hooks: used to override C runtime library (glibc) library functions related to file operations (such as: connect, poll, send, recv, etc.), so as to simulate various behaviors such as connection, polling, sending, and receiving related to the MySQL master node.
[0097] Virtual socket: The socket does not actually send and receive data in the network, but is only used for hooking various operations of the master-slave replication connection.
[0098] Virtual master node information: used to maintain the virtual master node address, publisher information, client context of the publishing system, binlog data cache, etc.
[0099] Client of the target storage system: used for data interaction between the target storage system and the slave node, and different client programs (usually provided by third-party vendors) can be used to adapt to different target storage systems (such as: Amazon S3, Alibaba Cloud OSS, Kafka, etc.).
[0100] The main process of the IO thread of the native MySQL database system is as follows: 1. The slave node calls the connect function to establish a connection with the master node, and according to the connection protocol of the MySQL system, executes the connection handshake step (that is, repeatedly sends and receives command packets for connection handshake); 2. Calls the send function to send a binlogdump request to the master node; 3. Loops to call the poll function and the recv function. If it is found that the binlog data transmitted by the master node, the binlog data is received and dumped to the relay log.
[0101] Continue to refer to Figure 10 , showing the interaction flow chart between the target storage system and the slave node. In this implementation, a special network address can be set to identify the virtual master library (such as Figure 10 The address shown: 1.1.1.1:1111). When starting the slave node through the "start slave" command of the MySQL system, the master node address of the master node is used as a parameter to call the connect library function. In the connect hook function, as long as the address parameter is compared with the address of the virtual master node, it can be judged whether it is necessary to hook the current socket.
[0102] On the slave node, the function hooks corresponding to each library function are all called in the binlog IO thread.
[0103] As Figure 10 shown, after using the hook method, the operation process of the IO thread will be adjusted as follows (the connection handshake step is ignored here because the virtual master node will not actually create a connection with the slave node, and only the response packet of the command packet needs to be simulated):
[0104] 1. Call the connect function to create a connection to the virtual master library:
[0105] Step ①: Compare the preset network parameters with the master node address of the virtual master node. If they are the same, it means connecting to the virtual master node;
[0106] Step ②: Connect to the publisher (a third-party storage system, such as a distributed object storage system or a message queue);
[0107] Step ③: If the connection to the publisher is successful, mark the current socket as the socket corresponding to the virtual master node.
[0108] 2. Call the send function to simulate sending a binlogdump request:
[0109] Step ④: The master node parses and caches the binlogdump request (including the starting position of the binlog and the GTID of the logs that have been landed on the slave node).
[0110] 3. Call the poll function to wait for incremental binlog data packets:
[0111] Step ⑤: Detect whether there is new binlog data on the publisher (the unsynchronized pre-written log data corresponding to the slave node).
[0112] 4. Call the recv function to receive new binlog data packets:
[0113] Step ⑥: Download the incremental binlog data from the publisher (if it is the first download, it needs to be merged with the content of the binlogdump data packet, and only the data that has not been landed locally is retained), and cache it.
[0114] Step ⑦: Copy the data from the binlog cache to the receive cache of recv.
[0115] It should be noted that in the above process, the poll function and the recv function need to be called repeatedly until the slave node is stopped.
[0116] Continue to refer to Figure 11 , which shows the flow diagram of the connect hook function (the function hook corresponding to the connection function).
[0117] Corresponding to the above steps ①-③, in some optional implementation manners of this embodiment, the above execution entity may execute the above step 201 in the following manner: in response to the slave node calling the original connection function, call the new connection function through the function hook corresponding to the original connection function, and in response to the preset network parameters being consistent with the master node address of the virtual master node, establish a virtual connection between the slave node and the virtual master node.
[0118] After establishing the virtual connection, a mapping relationship between the socket in the virtual connection and the virtual master node will be established.
[0119] Among them, the original connection function is the connect function in the system library function, which is used to establish a connection between the slave node and the actual master node. During the startup process of the application program of the slave node, the dynamic link library including the new connection function is loaded.
[0120] In this implementation manner, the process of replacing the original connection function with the new connection function adopts dynamic library injection and function hook technology, so that the replacement of the connection function is realized without invading the source code of the database system, and the convenience and flexibility of the function replacement process are improved.
[0121] In some optional implementation manners of this embodiment, the above execution entity may also execute the establishment process of the virtual connection in the following manner: First, according to the node type of the virtual master node, create a client adapted to the virtual master node in the slave node; then, through the client, establish a real connection between the slave node and the virtual master node.
[0122] In this implementation manner, for different types of virtual master nodes, different clients need to be established. Taking the node type of the virtual master node being a distributed storage system as an example, it is necessary to create a client adapted to the distributed storage system in the slave node to establish a connection between the slave node and the distributed storage system through the client.
[0123] Taking the node type of the virtual master node being a message queue as an example, it is necessary to create a client adapted to the message queue in the slave node to establish a connection between the slave node and the message queue through the client.
[0124] It should be noted that two connections are established here: one is the physical connection between the client and the target storage system, and the other is the virtual master-slave synchronization connection between the virtual master node and the slave node. Two network sockets will be generated respectively on the slave node. During the data reception process on the slave node, first read the data through the physical connection between the client and the target storage system, and then convert the data format into the database native master-slave synchronization data packet format and return it to the caller of the receive function (IO thread). The IO thread does not know the existence of the virtual master node and it thinks that the virtual connection is the actual master-slave replication connection.
[0125] In this implementation manner, a client adapted to the virtual master node is created according to the node type of the virtual master node, so as to establish a reliable and stable connection between the virtual master node and the slave nodes.
[0126] In this implementation manner, in response to the preset network parameters being consistent with the master node address of the virtual master node, the original connection function of the slave node is used for processing. During the creation of the virtual connection, if a connection failure occurs, an error code related to the virtual connection needs to be generated according to the error type to prompt relevant technical personnel to perform exception handling.
[0127] Continue to refer to Figure 12 , which shows a schematic flow diagram of the send hook function (the function hook corresponding to the send function).
[0128] Corresponding to step ④ above, in some optional implementation manners of this embodiment, the above-mentioned execution subject may execute the above-mentioned step 201 in the following manner:
[0129] First, in response to the slave node calling the original send function, the new send function is called through the function hook corresponding to the original send function, and a dump request is sent to the virtual master node.
[0130] Among them, the original send function is the send function in the system library function, which is used to send a dump request (binlog dump request) to the actual master node corresponding to the slave node. During the startup process of the application program of the slave node, the dynamic link library including the new send function is loaded.
[0131] In this implementation manner, it is determined whether the currently connected socket is a socket having a mapping relationship with the virtual master node; if so, in response to the slave node calling the original send function, the new send function is called through the function hook corresponding to the original send function, and a dump request is sent to the virtual master node; if not, information processing is performed through the original connection function corresponding to the slave node; in response to determining that the request sent by the slave node to the virtual master node is not a dump request, it is necessary to simulate the connection handshake protocol so that the slave node and the virtual master node are in a virtual connection state.
[0132] Second, the dump request is parsed to determine the pre-written log data that has been synchronized by the slave node relative to the virtual master node.
[0133] Specifically, the slave node may send a binlog dump request to the virtual master node through the virtual connection, and the virtual send function parses the binlog dump request to determine the pre-written log data that has been synchronized by the slave node relative to the virtual master node.
[0134] In this implementation manner, the process of replacing the original sending function with the new sending function adopts dynamic library injection and function hooking techniques, enabling the replacement of the sending function without intruding into the source code of the database system, and improving the convenience and flexibility of the function replacement process.
[0135] Corresponding to the above steps ⑤ - ⑦, in some optional implementation manners of this embodiment, the above execution subject may execute the above step 203 in the following manner:
[0136] First, in response to the slave node calling the original polling function, the new polling function is called through the function hook corresponding to the original polling function to determine whether the virtual master node includes unsynchronized pre-written log data other than the synchronized pre-written log data.
[0137] Continue to refer to Figure 13 , which shows the flowchart of the poll hook function (the function hook corresponding to the polling function).
[0138] Among them, the original polling function is the poll function in the system library function, which is used to determine whether unsynchronized pre-written log data sent by the actual master node is received in the receive buffer of the network connection between the slave node and the actual master node. During the startup process of the application program of the slave node, the dynamic link library including the new polling function is loaded.
[0139] Second, in response to determining that there is unsynchronized pre-written log data and the slave node calls the original receiving function, the new receiving function is called through the function hook corresponding to the original receiving function, and the unsynchronized pre-written log data is returned to the caller of the receiving function, that is, the binlog IO thread.
[0140] Continue to refer to Figure 14 , which shows the flowchart of the recv hook function (the function hook corresponding to the receiving function).
[0141] Among them, the original receiving function is the recv function in the system library function, which is used to receive data from the actual master node. During the startup process of the application program of the slave node, the dynamic link library including the new receiving function is loaded.
[0142] In this implementation manner, the process of replacing the original polling function with the new polling function and the process of replacing the original receiving function with the new receiving function adopt dynamic library injection and function hooking techniques, enabling the replacement of the polling function and the receiving function without intruding into the source code of the database system, and improving the convenience and flexibility of the function replacement process.
[0143] Continue to refer to Figure 13, in some alternative implementation manners of this embodiment, the above-mentioned execution entity may determine whether the virtual master node includes unsynchronized pre-written log data other than the synchronized pre-written log data in the following manner:
[0144] In response to determining that the virtual master node supports the polling operation, determine whether the virtual master node includes unsynchronized pre-written log data through the polling interface of the virtual master node.
[0145] In this manner, directly determine whether the virtual master node includes unsynchronized pre-written log data through the polling interface.
[0146] In response to determining that the virtual master node does not support the polling operation, determine whether the virtual master node includes unsynchronized pre-written log data according to the log file indicated by the latest generated metadata in the virtual master node.
[0147] Metadata is used to store the correspondence between the data publisher and the start and end points of the data GTID (Global Transaction Identifier). Among them, the data publisher is represented by Server-UUID (Universally Unique Identifier).
[0148] In this implementation manner, different methods can be used to determine whether the virtual master node includes unsynchronized pre-written log data in different situations, which improves the flexibility of the data determination process.
[0149] Continue to refer to Figure 14 , in some alternative implementation manners of this embodiment, the above-mentioned execution entity may synchronize the unsynchronized pre-written log data to the slave node in the following manner: First, store the unsynchronized pre-written log data in the receive buffer; then, dump the unsynchronized pre-written log data in the receive buffer as the relay log of the slave node.
[0150] Specifically, determine whether the currently connected socket is a socket mapped to the virtual master node; in response to determining that it is, determine whether the virtual master node and the slave node are in the handshake connection stage; in response to determining that it is, determine whether the slave node is receiving the binary log file from the virtual master node for the first time; in response to determining that it is, download the binary log file from the virtual master node, determine the data position indicated by the binlog dump request sent by the slave node to the virtual master node from the binary log file, and then remove the GTIDs included in the binlog dump from the binary log file. Place the remaining binary log file (unsynchronized pre-written log data) into the receive buffer of the virtual connection, and then dump it into a relay log. In the slave node, subsequent binlog log data on the slave node will be generated based on the replay process of the relay log.
[0151] In response to determining that the currently connected socket is not a socket mapped to the virtual master node, perform information processing through the original receive function corresponding to the slave node. In response to determining that the virtual master node and the slave node are in the handshake connection stage, simulate the handshake connection protocol to complete the connection handshake process between the virtual master node and the slave node.
[0152] In response to determining that the receive buffer of the virtual connection is empty, directly download data from the target storage system and return it to the caller of the function hook of the original receive function (binlog IO thread); in response to determining that the receive buffer of the virtual connection is not empty, return the data in the receive buffer of the virtual connection to the caller of the function hook of the original receive function (binlog IO thread).
[0153] Continue to refer to Figure 15 , which shows a schematic flowchart 1500 of another embodiment of the data synchronization method according to the present application, including the following steps:
[0154] Step 1501, in response to the slave node calling the original connection function, call the new connection function through the function hook corresponding to the original connection function. In response to the preset network parameters being consistent with the master node address of the virtual master node, establish a virtual connection between the slave node and the virtual master node.
[0155] Among them, the original connection function is used to establish a connection between the slave node and the actual master node. During the startup process of the application program of the slave node, the dynamic link library including the new connection function is loaded.
[0156] Step 1502, in response to the slave node calling the original send function, call the new send function through the function hook corresponding to the original send function, and send a dump request to the virtual master node.
[0157] Among them, the original sending function is used to send a dump request to the actual master node corresponding to the slave node. During the startup process of the application of the slave node, the dynamic link library including the new sending function is loaded.
[0158] Step 1503: Parse the dump request to determine the pre-written log data that has been synchronized by the slave node relative to the virtual master node.
[0159] Step 1504: In response to the slave node calling the original polling function, call the new polling function through the function hook corresponding to the original polling function to determine whether the virtual master node includes unsynchronized pre-written log data other than the pre-written log data that has been synchronized.
[0160] Among them, the original polling function is used to determine whether unsynchronized pre-written log data sent by the actual master node is received in the receive buffer of the network connection between the slave node and the actual master node. During the startup process of the application of the slave node, the dynamic link library including the new polling function is loaded
[0161] Step 1505: In response to determining that there is unsynchronized pre-written log data and the slave node calls the original receiving function, call the new receiving function through the function hook corresponding to the original receiving function to synchronize the unsynchronized pre-written log data to the slave node.
[0162] Among them, the original receiving function is used to receive data from the actual master node. During the startup process of the application of the slave node, the dynamic link library including the new receiving function is loaded.
[0163] It can be seen from this embodiment that compared with Figure 2 the corresponding embodiment, the process 1500 of the data synchronization method in this embodiment specifically illustrates the replacement process of the connection function, the sending function, the polling function, and the receiving function, so as to simulate the native master-slave replication mechanism of the database and realize the data synchronization process from the target storage system as the virtual master node to the slave node, enabling the slave node to be set to the read-only mode, and ensuring the data consistency between the slave node and the target storage system while ensuring the data storage effect of the target storage system.
[0164] Continuing to refer to Figure 16 , as an implementation of the methods shown in the above figures, an embodiment of a data synchronization device is provided in the present application. This device embodiment corresponds to Figure 2 the method embodiment shown, and this device can be specifically applied to various electronic devices.
[0165] As shown in Figure 16As shown, a data synchronization device includes: a connection unit 1601 configured to establish a connection between a slave node and a virtual master node with the target storage system as the virtual master node; a determination unit 1602 configured to determine the pre-written log data that has been synchronized in the slave node; and a synchronization unit 1603 configured to perform data synchronization between the virtual master node and the slave node according to the pre-written log data that has been synchronized.
[0166] In some examples, the above-mentioned connection unit 1601 is further configured to: in response to the slave node calling the original connection function, call the new connection function through the function hook corresponding to the original connection function, and establish a virtual connection between the slave node and the virtual master node in response to the preset network parameters being consistent with the master node address of the virtual master node, where the original connection function is used to establish a connection between the slave node and the actual master node, and during the startup process of the application program of the slave node, the dynamic link library including the new connection function is loaded.
[0167] In some examples, the above-mentioned connection unit 1601 is further configured to: create a client adapted to the virtual master node in the slave node according to the node type of the virtual master node; and establish a real connection between the slave node and the virtual master node through the client.
[0168] In some examples, the above-mentioned determination unit 1602 is further configured to: in response to the slave node calling the original sending function, call the new sending function through the function hook corresponding to the original sending function, and send a dump request to the virtual master node, where the original sending function is used to send a dump request to the actual master node corresponding to the slave node, and during the startup process of the application program of the slave node, the dynamic link library including the new sending function is loaded; parse the dump request to determine the pre-written log data that has been synchronized in the slave node relative to the virtual master node.
[0169] In some examples, the above-mentioned synchronization unit 1603 is further configured to: in response to the slave node calling the original polling function, call the new polling function through the function hook corresponding to the original polling function to determine whether the virtual master node includes pre-written log data that has not been synchronized other than the pre-written log data that has been synchronized, where the original polling function is used to determine whether the receive buffer of the network connection between the slave node and the actual master node has received the pre-written log data that has not been synchronized sent by the actual master node, and during the startup process of the application program of the slave node, the dynamic link library including the new polling function is loaded; in response to determining that there is pre-written log data that has not been synchronized and the slave node calls the original receiving function, call the new receiving function through the function hook corresponding to the original receiving function, and synchronize the pre-written log data that has not been synchronized to the slave node, where the original receiving function is used to receive data from the actual master node, and during the startup process of the application program of the slave node, the dynamic link library including the new receiving function is loaded.
[0170] In some examples, the above-mentioned synchronization unit 1603 is further configured to: in response to determining that the virtual master node supports the polling operation, determine whether the virtual master node includes unsynchronized pre-written log data through the polling interface of the virtual master node; in response to determining that the virtual master node does not support the polling operation, determine whether the virtual master node includes unsynchronized pre-written log data according to the log file indicated by the latest generated metadata in the virtual master node.
[0171] In some examples, the above-mentioned synchronization unit 1603 is further configured to: store the unsynchronized pre-written log data into the receive buffer; dump the unsynchronized pre-written log data in the receive buffer as the relay log of the slave node.
[0172] In some examples, the above-mentioned connection unit 1601 is further configured to: use the distributed storage system and / or message queue in the third-party storage system as the virtual master node.
[0173] In this embodiment, the connection unit in the data synchronization device uses the target storage system as the virtual master node to establish a connection between the slave node and the virtual master node; the determination unit determines the synchronized pre-written log data in the slave node; the synchronization unit performs data synchronization between the virtual master node and the slave node according to the synchronized pre-written log data, thereby simulating the native master-slave replication mechanism of the database. Without introducing new components between the target storage system and the slave node, the data synchronization process of using the target storage system as the virtual master node to the slave node can be realized, which is perfectly integrated with the native master-slave replication mechanism of the database. Without introducing additional high-availability mechanisms, the overall availability of the system is higher; the slave node can be set to read-only mode to ensure data consistency between the slave node and the target storage system.
[0174] Reference is made below to Figure 17 , which shows a schematic structural diagram of a computer system 1700 suitable for use in implementing the device of the embodiments of the present application (such as Figure 1 the devices 101, 102, 103, 105 shown). Figure 17 The devices shown are merely examples and should not impose any limitation on the functions and usage scope of the embodiments of the present application.
[0175] As Figure 17As shown, computer system 1700 includes a processor (e.g., CPU, central processing unit) 1701, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 1702 or a program loaded from storage section 1708 into random access memory (RAM) 1703. In RAM 1703, various programs and data required for the operation of system 1700 are also stored. The processor 1701, ROM 1702, and RAM 1703 are connected to each other via a bus 1704. Input / output (I / O) interface 1705 is also connected to bus 1704.
[0176] The following components are connected to I / O interface 1705: an input section 1706 including a keyboard, a mouse, etc.; an output section 1707 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 1708 including a hard disk, etc.; and a communication section 1709 including a network interface card such as a LAN card, a modem, etc. The communication section 1709 performs communication processing via a network such as the Internet. A drive 1710 is also connected to I / O interface 1705 as needed. A removable medium 1711, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on drive 1710 as needed so that a computer program read from it can be installed into storage section 1708 as needed.
[0177] Specifically, according to an embodiment of the present application, the processes described above with reference to the flowcharts can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program codes for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network via communication section 1709, and / or installed from removable medium 1711. When the computer program is executed by processor 1701, the above-described functions defined in the methods of the present application are performed.
[0178] It should be noted that the computer-readable medium of the present application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. In the present application, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, in which computer-readable program code is carried. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and this computer-readable medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.
[0179] The computer program code for performing the operations of the present application can be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages - such as Java, Smalltalk, C++, and also include conventional procedural programming languages - such as the "C" language or similar programming languages. The program code can be executed entirely on the client computer, partially on the client computer, executed as an independent software package, partially on the client computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the client computer through any type of network - including a local area network (LAN) or a wide area network (WAN) - or, alternatively, can be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).
[0180] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of devices, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0181] The units involved in the embodiments described in the present application can be implemented in software or in hardware. The described units can also be provided in a processor. For example, it can be described as: a processor including a connection unit, a determination unit, and a synchronization unit. Among them, the names of these units do not constitute a limitation on the units themselves in some cases. For example, the connection unit can also be described as "a unit that uses the target storage system as a virtual master node and establishes a connection between the slave node and the virtual master node".
[0182] On the other hand, the present application also provides a computer-readable medium, which can be included in the devices described in the above embodiments; or can exist separately without being assembled into the device. The above computer-readable medium carries one or more programs, and when the above one or more programs are executed by the device, the computer device is caused to: use the target storage system as a virtual master node and establish a connection between the slave node and the virtual master node; determine the synchronized pre-written log data in the slave node; and perform data synchronization between the virtual master node and the slave node according to the synchronized pre-written log data.
[0183] The above description is only a preferred embodiment of the present application and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of the invention involved in the present application is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, technical solutions formed by mutually replacing the above features with (but not limited to) technical features with similar functions disclosed in the present application.
Claims
1. A data synchronization method, comprising: Taking the target storage system as a virtual master node, establishing a connection between a slave node and the virtual master node, including: in response to the slave node calling an original connection function, calling a new connection function through a function hook corresponding to the original connection function, and in response to the preset network parameters being consistent with the master node address of the virtual master node, establishing a virtual connection between the slave node and the virtual master node, wherein the original connection function is used to establish a connection between the slave node and the actual master node, and during the startup process of the application of the slave node, a dynamic link library including the new connection function is loaded; the virtual master node stores pre-written log data to be synchronized from the actual master node; Determining the synchronized pre-written log data in the slave node; Data synchronization is performed between the virtual master node and the slave node according to the synchronized pre-written log data.
2. The method according to claim 1, wherein: The establishing of a connection between the slave node and the virtual master node further includes: According to the node type of the virtual master node, creating a client adapted to the virtual master node in the slave node; A real connection between the slave node and the virtual master node is established through the client.
3. The method according to claim 1, wherein: The determining the synchronized pre-written log data in the slave node includes: In response to the slave node calling the original sending function, calling the new sending function through the function hook corresponding to the original sending function, sending a dump request to the virtual master node, wherein the original sending function is used to send the dump request to the actual master node corresponding to the slave node, and during the startup process of the application program of the slave node, a dynamic link library including the new sending function is loaded; The dump request is parsed to determine the synchronized pre-written log data of the slave node relative to the virtual master node.
4. The method according to claim 1, wherein: The step of performing data synchronization between the virtual master node and the slave node according to the synchronized pre-written log data includes: In response to the slave node calling the original polling function, calling the new polling function through the function hook corresponding to the original polling function, determining whether the virtual master node includes unsynchronized pre-written log data other than the synchronized pre-written log data, wherein the original polling function is used to determine whether the receiving cache of the network connection between the slave node and the actual master node has received the unsynchronized pre-written log data sent by the actual master node, and during the startup process of the application of the slave node, the dynamic link library including the new polling function is loaded; In response to determining that the unsynchronized pre-written log data is included and the slave node calls the original receive function, a new receive function is called through a function hook corresponding to the original receive function to synchronize the unsynchronized pre-written log data to the slave node, wherein the original receive function is used to receive data from the actual master node, and during the startup process of the application of the slave node, a dynamic link library including the new receive function is loaded.
5. The method according to claim 4, wherein: The determining whether the virtual master node includes unsynchronized pre-written log data other than the synchronized pre-written log data comprises: In response to determining that the virtual master node supports a polling operation, determining, through a polling interface of the virtual master node, whether the virtual master node includes the unsynchronized pre-written log data; In response to determining that the virtual master node does not support the polling operation, determining whether the virtual master node includes the unsynchronized pre-written log data according to a log file indicated by the metadata most recently generated in the virtual master node.
6. The method according to claim 4, wherein: The step of synchronizing the unsynchronized pre-written log data to the slave node includes: Storing the unsynchronized pre-written log data in the receiving buffer; The unsynchronized pre-written log data in the receiving buffer is dumped as a relay log of the slave node.
7. The method according to claim 1, wherein: The step of using the target storage system as a virtual master node includes: The distributed storage system and / or the message queue in the third-party storage system is used as the virtual master node.
8. A data synchronization device, comprising: A connection unit is configured to use the target storage system as a virtual master node and establish a connection between a slave node and the virtual master node, including: in response to the slave node calling an original connection function, calling a new connection function through a function hook corresponding to the original connection function, and in response to a preset network parameter being consistent with a master node address of the virtual master node, establishing a virtual connection between the slave node and the virtual master node, wherein the original connection function is used to establish a connection between the slave node and an actual master node, and during the startup process of an application program of the slave node, a dynamic link library including the new connection function is loaded; the virtual master node stores pre-written log data to be synchronized from the actual master node; a determining unit configured to determine the synchronized pre-written log data in the slave node; The synchronization unit is configured to perform data synchronization between the virtual master node and the slave node according to the synchronized pre-written log data.
9. A computer readable medium having a computer program stored thereon, wherein: When the program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
10. An electronic device comprising: one or more processors; a storage device having one or more programs stored thereon, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 7.
11. A computer program product comprising: A computer program which, when executed by a processor, implements the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Database log synchronization method and device, computer equipment and readable storage medium
CN111008246A
Data synchronization method and device, electronic equipment and computer readable storage medium
CN117609381A