Data synchronization method and device, communication node and storage medium
By creating a persistent file in the storage node to store the replication backlog buffer data, and reading the data from the file at the master node end to send it to the slave node, the problem of replication backlog buffer full when the network is unstable is solved, and efficient partial data synchronization is achieved, reducing resource consumption and improving communication efficiency.
Patent Information
- Application Number
- CN202510121215.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-27
AI Technical Summary
In the unstable or complex network between master and slave nodes, the prior art when the main node update traffic pressure is huge, the fixed-sized replication backlog buffer is easily filled, resulting in only full data synchronization, waste of resources and low communication efficiency.
By creating a persistent file in the storage node, it is used to store data and related information of the replica backlog buffer, including identification, cluster identification, location information and time information. When the master node does not contain the requested data in the replication backlog buffer, it reads the data from the file and sends it to the slave node to realize partial data synchronization.
The frequency of full data synchronization is reduced, resource consumption is reduced, and communication efficiency between master and slave nodes is improved, avoiding the inefficient process of full synchronization after the overflow of replication backlog buffers.
Smart Images

Figure CN120045569A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to database technology, and in particular to a data synchronization method, apparatus, communication node, and storage medium. Background Art
[0002] Remote Dictionary Server can be abbreviated as Redis. The service mode of Redis includes a master-slave mode which includes a master node (also called a master node) and a slave node (also called a slave node). In this master-slave mode, a series of data is stored on the master node, and the data stored on the master node needs to be replicated to the slave node; subsequently, whenever there is an update to the data on the master node, the updated data needs to be synchronized to the slave node.
[0003] Currently, the slave node obtains the full amount of data on the master node through the full synchronization (Sync) command, so as to keep the data of the master and slave nodes consistent. However, when the connection between the master and slave nodes is interrupted, after the connection between the master and slave nodes is restored, the slave node will continue to obtain the full amount of data on the master node through the Sync command, which will cause waste of resources. To solve this problem, a new partial synchronization (Psync) command is proposed to achieve partial data synchronization. This solution requires the master node to generate a replication backlog buffer of a fixed size, and this replication backlog buffer is a circular buffer, that is, after the buffer is full of data, new data will continue to be written from the starting position of the buffer.
[0004] With the emergence of kubernete and cloud scenarios, the complexity of the network structure has increased. There are situations where there are differences in the network between the master and slave nodes and the network is unstable; in addition, in the case of a high query per second (QPS) of redis, the update traffic pressure on the master node is huge, and the replication backlog buffer of a fixed size will soon be full. In this case, full synchronization has to be performed between the master node and the slave node. Summary of the Invention
[0005] To solve the existing technical problems, embodiments of the present invention provide a data synchronization method, apparatus, electronic device, and storage medium.
[0006] To achieve the above object, the technical solution of the embodiments of the present invention is implemented as follows:
[0007] In a first aspect, an embodiment of the present invention provides a data synchronization method, which is applied to a first node. The method includes: the first node collects data in a replication backlog buffer, and writes the data and related information of the data into a first file in a storage node; wherein, the related information of the data includes: a first identifier representing the replication backlog buffer, a second identifier representing the cluster to which the first node belongs, location information of the data in the replication backlog buffer, and time information of the data; when receiving a first command sent by a second node for requesting partial data synchronization, in the case that the replication backlog buffer does not contain the first data requested by the first command, the first data corresponding to the first command is read from the first file in the storage node, and the first data is sent to the second node through a second command. The first node is the master node, and the second node is the slave node.
[0008] In a second aspect, an embodiment of the present invention further provides a data synchronization method, which is applied to a storage node. The method includes: the storage node receives the data sent by the first node and the related information of the data, and writes the data and the related information of the data into a first file; the related information of the data includes: a first identifier of the replication backlog buffer of the first node, a second identifier of the cluster to which the first node belongs, location information of the data in the replication backlog buffer, and time information of the data; the storage node reads the first data requested by the command from the first file based on the command of the first node, and sends the first data to the first node for the first node to send the first data to the second node. The first node is the master node, and the second node is the slave node.
[0009] In a third aspect, an embodiment of the present invention further provides a data synchronization method, which is applied to a second node. The method includes: the second node sends a first command to the first node, and the first command is used to request partial data synchronization; the first node is the master node, and the second node is the slave node; receive a second command sent by the first node, and the second command includes first data, and the first data is obtained by the first node reading from a first file in a storage node; the first file pre-stores the data in the replication backlog buffer written by the first node, and the data stored in the first file at least includes: a first identifier representing the replication backlog buffer, a second identifier representing the cluster to which the first node belongs, location information of the data in the replication backlog buffer, and time information of the data.
[0010] In a fourth aspect, an embodiment of the present invention further provides a data synchronization device, which is applied to a first node. The device includes: an acquisition unit, a loading unit, and a first communication unit; wherein,
[0011] The collection unit is used to collect the data in the replication backlog buffer and write the data and the related information of the data into a first file in the storage node; wherein, the related information of the data includes: a first identifier representing the replication backlog buffer, a second identifier representing the cluster to which the first node belongs, the position information of the data in the replication backlog buffer, and the time information of the data.
[0012] When the loading unit receives a first command for requesting partial data synchronization sent by a second node acting as a slave node, and when the first data requested by the first command is not included in the replication backlog buffer, the loading unit reads the first data corresponding to the first command from the first file in the storage node.
[0013] The first communication unit is used to send the first data to the second node through a second command. The first node is the master node and the second node is the slave node.
[0014] In a fifth aspect, an embodiment of the present invention further provides a data synchronization device, which is applied to a storage node. The device includes: a second communication unit, a processing unit, and a storage unit; wherein,
[0015] The second communication unit is used to receive the data sent by the first node and the related information of the data. The related information of the data includes: a first identifier of the replication backlog buffer of the first node, a second identifier of the cluster to which the first node belongs, the position information of the data in the replication backlog buffer, and the time information of the data.
[0016] The processing unit is used to write the data and the related information of the data into the first file; and is also used to read the first data requested by the command from the first file based on the command of the first node.
[0017] The second communication unit is further used to send the first data to the first node for the first node to send the first data to the second node; the first node is the master node and the second node is the slave node;
[0018] The storage unit is used to store the first file.
[0019] Sixth aspect, an embodiment of the present invention provides a data synchronization device, which is applied to a second node. The device includes: a third communication unit, configured to send a first command to a first node, where the first command is used to request partial data synchronization; and is further configured to receive a second command sent by the first node, where the second command includes first data, and the first data is obtained by the first node reading from a first file of a storage node; the first file prestores data of a replication backlog buffer written by the first node, and the data stored in the first file at least includes: a first identifier representing the replication backlog buffer, a second identifier representing a cluster to which the first node belongs, location information of data in the replication backlog buffer, and time information of the data; the first node is a master node, and the second node is a slave node.
[0020] Seventh aspect, an embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of the data synchronization method according to any one of the above aspects of the embodiments of the present invention are implemented.
[0021] Eighth aspect, an embodiment of the present invention further provides a communication node, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the data synchronization method according to any one of the above aspects of the embodiments of the present invention are implemented.
[0022] Ninth aspect, an embodiment of the present invention further provides a computer program product, including computer program instructions, and the computer program instructions cause a computer to execute the steps of the data synchronization method according to any one of the above aspects of the embodiments of the present invention.
[0023] The data synchronization method, device, electronic device, and storage medium provided by the embodiments of the present invention propose to uniquely identify the replication backlog buffer, so that the unique label of the data in the replication backlog buffer can be guaranteed regardless of system restart or buffer overflow; use the first file in the storage node to implement persistent storage of the data in the replication backlog buffer, thereby minimizing full data synchronization and reducing resource consumption; and, through a new communication command between the master node and the slave node, implement partial data synchronization by recording the corresponding data in the first file, overcome the inefficient communication process of only being able to perform full data synchronization after the replication backlog buffer overflows, and improve the communication efficiency. Description of the Drawings
[0024] Figure 1 It is a schematic diagram of the system structure to which the data synchronization method of the embodiment of the present invention is applied;
[0025] Figure 2 It is a flow diagram of the data synchronization method of the embodiment of the invention Figure 1 ;
[0026] Figure 3 Flow schematic of the data synchronization method for the invention embodiment Figure 2 ;
[0027] Figure 4 Flow schematic of the data synchronization method for the invention embodiment Figure 3 ;
[0028] Figure 5 Format schematic of the data collected by the first node in the invention embodiment;
[0029] Figure 6 Format schematic of the first file in the data synchronization method of the invention embodiment;
[0030] Figure 7 Data reading schematic of the storage node in the data synchronization method of the invention embodiment;
[0031] Figure 8 Interaction flow schematic of the data synchronization method of the invention embodiment Figure 1 ;
[0032] Figure 9 Interaction flow schematic of the data synchronization method of the invention embodiment Figure 2 ;
[0033] Figure 10 Composition structure schematic of the data synchronization device of the invention embodiment Figure 1 ;
[0034] Figure 11 Composition structure schematic of the data synchronization device of the invention embodiment Figure 2 ;
[0035] Figure 12 Composition structure schematic of the data synchronization device of the invention embodiment Figure 3 ;
[0036] Figure 13 Hardware composition structure schematic of the communication node in the invention embodiment. Detailed implementation manners
[0037] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0038] It should be understood that the term "and / or" in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this article generally represents an "or" relationship between the associated objects before and after.
[0039] The terms "first", "second", etc. in the description and claims of this application are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of this application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion.
[0040] Figure 1 It is a schematic structural diagram of the system to which the data synchronization method of the embodiment of the present invention is applied; as Figure 1 shown, the system to which the data synchronization method of this embodiment is applied at least includes: a master node (or master node), a slave node (or slave node), and a storage node; wherein, there are two types of files stored in the storage node: dump.rdb file and dump.rb file; wherein, the rdb file can also be called the Redis full amount data file (Redis database); the dump.rb file is the persistent file for incremental synchronization involved in the data synchronization method of the embodiment of the present invention, and all nodes in the cluster reuse the dump.rb file through the shared storage mode.
[0041] In this embodiment, the master node and the slave node can be physical nodes, or they can also be virtual nodes, that is, the master node can also be called the master node instance, and the slave node can also be called the slave node instance. The master node and the slave node form a cluster (the cluster can also be called a group, a cluster, or a cluster, etc.). It can be understood that there is one master node in the cluster, and there can be one or more slave nodes. Figure 1 Only one slave node is taken as an example for illustration.
[0042] Both the master node and the slave node have a loader and a collector; when the node is the master node, the loader and the collector are started; when the node is the slave node, the loader and the collector are paused, that is, only the master node performs read and write operations on the dump.rb file in the storage node. When the slave node switches to the master node, the loader and the collector are started.
[0043] Among them, the master node uses the loader to asynchronously collect the data in the replication backlog buffer and writes it into the dump.rb file in the storage node; where rb is the abbreviation of ringbuffer (circular buffer), indicating the persistent file of the replication backlog buffer. Persisting the ringbuffer has no impact on other parts of the redis system and has good compatibility.
[0044] The master node loads the corresponding data of the dump.rb file in the storage node using a loader under specific scenarios or conditions, and sends it to the slave node that needs to copy the data through a specific command. Among them, the specific command is a communication command different from the Psync command and the Sync command (such as the dsync command). Through this command, more efficient master-slave data replication can be completed, fundamentally avoiding the problem of full-volume data synchronization.
[0045] In other alternative embodiments, the system may further include a monitoring node, which is mainly responsible for monitoring the status and roles of each redis node; among them, the status of the node may include, for example: working status, downtime status, etc.; the role of the node is its role as a master node or a slave node. When the master node is started or centrally started, it will generate a first identifier of the replication backlog buffer, which can be denoted as ring_id. During the process of forming a cluster of nodes, a unique identifier of the cluster, that is, the cluster identifier, can be denoted as cluster_id. The monitoring node manages and maintains the first identifier (ring_id) of each replication backlog buffer and the cluster identifier (cluster_id), so that these two types of identifiers are shared in the corresponding cluster. When a node restarts, it can pull the first identifier (ring_id) of the corresponding replication backlog buffer from the monitoring node and save it into its own status information to ensure the uniqueness of the replication backlog buffer through this mechanism.
[0046] Based on this, an embodiment of the present invention provides a data synchronization method, and the method is applied to a first node. Figure 2 Schematic diagram of the process of the data synchronization method according to the embodiment of the invention Figure 1 ; as Figure 2 shown, the method includes:
[0047] Step 101: The first node collects the data of the replication backlog buffer, and writes the data and the related information of the data into a first file in the storage node; among them, the related information of the data includes: the first identifier representing the replication backlog buffer, the second identifier representing the cluster to which the first node belongs, the position information of the data in the replication backlog buffer, and the time information of the data.
[0048] Step 102: When receiving a first command sent by the second node for requesting partial data synchronization, in the case that the first data requested by the first command is not included in the replication backlog buffer, read the first data corresponding to the first command from the first file of the storage node, and send the first data to the second node through a second command.
[0049] Correspondingly, an embodiment of the present invention provides a data synchronization method, and the method is applied to a storage node. Figure 3 Schematic diagram of the process of the data synchronization method according to the embodiment of the inventionFigure 2 ; As Figure 3 shown, the method includes:
[0050] Step 201: The storage node receives the data sent by the first node and the related information of the data, and writes the data and the related information of the data into the first file; the related information of the data includes: the first identifier of the replication backlog buffer of the first node, the second identifier of the cluster to which the first node belongs, the position information of the data in the replication backlog buffer, and the time information of the data.
[0051] Step 202: Based on the command of the first node, the storage node reads the first data requested by the command from the first file, and sends the first data to the first node for the first node to send the first data to the second node.
[0052] Correspondingly, an embodiment of the present invention provides a data synchronization method, and the method is applied to the second node. Figure 4 The flowchart of the data synchronization method according to the embodiment of the invention Figure 3 ; As Figure 4 shown, the method includes:
[0053] Step 301: The second node sends a first command to the first node, and the first command is used to request partial data synchronization;
[0054] Step 302: Receive the second command sent by the first node, where the second command includes the first data, and the first data is obtained by the first node reading from the first file of the storage node; the data pre-stored in the first file includes at least: the first identifier representing the replication backlog buffer, the second identifier representing the cluster to which the first node belongs, the position information of the data in the replication backlog buffer, and the time information of the data.
[0055] In this embodiment, the first node is the master node, and the second node is the slave node. In other alternative embodiments, in a specific scenario, the roles of the first node and the second node may also change. For example, when the first node fails, the first node may also be switched to the slave node; the second node may also be switched from the slave node to the master node.
[0056] In this embodiment, the first file may be denoted as the dump.rb file, which is a persistent file in the storage node. The size of the first file is much larger than the size of the replication backlog buffer in the first node as the master node, so as to realize the persistent storage of the data in the replication backlog buffer of each first node.
[0057] For the first node side: In this embodiment, the first node collects data in the replication backlog buffer and writes the data and related information of the data into a first file in the storage node. Specifically, the data stored in the replication backlog buffer includes the original data and the position information of the original data in the replication backlog buffer; the first node obtains the first identifier of the replication backlog buffer, the second identifier of the cluster to which the first node belongs, and the time information of the data, and sends the above information to the storage node and writes it into the first file. Among them, the time information of the data is the generation time of the original data. The first identifier of the replication backlog buffer can be represented by a ring id. The second identifier of the cluster to which the first node belongs can be represented by a cluster id. Exemplarily, the position information of the original data in the replication backlog buffer can be represented by the data offset of the original data in the replication backlog buffer, and this position information or data offset can uniquely locate the position of the original data in the replication backlog buffer.
[0058] It should be noted that the ring id represents the identifier of the replication backlog buffer, and a value of ring_id (for example, ring1) can represent the first identifier of the replication backlog buffer of the first node. Similarly, the cluster id represents the identifier of the node cluster, and a value of cluster id (such as cluster1) can represent the identifier of the cluster to which the first node belongs.
[0059] In some alternative embodiments, the collector of the first node collects data in the replication backlog buffer and writes the data and related information of the data into a first file in the storage node. Exemplarily, Figure 5 is a schematic diagram of the format of the data collected by the first node in the embodiment of the present invention; the format of each piece of data collected by the collector of the first node is as Figure 5 shown, including: cluster identifier (cluster ID), ring identifier (ring ID), offset, opTimestamp (op timestamp), and data (original data); among them;
[0060] The cluster ID, that is, the identifier of the cluster to which the first node belongs (denoted as the second identifier), represents the unique identifier (ID) of the cluster to which the master device belongs, and is used to verify whether the saved data is generated by the nodes in this cluster;
[0061] The ring ID represents the unique identifier of the replication backlog buffer (denoted as the first identifier). Since the existing Psync command cannot identify the uniqueness of the data, the embodiment of the present invention designs a historical unique identification number to facilitate accurate positioning of the data replication position;
[0062] Offset, that is, the location information of the data in the replication backlog buffer, identifies the data offset of the Aof-data to be replicated in the current replication backlog buffer. This field can uniquely locate the location of the data in the current replication backlog buffer;
[0063] opTimestamp, i.e., the time information of the data; for example, the time when the Aof-data is generated can be combined with a 6-bit random number, and the time can be accurate to milliseconds; this field supports users to copy data by time position;
[0064] Data (Aof-data) represents the original data generated by redis.
[0065] On the storage node side: after receiving the data and the related information of the data sent by the first node, the storage node persists the data and the related information of the data into the first file.
[0066] In some optional embodiments, the first file includes an index part and a data part; the index part includes a first-level index, a second-level index and a third-level index; wherein,
[0067] The first-level index includes at least one or more groups of corresponding first fields, second fields, and third fields, the first field indicates an identifier of the replication backlog buffer, the second field indicates a position (such as represented by an offset) at which data corresponding to the identifier of the replication backlog buffer first appears in the first file, and the third field is used to indicate a position of a second-level index of data storage location information belonging to the identifier of the replication backlog buffer;
[0068] The second-level index includes at least one or more groups of corresponding fourth fields and fifth fields, the fourth field identifies the identifier of the replication backlog buffer and the location of the corresponding data in the replication backlog buffer, and the fifth field identifies the data storage location of the data belonging to the identifier of the replication backlog buffer in the first file;
[0069] The third-level index includes at least one or more groups of corresponding sixth fields and seventh fields, the sixth field identifies the time information of the data, and the seventh field identifies the data storage location corresponding to the time information in the first file.
[0070] For example, Figure 6 Schematic diagram of the format of the first file in the data synchronization method of an embodiment of the present invention; Figure 6 As shown, the first three lines are the index part of the first file, and the fourth line is the data part; the first line is the first-level index of the index part, the second line is the second-level index of the index part, and the third line is the third-level index of the index part; wherein,
[0071] The value in the ring_num field indicates how many ring values are included in the first file, that is, how much data in the replication backlog buffer is included in the first file. Exemplarily, this field can reserve 1 megabyte (M) of data to minimize the cost of frequent data movement.
[0072] ring1 and rin2 represent the values of the ring_id field, that is, the identifiers of the replication backlog buffer.
[0073] The value of the poffset field after ring1 indicates the offset (i.e., position) where the data of ring1 first appears in the first file, and the value of the index_offset field indicates the position of the second-level index corresponding to the offset of ring1.
[0074] (ring1, offset1) represents the value of the (ring ID, offset) field, where offset is the value of the offset carried in the data written by the first node; poffset represents the storage location of the data corresponding to ring1, offset1 in the first file, and this storage location can be represented by an offset.
[0075] The value of the timestamp quantity (which can also be expressed as Timestamp_num) field indicates the number of data items identified by timestamps. Exemplarily, in the embodiments of the present invention, calculations can be performed at fixed time intervals (e.g., 2 hours), and the total size can be controlled to 24, indicating that at most 2 days of data can be stored. Among them, opTimestamp1 represents the value of the opTimestamp field, and the corresponding poffset represents the storage location of the data corresponding to this timestamp in the first file, and this storage location can be represented by an offset.
[0076] (cluster_id, ring_id, offset, opTimestamp, aof-data) is the data part, directly from the first node, and the data is arranged in ascending order. Among them, cluster_id can identify the cluster to which the data belongs and has an identification function when the file is shared by subsequent clusters.
[0077] In some alternative embodiments, writing the data and related information of the data into the first file includes: writing the data and the related information of the data to the tail of the data part of the first file, and modifying the index part of the first file according to the related information of the data according to the following rules: if the first identifier is the identifier of a newly added replication backlog buffer in the first file, add a first-level index corresponding to the first identifier in the index information; the first-level index includes first index information and second index information, the first index information is used to indicate the position where the data corresponding to the first identifier first appears in the first file, and the second index information is used to indicate the position of the second-level index corresponding to the first identifier; add a second-level index corresponding to the first identifier in the index part; the second-level index includes third index information corresponding to a third identifier, the third identifier is used to represent the position information of the first identifier and the data in the replication backlog buffer, and the third index information is used to indicate the data storage position of the data corresponding to the third identifier in the first file; add a corresponding third-level index in the index part; the third-level index includes fourth index information corresponding to a fourth identifier, the fourth identifier is used to represent the time information of the data, and the fourth index information is used to indicate the data storage position of the data corresponding to the fourth identifier in the first file.
[0078] In this embodiment, for the data to be written by the first node (such as Figure 5 the data shown), the storage node can directly write the data to the tail of the data part of the first file, that is, it can directly update the data (cluster_id, ring_id, offset, opTimestamp, aof-data) to the tail of the data part of the file dump.rb. For the index part of the first file, it is necessary to add and adjust the index part according to specific circumstances. Specifically:
[0079] First, the storage node first determines whether the identifier of the replication backlog buffer (i.e., ring id) is a newly added identifier, that is, whether the index part of the first file includes the identifier of the replication backlog buffer (i.e., ring id). If the index part of the first file does not include the identifier of the replication backlog buffer (i.e., ring id), that is, if the identifier of the replication backlog buffer is a newly added identifier, add a ring index part (i.e., the first-level index) to the index part of the first file. For example Figure 6In the example shown, the added ring index part includes a first field (i.e., the ring id field), a second field (i.e., the poffset field or the first index information), and a third field (i.e., the index_offset field or the second index information); if the data to be stored is ring3, the ring id field can be added at the end of the first line, and the value of this ring id field is ring3; the corresponding poffset field and index_offset field are further added. Further, since the ring id is newly added, the value of ring_num needs to be incremented by 1. If the first identifier (i.e., the ring id) is not a newly added identifier, the value of ring_num is not processed.
[0080] Secondly, the storage node adds a second-level index in the index part according to the position information of the data in the replication backlog buffer, that is, adds a (ring ID, offset) field and the corresponding poffset.
[0081] Among them, as an optional implementation manner, the second-level index includes at least one group of third identifiers and third index information; among them, the position information of the data corresponding to each group of third identifiers represents a data size less than or equal to the first threshold.
[0082] Specifically, when the storage node determines that the size of the data to be written is greater than or equal to the first threshold according to the position information of the data in the replication backlog buffer, for example, calculates the difference between the offset1 in (Ring1, offset1) in the first file (assuming that offset1 here is the offset corresponding to the previously written Ring1) and the offset corresponding to the currently written data; if the difference is greater than 1M (i.e., the first threshold, which can be pre-configured or set), a (ring, offset) index field (i.e., the second-level index) is newly added to the index part of the first file, that is, a (ring, offset) field and the corresponding poffset field are newly added, and the values of the corresponding fields are adjusted.
[0083] Thirdly, the storage node adds a third-level index in the index part according to the time information of the data, that is, adds an opTimestamp field and the corresponding poffset.
[0084] Among them, as an optional implementation manner, the third-level index includes at least one group of fourth identifiers and fourth index information; among them, the time interval between two adjacent fourth identifiers is less than or equal to the second threshold.
[0085] Specifically, when the storage node determines that the time since the last write is greater than or equal to a second threshold based on the time information of the data, for example, it compares whether the difference between the opTimestamp in the current data and the maximum timestamp value corresponding to the existing first representation in the first file exceeds two hours (i.e., the second threshold, which can be pre-configured or set); if the difference exceeds two hours, a new opTimestamp field and the corresponding poffset field (i.e., the third-level index) are created, and the value of the timestamp quantity field is incremented by 1.
[0086] In some alternative embodiments, for the first node side, the first command at least includes: a first identifier indicating the replication backlog buffer, a second identifier indicating the cluster to which the second node belongs, and first time information; the reading of the first data from the first file of the storage node includes: the first node sending a third command to the storage node, the third command including the first identifier, the second identifier, and the first time information, for the storage node to obtain, based on the second command, the first data that belongs to the cluster corresponding to the second identifier, comes from the replication backlog buffer corresponding to the first identifier, and is generated after the first time information from the first file; receiving a fourth command sent by the storage node, the fourth command including the first data.
[0087] In this embodiment, the first node loads the corresponding data from the storage node according to the request or demand of the second node (or can be called the user), and sends the obtained data to the second node (or user) through a command.
[0088] Specifically, the second node can send a first command to the first node, and the first command is mainly used to inform which part of the data is requested. Among them, the first command may include four pieces of information: (cluster id, ring id, offset, opTimestamp). Among them, cluster id (i.e., the second identifier) is mainly used to verify whether it belongs to the same cluster, that is, to verify whether it has the qualification to request data. Ring id (i.e., the first identifier) is mainly used to search for the first-level index; offset (i.e., the first position information) is mainly used to search for the second-level index, and opTimestamp (i.e., the first time information) is mainly used to verify the timestamp. After receiving the first command, the first node can send a third command to the storage node to read or load the data corresponding to the third command. Among them, the third command may include the same four pieces of information (cluster id, ring id, offset, opTimestamp) in the first command, for the storage node to search in the first file according to these four pieces of information.
[0089] In some alternative embodiments, the method may further include: the first node determines whether the second node and the first node belong to the same cluster according to the second identifier of the cluster to which the second node belongs; in the case where the determination result is that the second node and the first node belong to the same cluster, reading the first data from the first file of the storage node; in the case where the determination result is that the second node and the first node do not belong to the same cluster, sending an error command to the second node.
[0090] The technical solution of this embodiment mainly aims at the synchronization of incremental data, that is, before the second node (or user) sends the first command, it has obtained the full amount of data at that time through full amount synchronization, but due to other reasons (such as communication interruption, etc.), it is necessary for the second node to synchronize incremental data. Therefore, for the second node, before sending the first command, it knows the position of the replicated backlog buffer of the data that has been synchronized in the first node. Therefore, this first position information may be the position of the data that has been synchronized and completed in the replicated backlog buffer of the first node.
[0091] In some alternative embodiments, for the storage node side, the storage node reads the first data requested by the command from the first file based on the command of the first node, including: the storage node receives a third command sent by the first node, and the third command includes: a first identifier representing the replicated backlog buffer, a second identifier representing the cluster to which the second node belongs, a first position information of the data in the replicated backlog buffer, and the first time information; determining whether the second identifier is consistent with the identifier of the cluster in the first file; in the case where the second identifier is consistent with the identifier of the cluster in the first file, finding the first-level index in the first file according to the first identifier, and determining the first index position corresponding to the first identifier; finding the second-level index according to the first index position, and determining the first data storage position corresponding to the first position information; finding the third-level index according to the first data storage position, and determining the corresponding second time information; if the second time information is compared and consistent with the first time information, reading the first data in the first file after the first data storage position.
[0092] In this embodiment, the storage node uses the first identifier, the second identifier, the first position information of the data in the replicated backlog buffer, and the first time information in the third command to find the index part in the first file, and determines the position of the first data in the first file according to the hierarchical index positioning, and then reads the first data.
[0093] Specifically, the storage node first determines whether the second identifier (cluster_id) of the cluster to which the second node belongs is consistent with the second identifier (cluster_id) of the cluster in the first file; if it is inconsistent, an error is reported and returned (the storage node sends an error command to the first node); if it is consistent, the position or offset (poffset) where the corresponding data first appears in the first file and the position or offset (index_offset) of the second-level index are obtained through the first identifier (ring id) representing the replication backlog buffer. For details, refer to Figure 7 as shown in (1) of
[0094] Secondly, the storage node searches for the second-level index through the first position information (i.e., offset) of the data in the replication backlog buffer. Exemplarily, when offset1 < offset < offset2, the data between (ring 1, offset1) and (ring 1, offset2) is read to accurately find the offset position poffset of the corresponding data file. For details, refer to Figure 7 as shown in (2) of
[0095] Finally, all the data after poffset in the first file is read as the first data and sent to the first node. For details, refer to Figure 7 as shown in (3) of
[0096] In some alternative embodiments, for the storage node side, the method further includes: the storage node clears the data in the first file according to a set retention time, deletes the data whose storage time exceeds the set retention time according to the time information corresponding to the data, and updates the index part of the header of the first file.
[0097] In this embodiment, at most the data with the set retention time is retained in the storage node, and the set retention time can be set or configured according to the actual situation (such as the size of the first file). Exemplarily, the set retention time is, for example, 2 days. Then the storage node clears the data in the first file according to the set retention time, deletes the data whose storage time exceeds the set retention time, and updates the content of the index part of the header of the first file.
[0098] Exemplarily, the storage node may set a timer, and the timing period of the timer is the set retention time. Then when the timing time of the timer arrives, the storage node may determine the data quantity according to the Timestamp_num field in the index part of the header of the first file; when the quantity of the data exceeds the maximum quantity corresponding to the set retention time, it indicates that there is data in the first file that exceeds the set retention time, and then the data that exceeds the set retention time may be deleted according to the timestamp information, and the content of the index part may be updated.
[0099] For the second node side, in some alternative embodiments, the method further includes: after the second node switches to the master node, the second node starts a replication backlog buffer and reads second data from the first file of the storage node.
[0100] In this embodiment, after the second node switches from the slave node to the master node, the replication backlog buffer is started, that is, the loader and the collector are started; the collector may be used to load the second data from the first file of the storage node according to the incremental synchronization scheme of the embodiment of the present invention.
[0101] The data synchronization method of the embodiment of the present invention will be described below with specific examples.
[0102] Figure 8 It is a schematic interaction process of the data synchronization method of the embodiment of the present invention Figure 1 ; in this example, the first node is used as the master node and the second node is used as the slave node for illustration. As Figure 8 shown, the method includes:
[0103] Step 401: The slave node sends a dsync?-1 command to the master node, and the dsync?-1 command is used to request the full amount of data in the replication backlog buffer.
[0104] This step is the initial synchronization process of the slave node. Among them, the dsync?-1 command is consistent with the psync command, and actively requests the master node to perform a full resynchronization.
[0105] Step 402: The master node generates an RDB (i.e., the Redis full amount data file) and sends the RDB to the slave node.
[0106] Step 403: The master node sends a command to the slave node to notify that the master-slave connection is temporarily interrupted.
[0107] Step 404: During the temporary interruption of the master-slave connection, the data in the replication backlog buffer in the master node is updated, and new data is written into the replication backlog buffer.
[0108] Step 405: When the master-slave connection is temporarily interrupted or the data in the master and slave nodes is inconsistent for other reasons, the slave node sends a first command to the master node to request partial data synchronization.
[0109] Here, the first command can be denoted as the dsync ringid offset opTimestamp command; the first command may include cluster id, ring id, offset, and opTimestamp; where cluster id represents the identifier of the cluster to which the slave node belongs; ring id represents the identifier of the replication backlog buffer requested by the slave node; offset represents the position information of the data requested by the slave node in the replication backlog buffer, identifying the data offset of the data in the replication backlog buffer; opTimestamp represents the time information corresponding to the data requested by the node.
[0110] Steps 406 to 407: When the replication backlog buffer in the master node contains the data requested by the slave node, it is processed according to the existing redis processing logic, that is, partial data is sent to the slave node using the Psync command to perform partial data synchronization.
[0111] Step 408: When the replication backlog buffer in the master node does not contain the data requested by the slave node, the master node triggers the loader to load the corresponding data in the first file (such as dump.rb) from the storage node.
[0112] Step 409: The master server sends a command to the slave node indicating to start the partial synchronization function of the storage node reading.
[0113] Here, the command can be denoted as the dfullsync ringid offset opTimestamp command, for example.
[0114] Step 410: The master node sends the data loaded from the first file to the slave node. Exemplarily, the master node can use the dsync command to send the data loaded from the first file to the slave node to perform partial data synchronization.
[0115] Figure 9 Schematic diagram of the interaction process of the data synchronization method according to the embodiment of the present invention Figure 2 ; as Figure 9 shown, the method includes:
[0116] When the master node fails, the master and standby nodes are switched, that is, the master node is switched to the slave node, and a slave node is switched to the master node. The new master node starts the collector and the loader, and the standby node stops using the collector and the loader. At this time, the value of the ring_id of the new master node is incremented, and the maintenance of the new replication backlog buffer data is started.
[0117] The new master node loads the corresponding data in the first file (dump.rb) of the storage node, and the specific process is the same as the process of partial data synchronization in the previous example. It should be noted that if the storage node determines that the cluster id carried in the first command for performing partial data synchronization is inconsistent with the cluster id in the first file, the new master node cannot load the data of the first file (dump.rb), and the process is aborted.
[0118] In this scenario, when a slave node requests partial data synchronization, the loader of the new master node can load the first file (dump.rb) in the storage node to maintain the request and pull its data to the new state. After the master-slave switch, the slave node performs incremental synchronization by sending the first command (such as the dsync command), instead of full synchronization, which can reduce the cost of data synchronization between the master-slave clusters.
[0119] Based on the above embodiments, an embodiment of the present invention further provides a data synchronization device, which is applied to the first node. Figure 10 Schematic diagram of the composition structure of the data synchronization device according to the embodiment of the present invention Figure 1 ; as Figure 10 shown, the device includes a collection unit 11, a loading unit 12, and a first communication unit 13; wherein,
[0120] The collection unit 11 is configured to collect data in the replication backlog buffer and write the data and related information of the data into the first file in the storage node; wherein, the related information of the data includes: a first identifier representing the replication backlog buffer, a second identifier representing the cluster to which the first node belongs, position information of the data in the replication backlog buffer, and time information of the data.
[0121] The loading unit 12 is configured to, when receiving a first command for requesting partial data synchronization sent by a second node acting as a slave node, read the first data corresponding to the first command from the first file of the storage node when the first data requested by the first command is not included in the replication backlog buffer.
[0122] The first communication unit 13 is configured to send the first data to the second node through a second command, where the first node is the master node and the second node is the slave node.
[0123] In some alternative embodiments of the present invention, the first command at least includes: a first identifier indicating the replication backlog buffer, a second identifier indicating the cluster to which the second node belongs, first position information of the data in the replication backlog buffer, and first time information; the loading unit 12 is configured to send a third command to the storage node, the third command including the first identifier, the second identifier, the first position information of the data in the replication backlog buffer, and the first time information, for the storage node to obtain, based on the third command, first data that belongs to the cluster corresponding to the second identifier, comes from the replication backlog buffer corresponding to the first identifier, and is generated after the first time information, from the first file; and is further configured to receive a fourth command sent by the storage node, the fourth command including the first data.
[0124] In the embodiments of the present invention, the acquisition unit 11 and the loading unit 12 in the device can be implemented by a central processing unit (CPU, Central Processing Unit), a digital signal processor (DSP, Digital Signal Processor), a microcontroller unit (MCU, Microcontroller Unit), or a field-programmable gate array (FPGA, Field-Programmable Gate Array) in combination with a communication interface in actual applications; the first communication unit 13 in the device can be implemented by a communication module (including: a basic communication suite, an operating system, a communication module, a standardized interface, and a protocol, etc.) and a transceiver antenna in actual applications.
[0125] The embodiments of the present invention further provide a data synchronization device, and the device is applied to a storage node. Figure 11 Schematic diagram of the composition structure of the data synchronization device according to the embodiments of the present invention Figure 2 ; as Figure 11 shown, the device includes: a second communication unit 21, a processing unit 22, and a storage unit 23; wherein,
[0126] The second communication unit 21 is configured to receive data sent by a first node and related information of the data, the related information of the data including: a first identifier of the replication backlog buffer of the first node, a second identifier of the cluster to which the first node belongs, position information of the data in the replication backlog buffer, and time information of the data;
[0127] The processing unit 22 is configured to write the data and the related information of the data into a first file; and is further configured to read, based on a command of the first node, first data requested by the command from the first file;
[0128] The second communication unit 21 is further configured to send the first data to the first node, so that the first node sends the first data to the second node; the first node is the master node, and the second node is the slave node;
[0129] The storage unit 23 is configured to store the first file.
[0130] In some alternative embodiments of the present invention, the processing unit 22 is configured to write the data and the related information of the data to the tail of the data part of the first file, and modify the index part of the first file according to the related information of the data according to the following rules:
[0131] If the first identifier is the identifier of a newly added replication backlog buffer in the first file, add a first-level index corresponding to the first identifier in the index information; the first-level index includes first index information and second index information, the first index information is used to indicate the position where the data corresponding to the first identifier first appears in the first file, and the second index information is used to indicate the position of the second-level index corresponding to the first identifier;
[0132] Add a second-level index corresponding to the first identifier in the index part; the second-level index includes third index information corresponding to a third identifier, the third identifier is used to represent the position information of the first identifier and the data in the replication backlog buffer, and the third index information is used to indicate the data storage position of the data corresponding to the third identifier in the first file;
[0133] Add a corresponding third-level index in the index part; the third-level index includes fourth index information corresponding to a fourth identifier, the fourth identifier is used to represent the time information of the data, and the fourth index information is used to indicate the data storage position of the data corresponding to the fourth identifier in the first file.
[0134] In some alternative embodiments of the present invention, the second-level index includes at least one group of third identifiers and third index information; wherein, the data size represented by the position information of the data corresponding to each group of third identifiers in the replication backlog buffer is less than or equal to a first threshold; and / or,
[0135] The third-level index includes at least one group of fourth identifiers and fourth index information; wherein, the time interval between two adjacent fourth identifiers is less than or equal to a second threshold.
[0136] In some alternative embodiments of the present invention, the processing unit 22 is configured to receive, via the second communication unit 21, a third command sent by the first node, where the third command includes: a first identifier indicating a replication backlog buffer, a second identifier indicating the cluster to which the second node belongs, first position information of data in the replication backlog buffer, and the first time information; determine whether the second identifier is consistent with the identifier of the cluster in the first file; in the case where the second identifier is consistent with the identifier of the cluster in the first file, find the first-level index in the first file according to the first identifier, and determine the first index position corresponding to the first identifier; find the second-level index according to the first index position, and determine the first data storage position corresponding to the first position information; find the third-level index according to the first data storage position, and determine the corresponding second time information; if the second time information is consistent with the first time information after comparison, read the first data in the first file after the first data storage position.
[0137] In some alternative embodiments of the present invention, the processing unit 22 is further configured to clean the data in the first file according to a set retention time, delete the data whose storage time exceeds the set retention time according to the time information corresponding to the data, and update the index part of the header of the first file.
[0138] In an embodiment of the present invention, the processing unit 22 in the device can be implemented by a CPU, a DSP, an MCU, or an FPGA in combination with a communication interface in actual application; the storage unit 23 in the device can be implemented by a memory in actual application; the second communication unit 21 in the device can be implemented by a communication module (including: a basic communication suite, an operating system, a communication module, a standardized interface, and a protocol, etc.) and a transceiver antenna in actual application.
[0139] An embodiment of the present invention further provides a data synchronization device, and the device is applied to a second node. Figure 12 Schematic diagram of the composition structure of the data synchronization device according to the embodiment of the present invention Figure 3 ; as Figure 12As shown in the figure, the device includes: a third communication unit 31, configured to send a first command to a first node, where the first command is used to request partial data synchronization; and further configured to receive a second command sent by the first node, where the second command includes first data, and the first data is read by the first node from a first file of a storage node; the first file prestores data of a replication backlog buffer written by the first node, and the data stored in the first file at least includes: a first identifier representing the replication backlog buffer, a second identifier representing the cluster to which the first node belongs, location information of the data in the replication backlog buffer, and time information of the data; the first node is a master node, and the second node is a slave node.
[0140] In some alternative embodiments of the present invention, the device further includes a processing unit 32 and a loading unit 33; wherein, the processing unit 32 is configured to start a replication backlog buffer after the second node switches to be the master node;
[0141] The loading unit 33 is configured to read second data from the first file of the storage node.
[0142] In an embodiment of the present invention, the processing unit 32 and the loading unit 33 in the device can both be implemented by a CPU, a DSP, an MCU or an FPGA in combination with a communication interface in practical applications; the third communication unit 31 in the device can be implemented by a communication module (including: a basic communication suite, an operating system, a communication module, a standardized interface and protocol, etc.) and a transceiver antenna in practical applications.
[0143] It should be noted that: when the data synchronization device provided in the above embodiment performs data synchronization, only the above division of each program module is used for illustration. In practical applications, the above processing can be allocated to different program modules according to needs, that is, the internal structure of the device is divided into different program modules to complete all or part of the above-described processing. In addition, the data synchronization device provided in the above embodiment and the data synchronization method embodiment belong to the same concept, and the specific implementation process is detailed in the method embodiment, which will not be elaborated here.
[0144] An embodiment of the present invention further provides a communication node, and the communication node is a first node, a second node or a storage node. Figure 13 It is a schematic diagram of the hardware composition structure of the communication node according to an embodiment of the present invention, as Figure 13 shown, the communication node includes a memory 42, a processor 41, and a computer program stored on the memory 42 and executable on the processor 41. When the processor 41 executes the program, it implements the steps of the data synchronization method applied to the first node, the second node or the storage node according to an embodiment of the present invention.
[0145] Optionally, at least one communication interface 43 may also be included in the communication node. Each component in the communication node is coupled together through a bus system 44. It can be understood that the bus system 44 is used to implement the connection and communication between these components. In addition to including a data bus, the bus system 44 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clear illustration, in Figure 13 all kinds of buses are labeled as the bus system 44.
[0146] It can be understood that the memory 42 can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a ferromagnetic random access memory (FRAM), a flash memory, a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD-ROM); the magnetic surface memory can be a disk memory or a tape memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static random access memory (SRAM), synchronous static random access memory (SSRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), sync link dynamic random access memory (SLDRAM), direct rambus random access memory (DRRAM).The memory 42 described in the embodiments of the present invention is intended to include, but not limited to, these and any other suitable types of memories.
[0147] The method disclosed in the above embodiments of the present invention can be applied to the processor 41 or implemented by the processor 41. The processor 41 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit in the hardware of the processor 41 or the instructions in the form of software. The above-mentioned processor 41 may be a general-purpose processor, DSP, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor 41 can implement or execute each method, step, and logic block diagram disclosed in the embodiments of the present invention. The general-purpose processor may be a microprocessor or any conventional processor, etc. Combining the steps of the method disclosed in the embodiments of the present invention can be directly embodied as being executed and completed by the hardware decoding processor, or executed and completed by the combination of the hardware and software modules in the decoding processor. The software module may be located in the storage medium, and this storage medium is located in the memory 42. The processor 41 reads the information in the memory 42 and combines its hardware to complete the steps of the foregoing method.
[0148] In an exemplary embodiment, the communication node may be implemented by one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), FPGAs, general-purpose processors, controllers, MCUs, microprocessors, or other electronic components for performing the foregoing method.
[0149] In an exemplary embodiment, the embodiments of the present invention also provide a computer-readable storage medium, such as the memory 42 including a computer program, and the above computer program can be executed by the processor 41 of the communication node to complete the steps of the foregoing method. The computer-readable storage medium may be a FRAM, ROM, PROM, EPROM, EEPROM, Flash Memory, magnetic surface memory, optical disc, or CD-ROM, etc.; or it may be various devices including one or any combination of the above memories.
[0150] The computer-readable storage medium provided by the embodiments of the present invention stores a computer program thereon, and when the program is executed by the processor, it implements the steps of the data synchronization method in which the embodiments of the present invention are applied to the first node, the second node, or the storage node.
[0151] The embodiments of the present application also provide a computer program product, including a computer program, which can be executed by a communication node (such as the processor 41 of the communication node) to complete the steps of any of the foregoing data synchronization methods.
[0152] The features disclosed in several method or device embodiments provided by the present application can be combined arbitrarily without conflict to obtain new method embodiments or device embodiments.
[0153] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are only illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined, or can be integrated into another system, or some features can be ignored, or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed with each other can be through some interfaces. The indirect coupling or communication connection of the devices or units can be electrical, mechanical, or other forms.
[0154] The units described as separate components above may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units; some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0155] In addition, each functional unit in the embodiments of the present invention can be all integrated in a processing unit, or each unit can be separately used as a unit, or two or more units can be integrated in a unit; the above integrated unit can be implemented in the form of hardware, or in the form of a combination of hardware and software functional units.
[0156] Those of ordinary skill in the art can understand that all or part of the steps of implementing the foregoing method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps including the foregoing method embodiments; and the foregoing storage medium includes: various media such as removable storage devices, ROM, RAM, magnetic disks, or optical discs that can store program codes.
[0157] Alternatively, if the above integrated units of the present invention are implemented in the form of software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present invention, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as removable storage devices, ROM, RAM, magnetic disks, or optical discs.
[0158] As described above, the above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A data synchronization method, characterized in that: The method is applied to a first node, and the method comprises: The first node collects data in the replication backlog buffer, and writes the data and related information of the data into a first file in the storage node; wherein the related information of the data includes: a first identifier representing the replication backlog buffer, a second identifier representing the cluster to which the first node belongs, location information of the data in the replication backlog buffer, and time information of the data; When a first command for requesting partial data synchronization is received from a second node, if the replication backlog buffer does not contain the first data requested by the first command, the first data corresponding to the first command is read from the first file of the storage node, and the first data is sent to the second node through a second command. The first node is a master node, and the second node is a slave node.
2. The method according to claim 1, characterized in that The first command includes at least: a first identifier indicating a replication backlog buffer, a second identifier indicating a cluster to which the second node belongs, first location information and first time information of data in the replication backlog buffer; and reading the first data from the first file of the storage node includes: The first node sends a third command to the storage node, where the third command includes the first identifier, the second identifier, first position information of data in the replication backlog buffer, and the first time information, so that the storage node obtains, based on the third command, first data from the first file that belongs to the cluster corresponding to the second identifier, comes from the replication backlog buffer corresponding to the first identifier, and is generated after the first time information; A fourth command sent by the storage node is received, where the fourth command includes the first data.
3. A data synchronization method, characterized in that: The method is applied to a storage node, and the method comprises: The storage node receives data sent by the first node and related information of the data, and writes the data and related information of the data into a first file; the related information of the data includes: a first identifier of a replication backlog buffer of the first node, a second identifier of a cluster to which the first node belongs, location information of the data in the replication backlog buffer, and time information of the data; The storage node reads the first data requested by the command from the first file based on the command of the first node, and sends the first data to the first node, so that the first node sends the first data to the second node. The first node is a master node and the second node is a slave node.
4. The method according to claim 3, characterized in that Writing the data and related information of the data into the first file includes: The data and the related information of the data are written into the tail of the data part of the first file, and the index part of the first file is modified according to the related information of the data in accordance with the following rules: If the first identifier is an identifier of a newly added copy backlog buffer in the first file, a first-level index corresponding to the first identifier is added to the index information; the first-level index includes first index information and second index information, the first index information is used to indicate the position where the data corresponding to the first identifier first appears in the first file, and the second index information is used to indicate the position of the second-level index corresponding to the first identifier; Adding a second-level index corresponding to the first identifier to the index part; the second-level index includes third index information corresponding to the third identifier, the third identifier is used to indicate the first identifier and the location information of the data in the copy backlog buffer, and the third index information is used to indicate the data storage location of the data corresponding to the third identifier in the first file; A corresponding third-level index is added to the index part; the third-level index includes fourth index information corresponding to the fourth identifier, the fourth identifier is used to represent the time information of the data, and the fourth index information is used to indicate the data storage location corresponding to the fourth identifier in the first file.
5. The method according to claim 4, characterized in that The second-level index includes at least one set of third identifiers and third index information; wherein the data size represented by the position information of the data corresponding to each set of third identifiers in the replication backlog buffer is less than or equal to the first threshold; and / or, The third-level index includes at least one set of fourth identifiers and fourth index information; wherein the time interval represented by two adjacent fourth identifiers is less than or equal to the second threshold.
6. The method according to any one of claims 3 to 5, characterized in that: The storage node reads first data requested by the command from the first file based on the command of the first node, including: The storage node receives a third command sent by the first node, where the third command includes: a first identifier indicating a replication backlog buffer, a second identifier indicating a cluster to which the second node belongs, first location information of data in the replication backlog buffer, and the first time information; Determine whether the second identifier is consistent with the identifier of the cluster in the first file; When the second identifier is consistent with the identifier of the cluster in the first file, searching the first level index in the first file according to the first identifier to determine the first index position corresponding to the first identifier; Searching the second-level index according to the first index position to determine a first data storage position corresponding to the first position information; Searching the third-level index according to the first data storage location to determine the corresponding second time information; If the second time information is consistent with the first time information, the first data in the first file that is located after the first data storage position is read.
7. The method according to claim 3, characterized in that The method further comprises: The storage node cleans up the data in the first file according to the set retention time, deletes the data whose storage time exceeds the set retention time according to the time information corresponding to the data, and updates the index part of the header of the first file.
8. A data synchronization method, characterized in that: The method is applied as a second node, and the method includes: The second node sends a first command to the first node, where the first command is used to request partial data synchronization; the first node is a master node, and the second node is a slave node; Receive a second command sent by the first node, the second command includes first data, the first data is read by the first node from a first file of a storage node; the first file pre-stores data of a replication backlog buffer written by the first node, the data stored in the first file includes at least: a first identifier representing the replication backlog buffer, a second identifier representing the cluster to which the first node belongs, location information of the data in the replication backlog buffer, and time information of the data.
9. The method according to claim 8, characterized in that The method further comprises: After the second node switches to the primary node, the second node starts to copy the backlog buffer area and reads the second data from the first file of the storage node.
10. A data synchronization device, characterized in that: The device is applied to a first node, and comprises: a collection unit, a loading unit and a first communication unit; wherein, The collection unit is used to collect data in the replication backlog buffer and write the data and related information of the data into a first file in the storage node; wherein the related information of the data includes: a first identifier representing the replication backlog buffer, a second identifier representing the cluster to which the first node belongs, location information of the data in the replication backlog buffer, and time information of the data; The loading unit is configured to, upon receiving a first command for requesting partial data synchronization sent by a second node as a slave node, read first data corresponding to the first command from a first file of the storage node when the replication backlog buffer does not contain the first data requested by the first command; The first communication unit is used to send the first data to the second node through a second command, the first node is a master node, and the second node is a slave node.
11. A data synchronization device, characterized in that: The device is applied to a storage node, and comprises: a second communication unit, a processing unit and a storage unit; wherein, The second communication unit is used to receive data sent by the first node and related information of the data, where the related information of the data includes: a first identifier of the replication backlog buffer of the first node, a second identifier of the cluster to which the first node belongs, location information of the data in the replication backlog buffer, and time information of the data; The processing unit is configured to write the data and related information of the data into a first file; and is further configured to read first data requested by a command from the first file based on a command of the first node; The second communication unit is further configured to send the first data to the first node, so that the first node sends the first data to the second node; the first node is a master node, and the second node is a slave node; The storage unit is used to store the first file.
12. A data synchronization device, characterized in that: The device is applied to the second node, and the device includes: a third communication unit, used to send a first command to the first node, the first command is used to request partial data synchronization; and also used to receive a second command sent by the first node, the second command includes first data, and the first data is read by the first node from a first file of a storage node; the first file pre-stores data of a replication backlog buffer written by the first node, and the data stored in the first file includes at least: a first identifier representing the replication backlog buffer, a second identifier representing the cluster to which the first node belongs, location information of the data in the replication backlog buffer, and time information of the data; the first node is a master node, and the second node is a slave node.
13. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method according to any one of claims 1 to 9 are implemented.
14. A communication node, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps of the method according to any one of claims 1 to 9 are implemented.
15. A computer program product, characterized in that The method comprises computer program instructions, which enable a computer to execute the steps of the method according to any one of claims 1 to 9.