Distributed file synchronization method, system and device and storage medium

By determining the node identity information based on the node startup time in a distributed system and adopting a decentralized file synchronization method, the problem of cumbersome file synchronization and centralized single point failure in the existing technology is solved, and efficient and stable file synchronization is achieved.

CN119988340AActive Publication Date: 2025-05-13BIGO TECH PTE LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510019886.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-07
Publication Date
2025-05-13
Estimated Expiration
2045-01-07

AI Technical Summary

Technical Problem

The file synchronization process between nodes in existing distributed systems is cumbersome and time-consuming, resulting in large consumption of network resources and the centralized synchronization method has a single point of failure.

Method used

By obtaining the node startup time of each regional node in the distributed system, the node identity information of each regional node is determined, including the regional follower node, the regional leader node and the main leader node, the file synchronization is carried out in a decentralized manner, and only the differential file information is transmitted.

Benefits of technology

It realizes efficient file synchronization between nodes under changing network environment, reduces network resource consumption, simplifies synchronization process, and improves stability and reliability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119988340A_ABST
    Figure CN119988340A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a distributed file synchronization method, system and device and a storage medium. According to the technical scheme provided by the embodiment of the invention, the node starting time of each regional node in the distributed system is acquired, the node identity information of each regional node is determined based on the node starting time, and the file synchronization of the distributed system is carried out based on the node identity information, so that decentralized file synchronization is realized; under the condition that the network environment changes, efficient file synchronization can be carried out through change of the node identity information, the tedious process of file synchronization between the nodes of the distributed system is simplified, and the stability and reliability of file synchronization are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of computer technology, and in particular to a distributed file synchronization method, system, device and storage medium. Background Art

[0002] Currently, in distributed system operation scenarios, file synchronization between nodes is often required. File synchronization technology is a technology that ensures the consistency of file content between multiple nodes by comparing and transmitting file data. In related file synchronization solutions, one node usually actively notifies other nodes to pull files from the data source to achieve file synchronization between nodes.

[0003] However, the related file synchronization scheme requires each node to frequently pull files from the data source, resulting in large network resource consumption, and then network congestion and delays. The whole process is cumbersome and time-consuming, affecting the efficiency of file synchronization. In addition, since the nodes use a centralized method to synchronize files, when the central node responsible for notifying other nodes to synchronize files has a single point of failure, it will cause file synchronization to fail, affecting the stability and reliability of file synchronization. Summary of the invention

[0004] The embodiments of the present application provide a distributed file synchronization method, system, device and storage medium, which can adapt to changes in the network environment to efficiently synchronize files between nodes and solve the technical problem that the file synchronization process between nodes in a distributed system is cumbersome and time-consuming.

[0005] In a first aspect, an embodiment of the present application provides a distributed file synchronization method, including:

[0006] Obtaining the node startup time of each regional node in the distributed system, and determining the node identity information of each regional node based on the node startup time, wherein the node identity information includes regional follower nodes and regional leader nodes corresponding to different target regions, and a master leader node corresponding to each regional leader node, wherein the regional leader node is used to communicate with other regional follower nodes in the target region to which it belongs, and the master leader node is used to communicate with each regional leader node;

[0007] In the case where the node identity information of the local machine is the primary leader node, receiving the update file of the data source and storing the update file in the local first file information, determining the first difference file information between the first file information and the second file information of each regional leader node, and synchronizing the first difference file information to the corresponding regional leader node;

[0008] In the case where the node identity information of the local machine is a regional leader node, determining second difference file information between the local second file information and the third file information of each regional follower node in the target region, and synchronizing the second difference file information to the regional follower node;

[0009] When the node identity information of the local computer is a regional follower node, the local third file information is reported to the regional leader node of the target area, and the second difference file information synchronized by the regional leader node is received.

[0010] In a second aspect, an embodiment of the present application provides a distributed file synchronization system, including:

[0011] an identity determination module, configured to obtain a node startup time of each regional node in the distributed system, and determine the node identity information of each regional node based on the node startup time, wherein the node identity information includes regional follower nodes and regional leader nodes corresponding to different target regions, and a master leader node corresponding to each regional leader node, wherein the regional leader node is used to communicate with other regional follower nodes in the target region to which it belongs, and the master leader node is used to communicate with each regional leader node;

[0012] A first synchronization module is configured to receive an update file of a data source and store the update file in a local first file information when the node identity information of the local machine is a primary leader node, determine first difference file information between the first file information and the second file information of each regional leader node, and synchronize the first difference file information to the corresponding regional leader node;

[0013] A second synchronization module is configured to determine second difference file information between the local second file information and the third file information of each regional follower node in the target region when the node identity information of the local machine is a regional leader node, and synchronize the second difference file information to the regional follower node;

[0014] The third synchronization module is configured to report the local third file information to the regional leader node of the target area and receive the second difference file information synchronized by the regional leader node when the node identity information of the local machine is a regional follower node.

[0015] In a third aspect, an embodiment of the present application provides a distributed file synchronization device, including:

[0016] memory and one or more processors;

[0017] The memory is configured to store one or more programs;

[0018] When the one or more programs are executed by the one or more processors, the one or more processors implement the distributed file synchronization method as described in the first aspect.

[0019] In a fourth aspect, an embodiment of the present application provides a non-volatile computer-readable storage medium, which stores computer-executable instructions, and the computer-executable instructions are configured to execute the distributed file synchronization method as described in the first aspect when executed by a computer processor.

[0020] In a fifth aspect, an embodiment of the present application provides a computer program product, which includes instructions. When the instructions are executed on a computer or a processor, the computer or the processor executes the distributed file synchronization method as described in the first aspect.

[0021] In an embodiment of the present application, by acquiring the node startup time of each regional node in the distributed system, the node identity information of each regional node is determined based on the node startup time. The node identity information includes regional follower nodes and regional leader nodes corresponding to different target regions, and a master leader node corresponding to each regional leader node. The regional leader node is used to communicate with other regional follower nodes in the target region to which it belongs, and the master leader node is used to communicate with each regional leader node. In the case where the node identity information of the local machine is the master leader node, an update file of a data source is received and the update file is stored in the local first file information, the first difference file information between the first file information and the second file information of each regional leader node is determined, and the first difference file information is synchronized to the corresponding regional leader node. In the case where the node identity information of the local machine is the regional leader node, the second difference file information between the local second file information and the third file information of each regional follower node in the target region to which it belongs is determined, and the second difference file information is synchronized to the regional follower node. In the case where the node identity information of the local machine is the regional follower node, the local third file information is reported to the regional leader node of the target region to which it belongs, and the second difference file information synchronized by the regional leader node is received. By adopting the above-mentioned technical means, the node identity information of each regional node is determined by the node startup time, and the file synchronization of the distributed system is performed based on the node identity information, thereby realizing decentralized file synchronization. When the network environment changes, the file can also be efficiently synchronized by changing the node identity information, simplifying the cumbersome process of file synchronization between nodes in the distributed system and improving the stability and reliability of file synchronization. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 is a flow chart of a distributed file synchronization method provided by an embodiment of the present application;

[0023] Figure 2It is an interaction flow chart of each module of the regional node in the embodiment of the present application;

[0024] Figure 3 It is a flowchart of file synchronization of nodes in each region of the distributed system in an embodiment of the present application;

[0025] Figure 4 It is a structural diagram of a distributed file synchronization system provided by an embodiment of the present application;

[0026] Figure 5 It is a structural diagram of a distributed file synchronization device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0027] In order to make the purpose, technical scheme and advantages of the present application clearer, the specific embodiments of the present application are further described in detail below in conjunction with the accompanying drawings. It is understood that the specific embodiments described herein are only used to explain the present application, rather than to limit the present application. It should also be noted that, for the convenience of description, only the part related to the present application but not all the contents are shown in the accompanying drawings. Before discussing the exemplary embodiments in more detail, it should be mentioned that some exemplary embodiments are described as processes or methods depicted as flow charts. Although the flow chart describes each operation (or step) as a sequential process, many of the operations therein can be implemented in parallel, concurrently or simultaneously. In addition, the order of each operation can be rearranged. The process can be terminated when its operation is completed, but it can also have additional steps not included in the accompanying drawings. The process can correspond to a method, a function, a procedure, a subroutine, a subprogram, etc.

[0028] The distributed file synchronization method provided in this application aims to determine the node identity information of each regional node through the node startup time, and perform file synchronization of the distributed system based on the node identity information, thereby realizing decentralized file synchronization.

[0029] In related file synchronization schemes, one node usually actively notifies other nodes to pull files from the data source. The disadvantages of this method are very obvious. Frequent file pulling operations will cause a large consumption of network resources on the data source, especially when there are a large number of nodes or frequent data updates, which may cause network congestion and delays. In addition, there is also a method of using incremental transmission to reduce the consumption of network resources, but the transmission node needs to actively obtain information from other nodes that need to be synchronized. The whole process lacks flexibility and it is difficult to handle network topology changes in real time.

[0030] In addition, related file synchronization solutions are usually centralized, often require complex configuration and maintenance, and have a series of inherent problems, such as single point failure, limited system expansion, and inability to effectively handle network partitions. These defects are particularly prominent in large-scale or changeable network environments. Although some distributed file synchronization solutions have emerged, which can overcome the shortcomings of centralized systems to a certain extent, these solutions often need to rely on a distributed component, which requires the introduction of additional management complexity and operation and maintenance costs.

[0031] Based on this, a distributed file synchronization method is provided in an embodiment of the present application to solve the technical problem that the file synchronization process between nodes in a distributed system is cumbersome and time-consuming.

[0032] Example:

[0033] Figure 1 A flow chart of a distributed file synchronization method provided in an embodiment of the present application is given. The distributed file synchronization method provided in this embodiment can be executed by a distributed file synchronization device, which can be implemented by software and / or hardware. The distributed file synchronization device can be composed of two or more physical entities, or can be composed of one physical entity. Generally speaking, the distributed file synchronization device can be a processing device such as a node server of a distributed system.

[0034] The following description is made by taking the distributed file synchronization device as an example to perform the distributed file synchronization method. Figure 1 , the distributed file synchronization method specifically includes:

[0035] S110, obtaining the node startup time of each regional node in the distributed system, and determining the node identity information of each regional node based on the node startup time, wherein the node identity information includes regional follower nodes and regional leader nodes corresponding to different target regions, and a master leader node corresponding to each regional leader node, wherein the regional leader node is used to communicate with other regional follower nodes in the target region to which it belongs, and the master leader node is used to communicate with each regional leader node;

[0036] This application achieves seamless discovery and efficient file synchronization of devices in the local area network through a decentralized design when synchronizing files in a distributed system. At the same time, it uses a decentralized approach to adapt to changes in the network environment in real time, supports seamless collaboration between multiple nodes, reduces the consumption of network resources, and thus improves overall performance and stability.

[0037] Among them, when performing file synchronization between nodes in each region of the distributed system, the node startup time of each regional node in the distributed system is obtained to determine the node identity information of each regional node based on the node startup time, and the node identity information is used to perform file synchronization operations. Each regional node in the distributed system records its own startup time when it starts. By obtaining these startup times, a corresponding role can be assigned to each node. The node with the earliest startup time is selected as the main leader node, and the node with the earliest startup time in each target area is used as the regional leader node of the area, and the remaining nodes correspond to each target area as regional follower nodes. According to actual needs, the distributed system can be divided into multiple areas, and different areas are configured with a corresponding number of regional nodes. The divided areas are defined as target areas.

[0038] For each regional node, the following modules are included:

[0039] Node discovery module: used to discover other nodes in the LAN in real time based on mDNS;

[0040] Node management module: used to maintain the file synchronization status and health status of other nodes, responsible for electing leader nodes, and providing users with relevant query interfaces;

[0041] File download module: used to download files from data sources;

[0042] File management module: used to maintain the file status of this node;

[0043] Transmission optimization module: used to support incremental synchronization and breakpoint resume, ensuring that bandwidth is saved as much as possible and efficiency is improved during file transfer;

[0044] Node state machine: responsible for the relevant logic when this node switches between states.

[0045] like Figure 2 As shown in the figure, after the regional node is started, its node discovery module will send an mDNS notification message to the multicast address of the specified IP and port to register a specific host name and the IP address of this node. mDNS (Multicast DNS) is a device discovery protocol in a local area network. It realizes automatic device discovery by broadcasting DNS requests and responses in the local area network. Then the node discovery module will periodically send mDNS query messages to query other devices with a specific host name to obtain other node information in the current local area network. When the mDNS query response is received, the queried node information is handed over to the node management module for processing.

[0046] The node management module will periodically request other regional nodes to obtain information such as system startup time, regional information, and health status of other regional nodes. Based on the node startup time of each regional node obtained, the node identity information of each regional node can be configured according to the node startup time, and the corresponding leader node can be elected to synchronize files in the distributed system.

[0047] By using the mDNS protocol, the distributed system does not need to actively obtain node information or introduce additional components such as registration centers, which can reduce the complexity of the system, simplify the connection process between nodes, and reduce dependence on centralized management, thereby improving the flexibility and reliability of file synchronization.

[0048] Specifically, the node identity information of the present application includes regional follower nodes and regional leader nodes corresponding to different target areas, and master leader nodes corresponding to each regional leader node. When determining the node identity information of each regional node based on the node startup time, it includes:

[0049] Construct a small top heap data structure based on the node startup time of each regional node;

[0050] The small top heap data structure is queried based on the pre-built comparator, and the node identity information of each regional node is determined according to the position of each regional node in the small top heap data structure.

[0051] The election of the leader node is realized by maintaining a small top heap sorted by the node startup time. In it, a comparator is defined to compare the nodes according to the node startup time, and the node with the smaller startup time is prioritized. If the startup time is the same, the node IP is compared, and the smaller IP is prioritized. Then, an election class is defined based on the priority queue (in C++, it can inherit the standard library std::priority_queue), and the above comparator is used as the sorting rule.

[0052] When electing a leader node, the root node is taken out from the priority queue of the small top heap as the regional leader node of the corresponding target area, and the head of the small top heap is the main leader node of each regional leader node. The remaining nodes are regional follower nodes under the corresponding regional leader node (root node). Therefore, according to the election results, node identity information is assigned to each node, including regional follower node, regional leader node and main leader node.

[0053] It is understandable that the regional leader node is the node with the earliest system startup time among the nodes in this region, and the main leader node is the node with the earliest system startup time among all nodes. When a node is offline, the leader node election will be re-performed. When other nodes in the target area become leader nodes, other regional nodes in this area will become regional follower nodes.

[0054] After building the small top heap data structure based on the node startup time of each regional node, it also includes:

[0055] When a new regional node is detected, the new regional node is inserted at the end of the queue of the small top heap data structure based on the node startup time, and the small top heap data structure is rebuilt;

[0056] When the removal of the specified area node is detected, if the specified area node is at the head of the mini-heap data structure, the specified area node is removed based on the node pop-up function; if the specified area node is not at the head of the mini-heap data structure, the specified area node is directly removed and the mini-heap data structure is rebuilt.

[0057] The node discovery module queries the node information of other regional nodes in the distributed system in real time. When a new regional node is detected in the distributed system, the new node is inserted at the end of the priority queue and the small top heap is rebuilt. In addition, when a regional node in the distributed implementation system is detected to be removed, the node corresponding to the IP of the regional node in the priority queue is searched. If the regional node is found at the head of the queue, the head node is popped out. Otherwise, it is directly removed from the queue and the small top heap is rebuilt.

[0058] The regional node will maintain the rotation of the node between various states through the node state machine. The node has three states: initial state, follower state, and leader state. When the node has not obtained the information of other nodes, it will be in the initial state. When the node obtains the information of other nodes, it can conduct a local leader node election to determine whether its own state is a leader or a follower, thereby determining the node identity information of the local machine.

[0059] In addition, the node management module is also used to view the IP address and current connection status of the regional node, check the progress and synchronization history of file synchronization, detect the health status of other nodes (such as whether they are online, the last synchronization time), suspend or resume the synchronization task of a specific node, and avoid file conflicts caused by network instability. When the health status of a node is abnormal, other nodes can promptly discover the abnormal node through mDNS broadcast and suspend synchronization operations to the abnormal node. When a new node is discovered, the node management module will also immediately obtain relevant information about the new node and add it to the management list. The distributed system can actively monitor the health status, synchronization progress, and recent file synchronization records of other nodes in the network without additional configuration. This makes file synchronization more convenient, efficient, and highly stable.

[0060] S120. When the node identity information of the local machine is the primary leader node, receive the update file of the data source and store the update file in the local first file information, determine the first difference file information between the first file information and the second file information of each regional leader node, and synchronize the first difference file information to the corresponding regional leader node.

[0061] Further, based on the configuration of the node identity information, if the node identity information of the local local node of the current regional node is the primary leader node, when the data source has updated files, the primary leader node receives these files and updates its first file information library. Subsequently, the primary leader node calculates the difference between the first file information and the second file information of each regional leader node, and defines this difference as the first difference file information. The primary leader node will send these difference file information to the corresponding regional leader node to synchronize files between the primary leader node and the regional leader node.

[0062] S130. When the node identity information of the local device is a regional leader node, determine the second difference file information between the local second file information and the third file information of each regional follower node in the target region, and synchronize the second difference file information to the regional follower node.

[0063] Based on the configuration of the node identity information, if the node identity information of the current regional node is the regional leader node, the regional leader node will receive the first difference file information of the main leader node and update its second file information library. Then, the regional leader node calculates the difference between the second file information and the third file information of each regional follower node in the region to which it belongs, that is, the second difference file information. The regional leader node sends these difference file information to the respective target regional follower nodes to synchronize files between the regional leader node and the regional follower nodes.

[0064] S140. When the node identity information of the local device is a regional follower node, report the local third file information to the regional leader node of the target region, and receive the second difference file information synchronized by the regional leader node.

[0065] Based on the configuration of the node identity information, if the node identity information of the local node of the current regional node is a regional follower node, the regional follower node will report its third file information to the regional leader node of the region to which it belongs regularly or as needed. When receiving the second difference file information of the regional leader node, the regional follower node updates its third file information library to synchronize files between the regional leader node and the regional follower node.

[0066] For example, refer to Figure 3In a distributed file system, after determining the node identity information based on the node startup time, node 1 is configured as the primary leader node, and also as the regional leader node of the region to which it belongs, for communicating with the two regional follower nodes, node 2 and node 3. Nodes 4 and 5 are regional leader nodes. When synchronizing files, node 1 downloads update files from the data source, and determines the file information that needs to be updated to nodes 4 and 5 by comparing the local file information of each regional leader node. Then, node 1, as the regional leader node, performs file synchronization by comparing the file differences between itself and nodes 2 and 3. Node 4 performs file synchronization by comparing the file differences between itself and nodes 5 and 6. Node 7 performs file synchronization by comparing the file differences between itself and nodes 8 and 9.

[0067] Each node in the distributed system significantly reduces the amount of data transmission, network resource consumption, network congestion and latency by transmitting only differential file information instead of the entire file. At the same time, the differentiated synchronization strategy reduces unnecessary data transmission and processing, speeds up file synchronization, and enables each node in the distributed system to achieve consistency of file content faster. In addition, the distributed leader node architecture reduces the risk of single point failure. Even if a leader node has a problem, the system can quickly recover and continue to run, ensuring the stability and reliability of file synchronization. In addition, the decentralized node identity configuration can dynamically adapt to changes in network conditions, such as bandwidth fluctuations, delay changes, etc., and optimize performance by adjusting the synchronization frequency and strategy. Therefore, through intelligent node role division and differentiated file synchronization strategies, the efficiency, stability and reliability of file synchronization in distributed systems are significantly improved.

[0068] Optionally, the first file information, the second file information and the third file information all include corresponding file types and file update times;

[0069] Determining first difference file information between the first file information and the second file information of each regional leader node includes:

[0070] Determine a first file increment between the first file information and the second file information of each regional leader node according to the file type and the file update time, and use the first file increment as the first difference file information between the first file information and the second file information of each regional leader node;

[0071] Determining second difference file information between the local second file information and the third file information of the follower nodes in each area of ​​the target area includes:

[0072] The second file increment between the second file information and the third file information of each regional follower node is determined according to the file type and the file update time, and the second file increment is used as the second difference file information between the second file information and the third file information of each regional follower node in the target area.

[0073] The file management module will periodically obtain the previous level leader node from the node discovery module, and then send the file information of this node, including file type, file update time, and file MD5. After the leader node receives the file information of the follower node, it will compare it with its own file information, and then synchronize the different file increments to the next level node.

[0074] Among them, for the main leader node, by comparing the file type and file update time, it is determined which files are in the first file information but not in the second file information of the regional leader node, or which files are in the first file information but have different contents (judged by update time), and these different or newly added file information constitute the first file increment. The first file increment is regarded as the first difference file information between the first file information and the second file information of each regional leader node.

[0075] For the regional leader node, the difference between the local second file information and the third file information of each regional follower node in the target region is determined. By comparing the file type and file update time, the file information that is present in the local second file information but not in the third file information of the regional follower node, or the file information that is present but has different content, is found to form the second file increment. The second file increment is regarded as the second difference file information between the second file information and the third file information of each regional follower node in the target region.

[0076] After receiving the second difference file information synchronized by the regional leader node, the method further includes:

[0077] An MD5 check is performed based on the second difference file information and the local third file information, and a synchronous update process of the third file information is performed based on the MD5 check result.

[0078] After receiving the synchronization file from the regional leader node, the regional follower node will verify whether its MD5 is correct and then write it to the disk. The file management module will also regularly check the MD5 of local files to prevent tampering. If the MD5 is found to be incorrect, it will ask the regional leader node to resynchronize. Similarly, for the regional leader node, when synchronizing files from the main leader node, the above MD5 verification process is also performed. In addition, after the update file of the main leader node is downloaded, its MD5 will also be verified to see if it is correct, and then handed over to the file management module for storage.

[0079] Optionally, synchronizing the first difference file information to the corresponding regional leader node further includes:

[0080] Detecting the node health status of the corresponding regional leader node, and pausing or resuming the synchronization task of the first difference file information according to the node health status;

[0081] Synchronizing the second difference file information to the regional follower node also includes:

[0082] Detect the node health status of the corresponding regional follower node, and suspend or resume the synchronization task of the second difference file information according to the node health status.

[0083] Before synchronizing the first difference file information to the corresponding regional leader node, the health status of these regional leader nodes will be detected first. The node health status may include different status information such as the node's online status, resource usage (such as CPU, memory, disk space, etc.), network connectivity, etc. Based on the detected node health status, the master leader node makes a decision, that is, whether to continue, suspend or resume the synchronization task of the first difference file information. If the node is detected to be in an unhealthy state (for example, excessive resource usage, unstable network or node downtime), the master leader node may suspend the synchronization task to avoid placing a greater burden on the node or synchronization failure. When the regional leader node returns to a healthy state, the master leader node automatically resumes the synchronization task.

[0084] Similarly, before synchronizing the second difference file information to the corresponding regional follower nodes, the regional leader node will also detect the health status of these regional follower nodes. Based on the detected node health status, the regional leader node decides whether to continue, pause, or resume the synchronization task of the second difference file information. If the regional follower node is in an unhealthy state, the regional leader node will pause the synchronization task. After the regional follower node recovers, the synchronization task can be resumed.

[0085] By adding node health status detection and synchronization task management, the synchronization task of differential file information can be handled more intelligently, thereby improving the success rate of file synchronization and avoiding unnecessary burden on unhealthy nodes.

[0086] In the above, by obtaining the node startup time of each regional node in the distributed system, the node identity information of each regional node is determined based on the node startup time, the node identity information includes regional follower nodes and regional leader nodes corresponding to different target areas, and master leader nodes corresponding to each regional leader node, the regional leader node is used to communicate with other regional follower nodes in the target area to which it belongs, and the master leader node is used to communicate with each regional leader node; in the case where the node identity information of the local machine is the master leader node, the update file of the data source is received and the update file is stored in the local first file information, the first difference file information between the first file information and the second file information of each regional leader node is determined, and the first difference file information is synchronized to the corresponding regional leader node; in the case where the node identity information of the local machine is the regional leader node, the second difference file information between the local second file information and the third file information of each regional follower node in the target area to which it belongs is determined, and the second difference file information is synchronized to the regional follower node; in the case where the node identity information of the local machine is the regional follower node, the local third file information is reported to the regional leader node of the target area to which it belongs, and the second difference file information synchronized by the regional leader node is received. By adopting the above-mentioned technical means, the node identity information of each regional node is determined by the node startup time, and the file synchronization of the distributed system is performed based on the node identity information, thereby realizing decentralized file synchronization. When the network environment changes, the file can also be efficiently synchronized by changing the node identity information, simplifying the cumbersome process of file synchronization between nodes in the distributed system and improving the stability and reliability of file synchronization.

[0087] Based on the above embodiments, Figure 4 A schematic diagram of the structure of a distributed file synchronization system provided by this application. Figure 4 The distributed file synchronization system provided in this embodiment specifically includes: an identity determination module 21, a first synchronization module 22, a second synchronization module 23 and a third synchronization module 24.

[0088] The identity determination module 21 is configured to obtain the node startup time of each regional node in the distributed system, and determine the node identity information of each regional node based on the node startup time. The node identity information includes regional follower nodes and regional leader nodes corresponding to different target regions, and a master leader node corresponding to each regional leader node. The regional leader node is used to communicate with other regional follower nodes in the target region to which it belongs, and the master leader node is used to communicate with each regional leader node.

[0089] The first synchronization module 22 is configured to receive the update file of the data source and store the update file in the local first file information when the node identity information of the local machine is the primary leader node, determine the first difference file information between the first file information and the second file information of each regional leader node, and synchronize the first difference file information to the corresponding regional leader node;

[0090] The second synchronization module 23 is configured to determine the second difference file information between the local second file information and the third file information of each regional follower node in the target region when the node identity information of the local node is the regional leader node, and synchronize the second difference file information to the regional follower node;

[0091] The third synchronization module 24 is configured to report the local third file information to the regional leader node of the target area and receive the second difference file information synchronized by the regional leader node when the node identity information of the local device is a regional follower node.

[0092] Specifically, the node identity information of each regional node is determined based on the node startup time, including:

[0093] Construct a small top heap data structure based on the node startup time of each regional node;

[0094] The small top heap data structure is queried based on the pre-built comparator, and the node identity information of each regional node is determined according to the position of each regional node in the small top heap data structure.

[0095] After building the small top heap data structure based on the node startup time of each regional node, it also includes:

[0096] When a new regional node is detected, the new regional node is inserted at the end of the queue of the small top heap data structure based on the node startup time, and the small top heap data structure is rebuilt;

[0097] When the removal of the specified area node is detected, if the specified area node is at the head of the mini-heap data structure, the specified area node is removed based on the node pop-up function; if the specified area node is not at the head of the mini-heap data structure, the specified area node is directly removed and the mini-heap data structure is rebuilt.

[0098] Specifically, the first file information, the second file information and the third file information all include corresponding file types and file update times;

[0099] Determining first difference file information between the first file information and the second file information of each regional leader node includes:

[0100] Determine a first file increment between the first file information and the second file information of each regional leader node according to the file type and the file update time, and use the first file increment as the first difference file information between the first file information and the second file information of each regional leader node;

[0101] Determining second difference file information between the local second file information and the third file information of the follower nodes in each area of ​​the target area includes:

[0102] The second file increment between the second file information and the third file information of each regional follower node is determined according to the file type and the file update time, and the second file increment is used as the second difference file information between the second file information and the third file information of each regional follower node in the target area.

[0103] Specifically, synchronizing the first difference file information to the corresponding regional leader node also includes:

[0104] Detecting the node health status of the corresponding regional leader node, and pausing or resuming the synchronization task of the first difference file information according to the node health status;

[0105] Synchronizing the second difference file information to the regional follower node also includes:

[0106] Detect the node health status of the corresponding regional follower node, and suspend or resume the synchronization task of the second difference file information according to the node health status.

[0107] Specifically, after receiving the second difference file information synchronized by the regional leader node, the method further includes:

[0108] An MD5 check is performed based on the second difference file information and the local third file information, and a synchronous update process of the third file information is performed based on the MD5 check result.

[0109] In the above, by obtaining the node startup time of each regional node in the distributed system, the node identity information of each regional node is determined based on the node startup time, the node identity information includes regional follower nodes and regional leader nodes corresponding to different target areas, and master leader nodes corresponding to each regional leader node, the regional leader node is used to communicate with other regional follower nodes in the target area to which it belongs, and the master leader node is used to communicate with each regional leader node; in the case where the node identity information of the local machine is the master leader node, the update file of the data source is received and the update file is stored in the local first file information, the first difference file information between the first file information and the second file information of each regional leader node is determined, and the first difference file information is synchronized to the corresponding regional leader node; in the case where the node identity information of the local machine is the regional leader node, the second difference file information between the local second file information and the third file information of each regional follower node in the target area to which it belongs is determined, and the second difference file information is synchronized to the regional follower node; in the case where the node identity information of the local machine is the regional follower node, the local third file information is reported to the regional leader node of the target area to which it belongs, and the second difference file information synchronized by the regional leader node is received. By adopting the above-mentioned technical means, the node identity information of each regional node is determined by the node startup time, and the file synchronization of the distributed system is performed based on the node identity information, thereby realizing decentralized file synchronization. When the network environment changes, the file can also be efficiently synchronized by changing the node identity information, simplifying the cumbersome process of file synchronization between nodes in the distributed system and improving the stability and reliability of file synchronization.

[0110] The distributed file synchronization system provided in the embodiment of the present application can be configured to execute the distributed file synchronization method provided in the above embodiment, and has corresponding functions and beneficial effects.

[0111] Based on the above practical example, the present application embodiment also provides a distributed file synchronization device, referring to Figure 5, the distributed file synchronization device includes: a processor 31, a memory 32, a communication module 33, an input device 34 and an output device 35. The memory, as a computer-readable storage medium, can be configured to store software programs, computer executable programs and modules, such as program instructions / modules corresponding to the distributed file synchronization method described in any embodiment of the present application (for example, an identity determination module, a first synchronization module, a second synchronization module and a third synchronization module in a distributed file synchronization system). The communication module is configured to perform data transmission. The processor executes various functional applications and data processing of the device by running the software programs, instructions and modules stored in the memory, that is, to implement the above-mentioned distributed file synchronization method. The input device can be configured to receive input digital or character information, and generate key signal input related to user settings and function control of the device. The output device may include a display device such as a display screen. The above-mentioned distributed file synchronization device can be configured to execute the distributed file synchronization method provided in the above-mentioned embodiment, and has corresponding functions and beneficial effects.

[0112] On the basis of the above embodiments, the embodiments of the present application further provide a non-volatile computer-readable storage medium, wherein the non-volatile computer-readable storage medium stores computer-executable instructions, wherein the computer-executable instructions are configured to execute a distributed file synchronization method when executed by a computer processor, and the storage medium may be any of various types of memory devices or storage devices. Of course, the non-volatile computer-readable storage medium provided in the embodiments of the present application, whose computer-executable instructions are not limited to the distributed file synchronization method described above, can also execute related operations in the distributed file synchronization method provided in any embodiment of the present application.

[0113] On the basis of the above embodiments, the embodiments of the present application also provide a computer program product. The technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product. The computer program product is stored in a storage medium, and includes a number of instructions for enabling a computer device, a mobile terminal or a processor therein to execute all or part of the steps of the distributed file synchronization method described in each embodiment of the present application.

Claims

1. A distributed file synchronization method, characterized in that: include: Obtaining the node startup time of each regional node in the distributed system, and determining the node identity information of each regional node based on the node startup time, wherein the node identity information includes regional follower nodes and regional leader nodes corresponding to different target regions, and a master leader node corresponding to each of the regional leader nodes, wherein the regional leader node is used to communicate with other regional follower nodes in the target region to which it belongs, and the master leader node is used to communicate with each of the regional leader nodes; In the case where the node identity information of the local machine is the primary leader node, receiving an update file from a data source and storing the update file in a local first file information, determining first difference file information between the first file information and the second file information of each of the regional leader nodes, and synchronizing the first difference file information to the corresponding regional leader node; In the case where the node identity information of the local device is the regional leader node, determining second difference file information between the local second file information and the third file information of each of the regional follower nodes in the target region, and synchronizing the second difference file information to the regional follower nodes; When the node identity information of the local computer is the regional follower node, the local third file information is reported to the regional leader node of the target area, and the second difference file information synchronized by the regional leader node is received.

2. The distributed file synchronization method according to claim 1, characterized in that: The determining the node identity information of each regional node based on the node startup time includes: Constructing a small top heap data structure based on the node startup time of each regional node; The small top heap data structure is queried based on a pre-built comparator, and the node identity information of each regional node is determined according to the position of each regional node in the small top heap data structure.

3. The distributed file synchronization method according to claim 2, characterized in that: After constructing the small top heap data structure based on the node startup time of each regional node, it also includes: When a new regional node is detected, the new regional node is inserted at the end of the queue of the small top heap data structure based on the node startup time, and the small top heap data structure is rebuilt; When the removal of a designated area node is detected, if the designated area node is located at the head of the mini-heap data structure, the designated area node is removed based on a node pop-up function; if the designated area node is not at the head of the mini-heap data structure, the designated area node is directly removed and the mini-heap data structure is reconstructed.

4. The distributed file synchronization method according to claim 1, characterized in that: The first file information, the second file information and the third file information all include corresponding file types and file update times; The determining of first difference file information between the first file information and the second file information of each of the regional leader nodes includes: Determine a first file increment between the first file information and the second file information of each of the regional leader nodes according to the file type and the file update time, and use the first file increment as first difference file information between the first file information and the second file information of each of the regional leader nodes; The determining of second difference file information between the local second file information and the third file information of each of the regional follower nodes in the target region includes: A second file increment between the second file information and the third file information of each of the regional follower nodes is determined according to the file type and the file update time, and the second file increment is used as the second difference file information between the second file information and the third file information of each of the regional follower nodes in the target area.

5. The distributed file synchronization method according to any one of claims 1 to 4, characterized in that: The step of synchronizing the first difference file information to the corresponding regional leader node further includes: Detecting the node health status of the corresponding regional leader node, and pausing or resuming the synchronization task of the first difference file information according to the node health status; The step of synchronizing the second difference file information to the area follower node further includes: Detect the node health status of the corresponding regional follower node, and suspend or resume the synchronization task of the second difference file information according to the node health status.

6. The distributed file synchronization method according to any one of claims 1 to 4, characterized in that: After receiving the second difference file information synchronized by the regional leader node, the method further includes: An MD5 check is performed based on the second difference file information and the local third file information, and a synchronous update process of the third file information is performed based on the MD5 check result.

7. A distributed file synchronization system, characterized in that: include: an identity determination module, configured to obtain a node startup time of each regional node in the distributed system, and determine node identity information of each regional node based on the node startup time, wherein the node identity information includes regional follower nodes and regional leader nodes corresponding to different target regions, and a primary leader node corresponding to each of the regional leader nodes, wherein the regional leader node is used to communicate with other regional follower nodes in the target region to which it belongs, and the primary leader node is used to communicate with each of the regional leader nodes; A first synchronization module is configured to, when the node identity information of the local machine is the primary leader node, receive an update file from a data source and store the update file in a local first file information, determine first difference file information between the first file information and the second file information of each of the regional leader nodes, and synchronize the first difference file information to the corresponding regional leader node; A second synchronization module is configured to determine second difference file information between the local second file information and the third file information of each of the regional follower nodes in the target region when the node identity information of the local device is the regional leader node, and synchronize the second difference file information to the regional follower node; The third synchronization module is configured to report the local third file information to the regional leader node of the target area when the node identity information of the local machine is the regional follower node, and receive the second difference file information synchronized by the regional leader node.

8. A distributed file synchronization device, characterized in that: include: memory and one or more processors; The memory is configured to store one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the distributed file synchronization method as described in any one of claims 1-6.

9. A non-volatile computer-readable storage medium, characterized in that: The non-volatile computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are executed by a computer processor, they are configured to execute the distributed file synchronization method according to any one of claims 1 to 6.

10. A computer program product, characterized in that The computer program product includes instructions, and when the instructions are executed on a computer or a processor, the computer or the processor executes the distributed file synchronization method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Data synchronization method and device, electronic equipment and computer readable storage medium

    CN111259072A

  • Containerized resource dynamic allocation method for power distribution network system and electronic equipment

    CN119127388A

  • Network system

    JP2012028931A

  • Distributed processing system, distributed processing device, distributed processing method, and distributed processing program

    US20160070591A1