File synchronization method, device and medium
By taking a synchronization snapshot of the main directory files, confirming the synchronization type and calling multi-threaded synchronization, the problem of low synchronization efficiency in the distributed file system is solved, and efficient data replication and data consistency are achieved.
Patent Information
- Application Number
- CN202310018869.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-01-06
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2043-01-06
AI Technical Summary
In the existing distributed file system, a single-threaded approach is used when synchronizing difference points between a master and a slave cluster, resulting in low synchronization efficiency and failing to meet the requirements of efficient data replication.
By taking a synchronization snapshot of the files in the main directory, we get the target snapshot and compare it with the most recently synchronized baseline snapshot to confirm the synchronization type. We then call multiple threads to synchronize the files and retry the synchronization of the failed differences until all are successful.
It significantly improves the efficiency of data synchronization, ensures the consistency and availability of master-slave cluster data, and enhances the continuity and recoverability of data storage.
Smart Images

Figure CN116383161B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of distributed storage, and in particular to a file synchronization method, device, and medium. Background Art
[0002] With the digitization of information, data has become the core of users' businesses, and users' requirements for the stability of the storage systems that carry data are gradually increasing. Although most storage vendors can provide users with highly stable storage devices, they cannot prevent various natural disasters from causing irreparable damage to production systems. There are many ways to protect data, the most common of which is to replicate our data and store it in multiple locations. If a failure in one location causes data loss, we can still recover the data from the other locations, thus ensuring the security of our data. To ensure the continuity, recoverability, and high availability of data storage, remote replication technology has emerged.
[0003] Remote replication requires the fastest possible data transfer speed to improve replication efficiency. Concurrent replication can significantly improve replication efficiency, and the current mainstream method is to transfer snapshots. A snapshot is defined as a fully usable copy of a specified data set, containing an image of the data at a specific point in time (the point in time when the copy is started).
[0004] Each remote replication operation essentially synchronizes a snapshot of the primary cluster to the secondary cluster. The difference between the primary and secondary clusters is essentially the difference between the two snapshots. Because distributed file system services have a strict time sequence, synchronizing differences between primary and secondary clusters using a completely sequential method requires only a single thread, resulting in low synchronization efficiency.
[0005] It can be seen that how to improve synchronization efficiency during data synchronization is an urgent problem to be solved by those skilled in the art. Summary of the Invention
[0006] The purpose of this application is to provide a file synchronization method, device and medium for improving synchronization efficiency during data synchronization.
[0007] To solve the above technical problems, the present application provides a file synchronization method, including:
[0008] Synchronize the files in the home directory to get the target snapshot;
[0009] Determine the synchronization type based on the correspondence between the target snapshot and the reference snapshot; the reference snapshot is a data backup at the most recent successful synchronization point in time;
[0010] According to the synchronization type, multiple threads are called to synchronize the files.
[0011] Preferably, the determining the synchronization type according to the correspondence between the target snapshot and the reference snapshot includes:
[0012] Identify the difference between the target snapshot and the baseline snapshot; wherein the difference is the index node of the changed file;
[0013] The synchronization type is determined according to the distribution of the difference points in the target snapshot and the reference snapshot.
[0014] Preferably, it also includes:
[0015] If synchronization fails at the difference point, the difference point at which synchronization fails is recorded;
[0016] After waiting for the current synchronization to be completed, retry the synchronization of the difference points that failed to be synchronized until all the difference points are synchronized.
[0017] Preferably, determining the synchronization type according to the distribution of the difference points in the target snapshot and the reference snapshot includes:
[0018] If the difference point does not exist in the reference snapshot but exists in the target snapshot, confirming that the synchronization type is a new operation;
[0019] If the difference point exists in the reference snapshot but does not exist in the target snapshot, confirming that the synchronization type is a delete operation;
[0020] If the difference point exists in both the base snapshot and the target snapshot, the synchronization type is determined to be a modification operation.
[0021] Preferably, after all the difference points are synchronized, the method further includes:
[0022] Take a snapshot of the files in the master directory of the slave cluster to obtain a slave directory snapshot;
[0023] Determine whether the name of the secondary directory snapshot is consistent with the name of the target snapshot;
[0024] If they are consistent, the synchronization is successful.
[0025] Preferably, if the name of the secondary directory snapshot is inconsistent with the name of the target snapshot, the method further includes:
[0026] The difference in name inconsistency is recorded and an alarm message is sent.
[0027] Preferably, it also includes:
[0028] The recorded differences are stored in a cloud server.
[0029] To solve the above technical problems, the present application also provides a file synchronization device, comprising:
[0030] A first snapshot module is used to synchronize snapshots of files in the main directory to obtain a target snapshot;
[0031] A confirmation module, configured to confirm a synchronization type based on a correspondence between the target snapshot and a reference snapshot; the reference snapshot is a data backup at the most recent successful synchronization point in time;
[0032] The synchronization module is used to call multiple threads to synchronize the files according to the synchronization type.
[0033] To solve the above technical problems, the present application also provides another file synchronization device, comprising a memory for storing a computer program;
[0034] The processor is configured to implement the steps of the above-mentioned file synchronization method when executing the computer program.
[0035] In order to solve the above technical problems, the present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the file synchronization method as described above are implemented.
[0036] The file synchronization method provided by the embodiment of the present application obtains a target snapshot by taking a synchronization snapshot of the files in the main directory; confirms the synchronization type based on the correspondence between the target snapshot and the baseline snapshot; the baseline snapshot is a data backup at the time point of the most recent successful synchronization; and calls multiple threads to synchronize the files based on the synchronization type. Compared with the current technology, which uses a single thread for data synchronization and results in slow synchronization efficiency, the present technical solution takes a snapshot of the files in the main cluster and compares it with the baseline snapshot obtained by the data backup at the time point of the most recent successful synchronization, thereby confirming the correspondence between the target snapshot and the baseline snapshot, confirming the synchronization type based on the correspondence between the target snapshot and the baseline snapshot, and calling multiple threads for data synchronization based on the synchronization type, thereby improving synchronization efficiency.
[0037] In addition, the file synchronization device and medium provided in this application correspond to the above-mentioned file synchronization method and have the same effect as above. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0039] Figure 1A flowchart of a file synchronization method provided in an embodiment of the present application;
[0040] Figure 2 A synchronization method for synchronization failure difference points provided in an embodiment of the present application;
[0041] Figure 3 A structural diagram of a file synchronization device provided in an embodiment of the present application;
[0042] Figure 4 A structural diagram of another file synchronization device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0043] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0044] With the digitization of information, data has become the core of users' businesses, and users' requirements for the stability of the storage systems that carry data are gradually increasing. Although most storage vendors can provide users with highly stable storage devices, they cannot prevent various natural disasters from causing irreparable damage to production systems. There are many ways to protect data, the most common of which is to replicate our data and store it in multiple locations. If a failure in one location causes data loss, we can still recover the data from the other locations, thus ensuring the security of our data. To ensure the continuity, recoverability, and high availability of data storage, remote replication technology has emerged.
[0045] Remote replication requires the fastest possible data transfer speed to improve replication efficiency. Concurrent replication can significantly improve replication efficiency, and the current mainstream method is to transfer snapshots. A snapshot is defined as a fully usable copy of a specified data set, containing an image of the data at a specific point in time (the point in time when the copy is started).
[0046] Each remote replication operation essentially synchronizes a snapshot of the master cluster to the slave cluster. The differences between the master and slave clusters are essentially the differences between the two snapshots. Because distributed file system services have strict time order, synchronizing differences between master and slave clusters using a completely sequential approach requires only a single thread, resulting in low synchronization efficiency. Multi-threaded synchronization, due to its concurrent nature, cannot achieve completely sequential synchronization of differences, resulting in synchronization failures for many differences that require sequential operations.
[0047] It can be seen that how to improve synchronization efficiency during data synchronization is an urgent problem to be solved by those skilled in the art.
[0048] The core of this application is to provide a file synchronization method, device and medium for improving synchronization efficiency during data synchronization.
[0049] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.
[0050] Figure 1 A flowchart of a file synchronization method provided in an embodiment of the present application is shown in FIG. Figure 1 As shown, the method includes:
[0051] S10: Synchronize snapshots of files in the home directory to obtain a target snapshot;
[0052] S11: Confirm the synchronization type based on the correspondence between the target snapshot and the baseline snapshot; the baseline snapshot is the data backup at the most recent successful synchronization point in time;
[0053] S12: Based on the synchronization type, multiple threads are called to synchronize the files.
[0054] The definition of a snapshot is: a fully usable copy of a specified data set, which includes an image of the corresponding data at a certain point in time (the point in time when the copy starts). In this application, when each file in the file system is changed, a difference point is recorded. The remote replication function performs remote synchronization based on the difference point, and the synchronization order is not strictly sequential. When synchronizing certain file operations that must be performed sequentially, if a failure occurs, the failed difference point is recorded and a retry is performed. If the first retry fails, the failed difference point will continue to be recorded and wait for the second retry until all difference points are successfully synchronized.
[0055] In this application, when a file in the file system changes, such as adding, deleting, modifying, etc., the inode of the changed file will be recorded as a difference point. These difference points are logically sequential. For example, if a folder is added and a file is added under this folder, the difference point of creating the folder must be before creating the file.
[0056] In this embodiment, a target snapshot is obtained by synchronizing snapshots of files in the main directory. The target snapshot is the modified snapshot of the files that need to be synchronized to other clusters. When synchronization starts, multiple threads read the difference points concurrently, and select the files that need to be synchronized based on the difference points and synchronize them to the remote end. The difference point is the difference between the baseline snapshot and the target snapshot. The baseline snapshot is a data backup at the time of the last successful synchronization. The latest status of the slave end is consistent with the baseline snapshot of the master end. The purpose of synchronization is to calculate the synchronization type based on the difference points between the snapshots, and then synchronize. When the difference point does not exist on the baseline snapshot but exists on the target snapshot, it is explained as a new operation. When the difference point exists on the baseline snapshot but does not exist on the target snapshot, it is explained as a delete operation. When the difference point exists on both the baseline snapshot and the target snapshot, it is explained as a modification operation.
[0057] In addition, in the specific implementation, after the synchronization is successful, a snapshot can be taken on the slave side. The snapshot name is consistent with the target snapshot name on the master side. At this time, the master-slave cluster achieves data consistency.
[0058] In practice, during concurrent remote synchronization, strict order cannot be guaranteed, and differences may be synchronized prematurely. For example, synchronizing a file first and then creating its parent directory will cause synchronization to fail because the parent directory does not exist. When this fails, the differences should be recorded and retried after the first round of synchronization completes. After the first round of retries, there may still be differences that failed to synchronize. In this case, these differences should be recorded and awaited in the second round of retries. If this still fails, proceed to the next round. Because the differences with the highest business order in each retry round will always synchronize successfully, the number of differences in each retry round will decrease, ultimately leading to success.
[0059] Synchronization failures can occur in a variety of situations, including those described above. For example, if a directory still contains files, the deletion will fail and a difference will be recorded. This is not discussed here. In principle, add operations must wait for the parent directory to complete synchronization, and delete operations must wait for the child files to be successfully deleted.
[0060] The file synchronization method proposed in this application uses the index node of the file that records the business changes on the master cluster as the difference point, and performs concurrent synchronization on the slave cluster based on the difference point. When a difference point fails to synchronize, it is recorded and waits for the next round of synchronization retry. The synchronization may go through multiple rounds, and the synchronization is completed after all the difference points are synchronized successfully. This method can significantly improve the synchronization efficiency, thereby improving the efficiency of incremental replication. The inode that records the file changes is the difference point, and it can be retried when the master-slave synchronization difference point fails. Finally, after retrying, the master-slave data tends to be consistent.
[0061] The file synchronization method provided by the embodiment of the present application obtains a target snapshot by taking a synchronization snapshot of the files in the main directory; confirms the synchronization type based on the correspondence between the target snapshot and the baseline snapshot; the baseline snapshot is a data backup at the time point of the most recent successful synchronization; and calls multiple threads to synchronize the files based on the synchronization type. Compared with the current technology, which uses a single thread for data synchronization and results in slow synchronization efficiency, the present technical solution takes a snapshot of the files in the main cluster and compares it with the baseline snapshot obtained by the data backup at the time point of the most recent successful synchronization, thereby confirming the correspondence between the target snapshot and the baseline snapshot, confirming the synchronization type based on the correspondence between the target snapshot and the baseline snapshot, and calling multiple threads for data synchronization based on the synchronization type, thereby improving synchronization efficiency.
[0062] When files in the file system are added, deleted, or modified, the inode of the changed file is recorded as a difference point. These difference points are logically sequential. For example, if a folder is added and a file is added to this folder, the difference point of creating the folder must be before creating the file.
[0063] As described in the above embodiment, when a file in the file system is changed, the difference points are recorded. The distribution of the difference points can be confirmed by comparing the baseline snapshot and the target snapshot. In specific implementations, different distributions of the difference points represent different change operations, which will also be different during synchronization.
[0064] Therefore, in this embodiment, determining the synchronization type according to the correspondence between the target snapshot and the reference snapshot includes:
[0065] Identify the differences between the target snapshot and the baseline snapshot; the differences are the inodes of the changed files;
[0066] Determine the synchronization type based on the distribution of differences between the target snapshot and the baseline snapshot.
[0067] Specifically, the synchronization type is determined based on the distribution of differences between the target snapshot and the baseline snapshot. The following types are used:
[0068] If the difference point does not exist on the baseline snapshot but exists on the target snapshot, confirm that the synchronization type is a new operation;
[0069] If the difference point exists on the baseline snapshot but not on the target snapshot, confirm that the synchronization type is a delete operation;
[0070] If differences exist on both the base snapshot and the target snapshot, confirm that the synchronization type is a modification operation.
[0071] This technical solution determines the distribution of difference points. If a difference point does not exist in the baseline snapshot but exists in the target snapshot, it indicates an add operation. If a difference point exists in the baseline snapshot but not in the target snapshot, it indicates a delete operation. If a difference point exists in both the baseline and target snapshots, it indicates a modify operation, thus ensuring accurate operation determination.
[0072] In practice, each remote replication operation essentially synchronizes a snapshot of the primary cluster to the secondary cluster. The differences between the primary and secondary clusters are essentially the differences between the two snapshots. Because distributed file system services have strict time order, synchronizing differences between primary and secondary clusters using a completely sequential approach requires only a single thread, resulting in far lower synchronization efficiency than a multi-threaded approach. Multi-threaded synchronization, due to its simultaneous nature, cannot achieve completely sequential synchronization of differences, resulting in synchronization failures for many differences that must be processed sequentially.
[0073] Therefore, based on the above embodiment, this embodiment further includes:
[0074] If synchronization fails due to a difference, the difference point where synchronization fails will be recorded;
[0075] After the current synchronization is completed, retry the synchronization of the difference points that failed to be synchronized until all the difference points are synchronized.
[0076] Different file modification operations are synchronized differently. During concurrent synchronization, strict order cannot be guaranteed, and synchronization may occur earlier than the previous one. For example, if you synchronize the file first and then create the file's parent directory, synchronization will fail because the file's parent directory does not exist. Figure 2 A synchronization method for synchronization failure difference points provided in an embodiment of the present application is as follows Figure 2 As shown, when a failure occurs, the difference point of the synchronization failure is recorded, and the difference point is retried after the first round of synchronization is completed.
[0077] After the first round of retries, there may still be discrepancies that failed synchronization. These should be recorded and awaited in the second round of retries. If there are still failures, the next round will proceed. Because the most prioritized discrepancies in each retry round will always succeed, the number of discrepancies decreases with each retry, ultimately leading to full success.
[0078] Synchronization failures can occur in a variety of situations, including those described above. For example, if a directory still contains files, the deletion will fail and a difference will be recorded. In principle, additions must wait for the parent directory to be synchronized, and deletions must wait for the child files to be successfully deleted.
[0079] The file synchronization method provided in this embodiment records any differences in synchronization failures and waits for the next synchronization retry. Synchronization may proceed through multiple rounds, completing the synchronization process only after all differences have been successfully synchronized. This approach significantly improves synchronization efficiency, thereby increasing the efficiency of incremental replication. Changed inodes are recorded as differences, allowing retries to be made when master-slave synchronization fails. Ultimately, after retries, the master-slave data converges.
[0080] In a specific implementation, after the synchronization is completed, in order to ensure the completion of the synchronization and avoid data loss, the synchronization result can also be verified. In this embodiment, after all the difference points are synchronized, it also includes:
[0081] Take a snapshot of the files in the master directory of the slave cluster to obtain a slave directory snapshot;
[0082] Determine whether the name of the secondary directory snapshot is consistent with the name of the target snapshot;
[0083] If they are consistent, the synchronization is successful.
[0084] Furthermore, in a specific implementation, if the name of the slave directory snapshot is inconsistent with the name of the target snapshot, the method further includes: recording the name inconsistency difference and sending an alarm message; and further includes: storing the recorded difference in a cloud server.
[0085] In the above embodiments, the file synchronization method is described in detail. This application also provides corresponding embodiments of the file synchronization device. It should be noted that this application describes the embodiments of the device from two perspectives: one is based on the functional module perspective, and the other is based on the hardware perspective.
[0086] Figure 3 A structural diagram of a file synchronization device provided in an embodiment of the present application is shown as follows: Figure 3 As shown, the device includes:
[0087] A first snapshot module 10 is used to synchronize snapshots of files in the main directory to obtain a target snapshot;
[0088] Confirmation module 11, used to confirm the synchronization type according to the corresponding relationship between the target snapshot and the reference snapshot; the reference snapshot is the data backup at the time point of the most recent successful synchronization;
[0089] The synchronization module 12 is used to call multiple threads to synchronize files according to the synchronization type.
[0090] A snapshot is a fully usable copy of a specified data set, including an image of the corresponding data at a specific point in time (the point in time when the copy begins). The present invention records the difference points when each file in the file system is modified. The remote replication function performs remote synchronization based on the difference points, and the synchronization order is not strictly sequential. When synchronizing certain file operations that must be executed sequentially, if a failure occurs, the failure difference points are recorded and a retry is performed. If the first retry fails, the failure difference points are recorded again and a second retry is performed until all difference points are successfully synchronized.
[0091] When files in the file system are added, deleted, or modified, the inode of the changed file is recorded as a difference point. These difference points are logically sequential. For example, if a folder is added and a file is added to this folder, the difference point of creating the folder must be before creating the file.
[0092] At the start of synchronization, multiple threads concurrently read the difference points and select the files to be synchronized to the remote server based on the difference points. The difference points are the differences between the baseline snapshot and the target snapshot. The baseline snapshot is a backup of the data at the time of the last successful synchronization. The latest state of the slave is consistent with the baseline snapshot on the master. The purpose of synchronization is to calculate the synchronization type based on the difference points between the snapshots and then synchronize. After the synchronization is successful, a snapshot is created on the slave. The snapshot name matches the target snapshot name on the master, and the master and slave clusters are now data consistent.
[0093] Since the embodiments of the apparatus part correspond to the embodiments of the method part, please refer to the description of the embodiments of the method part for the embodiments of the apparatus part, and they will not be repeated here.
[0094] The file synchronization device provided by the embodiment of the present application obtains a target snapshot by taking a synchronization snapshot of the files in the main directory; confirms the synchronization type based on the correspondence between the target snapshot and the baseline snapshot; the baseline snapshot is a data backup at the time point of the most recent successful synchronization; and calls multiple threads to synchronize the files based on the synchronization type. Compared with the current technology, which uses a single thread for data synchronization and results in slow synchronization efficiency, the present technical solution takes a snapshot of the files in the main cluster and compares it with the baseline snapshot obtained by the data backup at the time point of the most recent successful synchronization, thereby confirming the correspondence between the target snapshot and the baseline snapshot, confirming the synchronization type based on the correspondence between the target snapshot and the baseline snapshot, and calling multiple threads for data synchronization based on the synchronization type, thereby improving synchronization efficiency.
[0095] Figure 4 A structural diagram of another file synchronization device provided in an embodiment of the present application, such as Figure 4 As shown, the device includes: a memory 20 for storing computer programs;
[0096] The processor 21 is configured to implement the steps of the file synchronization method of the above embodiment when executing a computer program.
[0097] The file synchronization device provided in this embodiment may include but is not limited to a smart phone, a tablet computer, a laptop computer, or a desktop computer.
[0098] Among them, the processor 21 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 21 can be implemented in at least one hardware form of a digital signal processor (DSP), a field programmable gate array (FPGA), and a programmable logic array (PLA). The processor 21 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the awake state, also known as a central processing unit (CPU); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 21 may be integrated with a graphics processing unit (GPU), which is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 21 may also include an artificial intelligence (AI) processor, which is used to process computing operations related to machine learning.
[0099] The memory 20 may include one or more computer-readable storage media, which may be non-transitory. The memory 20 may also include a high-speed random access memory, and a non-volatile memory, such as one or more disk storage devices, flash memory storage devices. In this embodiment, the memory 20 is at least used to store the following computer program 201, wherein, after the computer program is loaded and executed by the processor 21, it can implement the relevant steps of the file synchronization method disclosed in any of the aforementioned embodiments. In addition, the resources stored in the memory 20 may also include an operating system 202 and data 203, etc., and the storage method may be temporary storage or permanent storage. Among them, the operating system 202 may include Windows, Unix, Linux, etc. The data 203 may include but is not limited to target snapshots, baseline snapshots, etc.
[0100] In some embodiments, the file synchronization device may further include a display screen 22 , an input / output interface 23 , a communication interface 24 , a power supply 25 , and a communication bus 26 .
[0101] Those skilled in the art will understand that Figure 4 The structure shown in the figure does not constitute a limitation to the file synchronization device, and may include more or fewer components than shown in the figure.
[0102] The file synchronization device provided in an embodiment of the present application includes a memory and a processor. When the processor executes a program stored in the memory, it can implement the following method: synchronize snapshots of files in the main directory to obtain a target snapshot; confirm the synchronization type based on the correspondence between the target snapshot and the baseline snapshot; the baseline snapshot is a data backup at the time point of the most recent successful synchronization; and call multiple threads to synchronize files according to the synchronization type.
[0103] The file synchronization device provided by the embodiment of the present application obtains a target snapshot by taking a synchronization snapshot of the files in the main directory; confirms the synchronization type based on the correspondence between the target snapshot and the baseline snapshot; the baseline snapshot is a data backup at the time point of the most recent successful synchronization; and calls multiple threads to synchronize the files based on the synchronization type. Compared with the current technology, which uses a single thread for data synchronization and results in slow synchronization efficiency, the present technical solution takes a snapshot of the files in the main cluster and compares it with the baseline snapshot obtained by the data backup at the time point of the most recent successful synchronization, thereby confirming the correspondence between the target snapshot and the baseline snapshot, confirming the synchronization type based on the correspondence between the target snapshot and the baseline snapshot, and calling multiple threads for data synchronization based on the synchronization type, thereby improving synchronization efficiency.
[0104] Finally, the present application also provides an embodiment corresponding to a computer-readable storage medium. The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps described in the above method embodiment.
[0105] It is understandable that if the method in the above embodiment is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and executes all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0106] The computer-readable storage medium provided by the embodiment of the present application obtains a target snapshot by taking a synchronization snapshot of the files in the main directory; confirms the synchronization type according to the correspondence between the target snapshot and the baseline snapshot; the baseline snapshot is a data backup at the time point of the most recent successful synchronization; and calls multiple threads to synchronize the files according to the synchronization type. Compared with the current technology, which uses a single thread for data synchronization and results in slow synchronization efficiency, the present technical solution takes a snapshot of the files in the main cluster and compares it with the baseline snapshot obtained by the data backup at the time point of the most recent successful synchronization, thereby confirming the correspondence between the target snapshot and the baseline snapshot, confirming the synchronization type according to the correspondence between the target snapshot and the baseline snapshot, and calling multiple threads for data synchronization according to the synchronization type, thereby improving synchronization efficiency.
[0107] The above is a detailed introduction to the file synchronization method, device and medium provided by the present application. The various embodiments in the specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same and similar parts between the various embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part description. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of this application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the scope of protection of the claims of this application.
[0108] It should also be noted that, in this specification, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variants thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus comprising the element.
Claims
1. A file synchronization method, characterized in that: include: Synchronize the files in the home directory to get the target snapshot; The target snapshot is a modified file snapshot that needs to be synchronized to other slave clusters; Determining a synchronization type based on a correspondence between the target snapshot and the baseline snapshot; The benchmark snapshot is a data backup at the time point of the most recent successful synchronization from the cluster; According to the synchronization type, calling multiple threads to synchronize the files; The step of determining the synchronization type according to the corresponding relationship between the target snapshot and the reference snapshot includes: Identify the difference between the target snapshot and the baseline snapshot; wherein the difference is the index node of the changed file; Determining the synchronization type according to the distribution of the difference points in the target snapshot and the reference snapshot; Determining the synchronization type according to the distribution of the difference points in the target snapshot and the reference snapshot includes: If the difference point does not exist in the reference snapshot but exists in the target snapshot, confirming that the synchronization type is a new operation; If the difference point exists in the reference snapshot but does not exist in the target snapshot, confirming that the synchronization type is a delete operation; If the difference point exists in both the baseline snapshot and the target snapshot, confirming that the synchronization type is a modification operation; Also includes: In the process of calling multiple threads to synchronize the files, if there is a difference point synchronization failure, the difference point that failed synchronization is recorded, and multiple rounds of retries are initiated after the synchronization is completed until all the difference points are synchronized; in each round of retry, the difference point with the highest business order is processed first; After all the difference points are synchronized, the following steps are also included: Take a snapshot of the files in the master directory of the slave cluster to obtain a slave directory snapshot; Determine whether the name of the secondary directory snapshot is consistent with the name of the target snapshot; If they are consistent, the synchronization is confirmed to be successful; If the name of the secondary directory snapshot is inconsistent with the name of the target snapshot, the method further includes: Record the inconsistency in the names and send an alarm message; Also includes: The recorded differences are stored in a cloud server.
2. A file synchronization device, characterized in that: include: A first snapshot module is used to synchronize snapshots of files in the main directory to obtain a target snapshot; The target snapshot is a modified file snapshot that needs to be synchronized to other slave clusters; A confirmation module, configured to confirm a synchronization type according to a correspondence between the target snapshot and the reference snapshot; The benchmark snapshot is a data backup at the time point of the most recent successful synchronization from the cluster; A synchronization module, configured to call multiple threads to synchronize the files according to the synchronization type; The confirmation module is configured to: confirm the difference between the target snapshot and the reference snapshot; wherein the difference is the index node of the changed file; and determine the synchronization type based on the distribution of the difference between the target snapshot and the reference snapshot; The confirmation module is configured to: if the difference point does not exist on the baseline snapshot but exists on the target snapshot, confirm that the synchronization type is a new operation; if the difference point exists on the baseline snapshot but does not exist on the target snapshot, confirm that the synchronization type is a delete operation; if the difference point exists on both the baseline snapshot and the target snapshot, confirm that the synchronization type is a modify operation; The synchronization module is further configured to: in the process of calling multiple threads to synchronize the files, if any difference point fails to be synchronized, record the difference point that failed to be synchronized, and initiate multiple rounds of retries after the synchronization is completed until all the difference points are synchronized; wherein each round of retries prioritizes the difference point with the highest service order; The file synchronization device is also used to: after all the difference points are synchronized, take a snapshot of the files in the main directory of the slave cluster to obtain a slave directory snapshot; determine whether the name of the slave directory snapshot is consistent with the name of the target snapshot; if they are consistent, confirm that the synchronization is successful; if the name of the slave directory snapshot is inconsistent with the name of the target snapshot, record the difference point with inconsistent name and send an alarm message; and store the recorded difference point in the cloud server.
3. A file synchronization device, characterized in that: including a memory for storing a computer program; A processor, configured to implement the steps of the file synchronization method according to claim 1 when executing the computer program.
4. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which implements the steps of the file synchronization method according to claim 1 when executed by a processor.
Citation Information
Patent Citations
Method, device and system for performing data synchronization with Redis server
CN104881494A
Storage cluster data backup method and device, equipment and storage medium
CN114996054A