Fast backup method and system for differential mirror images during file traversal

By using the method of differential mirroring while traversing during data backup, and using thread collaboration between the worker and the disaster recovery machine, the problem of long traversal and transmission time in traditional backup methods is solved, and fast and efficient data backup is achieved.

CN120179462APending Publication Date: 2025-06-20INFORMATION2 SOFTWARE SHANGHAI
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510270951.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

Traditional data backup methods take a long time to traverse and transfer file directory lists, and also transmit data for a long time, which affects backup efficiency.

Method used

A quick backup method of differential mirroring during file traversal is adopted. Through thread collaboration between the worker and the disaster recovery machine, differential mirroring is saved while traversing, saving traversal and transmission time. The specific steps include the worker traversing the thread to scan the file list and sending file information, the disaster recovery machine receives and judges the file differences, transmits the different file blocks, and compares and transmits the differential content on the worker side.

Benefits of technology

It realizes fast data backup with differential mirroring while traversing, significantly saving time in traversing file directory lists and transferring data, and improving backup efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120179462A_ABST
    Figure CN120179462A_ABST
Patent Text Reader

Abstract

The invention discloses a quick backup method and system for a difference mirror image during file traversal, and the method comprises the steps: S1, scanning all file lists under an original directory through a working machine, sequentially obtaining the file information of each to-be-transmitted file, and transmitting the file information to a disaster recovery machine; s2, the disaster recovery machine stores the file information in a file information list; s3, the disaster recovery machine reads the content of the file information list in sequence, obtains MD5 values of all fixed blocks of the different files according to the sizes of the fixed blocks in sequence, and transmits the MD5 values to the working machine; s4, the working machine receives and stores the MD5 value of each fixed block of the same-name file sent by the disaster recovery machine, starts a transmission data thread to sequentially read the received MD5 value of each fixed block, compares the MD5 value with the MD5 value of the corresponding block of the corresponding file, and transmits the blocks with different comparison results to the disaster recovery machine; and S5, the disaster recovery machine receives and writes the same block coverage of the corresponding file.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer disaster recovery backup, and particularly to a fast backup method and system for differential mirroring during file traversal. Background Art

[0002] Data backup is the basis of disaster recovery, which refers to the process of copying all or part of the data set from the hard disk or array of the application host to other storage media to prevent data loss caused by operation errors or system failures in the system.

[0003] Traditional data backup modes mainly include full backup, incremental backup, and differential backup.

[0004] Full backup means backing up a complete set of data each time. When restoring, a full backup data is directly selected for restoration. This backup method only requires restoring one backup during restoration, but the backup time is very long each time because multiple complete data need to be backed up, and it also occupies the disk space of the standby machine.

[0005] Incremental backup only backs up the newly added and different files based on the previous backup each time, without data duplication.

[0006] Differential backup means that except for the first backup which is a full backup, the data of each subsequent backup is the differential data between the current data and the data of the first backup.

[0007] However, no matter which backup method is used, a data comparison and transmission mechanism is required to quickly transmit and backup the files of the working machine to the disaster recovery end. Summary of the Invention

[0008] The purpose of the present invention is to provide a fast backup method and system for differential mirroring during file traversal, which realizes a fast data backup method of differential mirroring while traversing, saving the time used to traverse the file directory list of the working machine end and reducing the time for transmitting data.

[0009] To achieve the above object, the present invention provides a fast backup method for differential mirroring during file traversal, including the following steps: Step S1, the working machine traversal thread scans all file lists in the original directory of the working machine, and sequentially obtains the file information of each file to be transmitted and sends it to the disaster recovery machine; Step S2, the disaster recovery machine receiving thread saves the file information of each file received from the working machine end into the file information list on the disaster recovery machine; In step S3, while the disaster recovery machine writes the file information into the file information list, the disaster recovery machine simultaneously starts a read thread to sequentially read the content of the file information list, determines the corresponding file difference file according to the read content. If it is a difference file, the MD5 values of each fixed block of the file at the disaster recovery machine end are obtained in sequence according to a fixed block size and transmitted to the working machine. In step S4, the receiving thread of the working machine receives and saves the MD5 values of each fixed block of the same-name file sent by the disaster recovery machine. Meanwhile, it starts a data transmission thread to sequentially read the received MD5 values of each fixed block of the same-name file, compares them with the MD5 values of the corresponding block contents of the corresponding file on the working machine, and transmits the contents of the blocks with different comparison results to the disaster recovery machine. In step S5, the receiving thread of the disaster recovery machine receives the data transmitted by the working machine and writes it to overwrite the same block of the corresponding file.

[0010] Preferably, in step S1, the obtained file information at least includes the file path, name, and file attributes, and the file attributes include the file size and modification time.

[0011] Preferably, step S3 further includes: In step S300, while writing the file information into the file information list, start a read thread to sequentially read the content of the file information list. In step S301, for each file information read, obtain the same-name file of the file under the corresponding path of the disaster recovery machine, and determine whether there is a same-name file. If there is no same-name file, directly send the disaster recovery machine file name and file size to the working machine. If there is a same-name file, go to step S302. In step S302, compare whether the modification time and size of the same-name file in the disaster recovery machine are the same as the modification time and size in the read file information. If they are the same, return the result that the files are the same to the working machine. If they are different, obtain the MD5 values of each fixed block of the same-name file in the disaster recovery machine in the order of the fixed block length and transmit them to the working machine.

[0012] Preferably, in step S301, if there is no same-name file, directly return the result that the file size at the standby end is 0 to the working machine.

[0013] Preferably, in step S302, if the modification time and size of the same-name file in the disaster recovery machine are different from the modification time and size in the read file information, further sequentially read the content of the same-name file in the disaster recovery machine in the order of the fixed block length, calculate the MD5 values of each fixed block, and transmit the calculated MD5 values of each data block to the working machine.

[0014] Preferably, in step S4, after the working machine receiving thread receives the MD5 value of the file with the same name on the disaster recovery machine, it saves the MD5 values to the MD5 list in sequence. Meanwhile, the working machine transfer data thread is started to sequentially read the MD5 list, and the MD5 value of each MD5 value is compared with the MD5 value of the file content of the corresponding file on the working machine with the same displacement size. If the MD5 values are different, the data block of this segment on the working machine is transferred to the disaster recovery machine.

[0015] Preferably, in step S5, after the working machine transfers data to the disaster recovery machine, the disaster recovery machine first determines whether the corresponding file is open. If it is not open, the file is opened first; if it is determined that the file is open, the received data block is directly written at the current offset position; when a file block is written, the working machine sends an end message, and the disaster recovery machine closes the written file and then processes the next differential file.

[0016] Preferably, in step S5, if the directory where the file is located does not exist, the parent directories where the file is located are created in sequence, and then the file is opened in read-write mode.

[0017] Preferably, when the working machine traversal thread finishes traversing, it sends a traversal end signal to the disaster recovery machine. If the read thread has finished reading the file information list at this time, the read task ends; otherwise, the read thread continues to execute.

[0018] To achieve the above object, the present invention further provides a fast backup system for differential mirroring during file traversal, including: A working machine, which is used to start a traversal thread to scan all file lists under the original directory of the working machine, sequentially obtain the file information of each file to be transferred and send it to the disaster recovery machine; receive and save the MD5 values of each fixed block of the file with the same name sent by the disaster recovery machine, sequentially read the MD5 values of each fixed block of the received file with the same name, and compare them with the MD5 values of the corresponding file content of the corresponding file on the working machine, and transfer the content of the blocks with different comparison results to the disaster recovery machine; A disaster recovery machine, which is used to use the receiving thread to save the file information of each file received from the working machine end to the file information list on the disaster recovery machine; while the receiving thread writes the file information into the file information list, at the same time, a read thread is started to sequentially read the content of the file information list, judge the corresponding file differential file according to the read content, if it is a differential file, sequentially obtain the MD5 values of each fixed block of the file on the disaster recovery machine end according to the fixed block size, and transfer them to the working machine; when the receiving thread receives the data block data transferred by the working machine, it writes it into the same block of the corresponding file to overwrite.

[0019] Compared with the prior art, the fast backup method and system for differential mirroring during file traversal of the present invention realizes fast data backup of differential mirroring while traversing, saves the time used to traverse the file directory list on the working machine end and reduces the time for transferring data. Description of the Drawings

[0020] Figure 1 It is a flowchart of the steps of a fast backup method for differential mirroring during file traversal according to the present invention; Figure 2 It is a system structure diagram of a fast backup system for differential mirroring during file traversal according to the present invention; Figure 3 It is a flowchart of traversing data and transmitting file attribute information in an embodiment of the present invention; Figure 4 It is a flowchart of comparing file contents in an embodiment of the present invention. Detailed Embodiments

[0021] The following describes the embodiments of the present invention through specific specific examples in combination with the drawings. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific examples, and various details in this specification can also be modified and changed based on different viewpoints and applications without departing from the spirit of the present invention.

[0022] Figure 1 It is a flowchart of the steps of a fast backup method for differential mirroring during file traversal according to the present invention. As Figure 1 shown, a fast backup method for differential mirroring during file traversal according to the present invention includes the following steps: Step S1, the traversal thread of the working machine scans all file lists in the original directory of the working machine, sequentially obtains the file information of each file to be transmitted, and sequentially sends it to the disaster recovery machine.

[0023] Specifically, before starting to back up data, the traversal thread of the working machine first scans all file lists in the original directory of the working machine, sequentially obtains the file information of each file to be transmitted, and the file information at least includes file path, name, and file attributes, etc., and sends these file information to the disaster recovery machine in sequence.

[0024] For example, for the original directory A of the working machine, the traversal thread traverses all files in directory A, sequentially obtains the file name, path, and attribute information of each file, and sends them to the disaster recovery machine.

[0025] Step S2, the receiving thread of the disaster recovery machine saves the file information of each file received from the working machine end into the file information list on the disaster recovery machine.

[0026] When receiving the file information sent from the working machine end, the receiving thread of the disaster recovery machine saves the received file name, path, and attribute information into the namelist file of the file information list on the disaster recovery machine.

[0027] Step S3, while writing the file information into the file information list, the disaster recovery machine simultaneously starts a read thread to sequentially read the content of the file information list, determines the corresponding file difference file according to the read content. If it is a difference file, the MD5 values of each fixed block of the file at the disaster recovery machine end are sequentially obtained according to a fixed block size and transmitted to the working machine.

[0028] Specifically, step S3 further includes: Step S300, while writing the file information into the file information list, start a read thread to sequentially read the content of the file information list.

[0029] That is to say, while writing the file information into the file information list, the disaster recovery machine starts another read thread to continuously read the file information in the namelist file of the file information list. The read content includes the file name, path, file size, time and other attribute information of the working end until the working machine finishes sending all.

[0030] Step S301, for each file information read, obtain the file with the same name under the corresponding path of the disaster recovery machine for this file, and determine whether there is a file with the same name. If there is no file with the same name, directly return the corresponding result to the working machine. If there is a file with the same name, enter step S302.

[0031] Specifically, for each record read from the file information list namelist, it is the file information corresponding to a file at the working machine end. Obtain the file with the same name under the corresponding path of the disaster recovery machine for this file, and determine whether the file with the same name exists. If the file with the same name does not exist, directly return the result that the backup end file size is 0 to the working machine. At this time, it is prompted that the working machine needs to perform a full backup.

[0032] Step S302, compare whether the modification time and size of the file with the same name at the disaster recovery machine are the same as the modification time and size in the read file information. If they are the same, return the result that the files are the same to the working machine; if they are different, obtain the MD5 values of each fixed block of the file with the same name in the disaster recovery machine in the order of the fixed block length and transmit them to the working machine.

[0033] That is to say, if there is a file with the same name on the disaster recovery machine locally, further compare the file size and modification time. If they are the same, it is considered that the file is consistent, and a result indicating that the files are the same is returned to the working machine, indicating that the file has not changed and there is no need to synchronize, so this file can be ignored. If the file size or modification time is different, then sequentially read the content of the file with the same name on the disaster recovery machine in the order of the fixed block length blk_size (for example, blk_size is 64K), calculate the MD5 value of each fixed block, and transmit the calculated MD5 values of each fixed block to the working machine. Among them, before the disaster recovery machine transmits the md5, it needs to first send the file name information to the working machine, then sequentially read the content of a fixed size (64k) from this file and calculate the md5, and sequentially send the md5 to the working machine and save it to the md5 list of the working machine.

[0034] Step S4: The receiving thread of the working machine receives and saves the MD5 values of each fixed block of the file with the same name sent by the disaster recovery machine. At the same time, start the data transmission thread to sequentially read the MD5 values of each fixed block of the file with the same name received, and compare them with the MD5 values of the corresponding file content of the corresponding blocks on the working machine, and transmit the content of the blocks with different comparison results to the disaster recovery machine.

[0035] Specifically, after the receiving thread of the working machine receives the MD5 values of the file with the same name from the disaster recovery machine, it saves them to the md5 list in order. At the same time, the data transmission thread of the working machine sequentially reads this md5 list, and compares each md5 value with the md5 value of the file content of the same displacement blk_size on the corresponding file on the working machine. If the md5 values are different, then transmit the data block of this fragment of the file to the disaster recovery machine.

[0036] Preferably, if the receiving thread of the working machine receives the result that the size of the standby file is 0, it means that a full backup is required at this time, and then transmit the entire file content corresponding to the file name on the working machine side to the disaster recovery machine; if the receiving thread of the working machine receives the result that the files are the same, ignore this file and do not perform backup processing.

[0037] Step S5: The receiving thread of the disaster recovery machine receives the data transmitted by the working machine and writes it to the same block of the corresponding file to overwrite.

[0038] Specifically, after the working machine transmits data to the disaster recovery machine, the disaster recovery machine first determines whether the file is open. If it is not open, it first opens the file. If the directory where the file is located does not exist, it sequentially creates the parent directory where the file is located, and then opens the file in read-write mode. If it is determined that the file is open, directly write the received data block at the current offset position. When a file block is written, the working machine will send an end message, and the disaster recovery machine closes the written file and then processes the next differential file.

[0039] Preferably, when the traversal thread of the working machine finishes traversing, it sends a signal indicating the end of traversal to the disaster recovery machine. If the read thread has finished reading the file information list namelist at this time, the read task ends; otherwise, the read thread continues to execute.

[0040] Figure 2 This is the system structure diagram of a fast backup system for differential mirroring during file traversal according to the present invention. As Figure 2 shown, a fast backup system for differential mirroring during file traversal according to the present invention includes: A working machine 20, which is used to start a traversal thread to scan all file lists in the original directory of the working machine, sequentially obtain the file information of each file to be transmitted, and send it to the disaster recovery machine; receive and save the MD5 values of each fixed block of the same-name file sent by the disaster recovery machine, sequentially read the MD5 values of each fixed block of the received same-name file, and compare them with the MD5 values of the corresponding file contents of the corresponding blocks on the working machine, and transmit the contents of the blocks with different comparison results to the disaster recovery machine.

[0041] Specifically, the working machine 20 further includes: A traversal thread start unit 201, which is used to start a traversal thread to scan all file lists in the original directory of the working machine, sequentially obtain the file information of each file to be transmitted, and sequentially send it to the disaster recovery machine.

[0042] Specifically, before starting to back up data, the traversal thread start unit 201 of the working machine starts a traversal thread to first scan all file lists in the original directory of the working machine, sequentially obtain the file information of each file to be transmitted, where the file information at least includes file path, name, and file attributes, etc., and sequentially send this file information to the disaster recovery machine.

[0043] For example, for the original directory A of the working machine, the traversal thread traverses all files in directory A, sequentially obtains the file name, path, and attribute information of each file, and sends it to the disaster recovery machine.

[0044] A working machine receiving thread start unit 202, which is used to start a receiving thread to receive and save the MD5 values of each fixed block of the same-name file sent by the disaster recovery machine.

[0045] Specifically, when receiving the MD5 values of each data block of the same-name file sent by the disaster recovery machine, the receiving thread start unit 202 starts a receiving thread to receive the MD5 values of the same-name file of the disaster recovery machine, and saves them to the md5 list in order.

[0046] A data transmission thread start unit 203, which is used to start a data transmission thread to sequentially read the MD5 values of each fixed block of the received same-name file, compare them with the MD5 values of the corresponding file contents of the corresponding blocks on the working machine, and transmit the contents of the blocks with different comparison results to the disaster recovery machine.

[0047] While the receiving thread receives and saves the MD5 value of the file with the same name on the disaster recovery machine to the MD5 list, the transmission data thread startup unit 203 starts the transmission data thread to sequentially read the MD5 list, and compares each MD5 value with the MD5 value of the file content with the same displacement blk_size of the corresponding file on the working machine. If the MD5 values are different, the data block of this fragment of the file is transmitted to the disaster recovery machine.

[0048] Preferably, if the working machine receiving thread receives the result that the size of the standby file is 0, it means that a full backup is required at this time, and the transmission data thread transmits the entire file content corresponding to the file name on the working machine side to the disaster recovery machine; if the working machine receiving thread receives the result that the files are the same, the file is ignored and no backup processing is performed.

[0049] The disaster recovery machine 21 uses the receiving thread to save the file information of each file received from the working machine side to the file information list on the disaster recovery machine; while the receiving thread writes the file information into the file information list, the read thread is started at the same time to sequentially read the content of the file information list, and judge the corresponding file difference file according to the read content. If it is a difference file, the MD5 value of each fixed block of the file on the disaster recovery machine side is obtained in sequence according to the fixed block size and transmitted to the working machine; when the receiving thread receives the data block data transmitted by the working machine, it writes it into the same block of the corresponding file to overwrite.

[0050] Specifically, the disaster recovery machine 21 further includes: The disaster recovery machine receiving thread startup unit 210 is used to start the receiving thread, and use the receiving thread to save the file information of each file received from the working machine side to the file information list on the disaster recovery machine; when the receiving thread receives the data block data transmitted by the working machine, it writes it into the same block of the corresponding file to overwrite.

[0051] Specifically, when receiving the file information sent from the working machine side, the disaster recovery machine receiving thread startup unit 201 starts the receiving thread to save the received file name, path and attribute information to the namelist file of the file information list on the disaster recovery machine.

[0052] After the working machine transmits the data block data of the corresponding file to the disaster recovery machine, the receiving thread first judges whether the file is open. If it is not open, the file is first opened. If the directory where the file is located does not exist, the parent directory where the file is located is created in sequence, and then the file is opened in read-write mode; if it is judged that the file is open, the received data block is directly written at the current offset position; when a file block is written, the working machine will send an end message, and the disaster recovery machine side closes the written file and then processes the next difference file.

[0053] A read thread startup unit 211, which is used to start a read thread to sequentially read the content of the file information list while the receiving thread writes file information into the file information list, determine the corresponding file difference file according to the read content, and if it is a difference file, sequentially obtain the MD5 values of each fixed block of the file at the disaster recovery machine end according to a fixed block size, and transmit them to the working machine.

[0054] Specifically, the read thread startup unit 211 includes: A read thread startup module, which is used to start a read thread to sequentially read the content of the file information list while writing file information into the file information list.

[0055] That is to say, while writing file information into the file information list, the disaster recovery machine starts another read thread to continuously read the file information in the namelist file of the file information list. The read content includes file name, path, file size, time and other attribute information of the working end until the working machine finishes sending all.

[0056] A same-name file judgment module, which is used to obtain the same-name file of the file at the corresponding path of the disaster recovery machine every time it reads the file information of a file, judge whether there is a same-name file. If there is no same-name file, directly return the corresponding result to the working machine. If there is a same-name file, enter the same-name file processing module.

[0057] Specifically, every time a record is read from the file information list namelist, it is the file information corresponding to a file at the working machine end. Obtain the same-name file of the file at the corresponding path of the disaster recovery machine, and judge whether the same-name file exists. If the same-name file does not exist, directly return the result that the backup end file size is 0 to the working machine. At this time, it is prompted that the working machine needs to perform a full backup.

[0058] A difference judgment and processing module, which is used to compare whether the modification time and size of the same-name file at the disaster recovery machine are the same as the modification time and size in the read file information. If they are the same, return the result that the files are the same to the working machine; if they are different, obtain the MD5 values of each fixed block of the same-name file in the disaster recovery machine in the order of the fixed block length, and transmit them to the working machine.

[0059] That is to say, if there is a same-name file locally at the disaster recovery machine, further compare the file size and modification time. If they are the same, it is considered that the file is consistent, and return the result that the files are the same to the working machine, indicating that the file has not changed and there is no need to synchronize, and this file can be ignored; if the file size or modification time is different, further sequentially read the content of the same-name file in the disaster recovery machine in the order of the fixed block length blk_size (for example, blk_size is 64K), calculate the MD5 values of each fixed block, and transmit the calculated MD5 values of each fixed block to the working machine.

[0060] Preferably, when the traversal thread of the working machine finishes traversing, it sends a traversal end signal to the disaster recovery machine. If the read thread has finished reading the file information list namelist at this time, the read task ends; otherwise, the read thread continues to execute. Embodiment

[0061] Figure 3 The flowchart of traversing data and transmitting file attribute information in the embodiment of the present invention.

[0062] As Figure 3 As shown, the traversal thread wk_thd1 of the working machine traverses directory A on the working machine, sequentially obtains the file name, path, and attribute information of each file, and sends the file name, path, and attribute information to the disaster recovery machine.

[0063] The receiving thread bk_thd1 of the disaster recovery machine saves the received file name, path, and attribute information of the working machine in the file information list namelist file of the disaster recovery machine.

[0064] The disaster recovery machine simultaneously starts another read thread bk_thd2 to continuously read the file information in the file information list namelist until the working machine has finished sending all the information.

[0065] For each file information read from the namelist by the disaster recovery machine, it compares the file information with the corresponding information of the local file of the disaster recovery machine. If there is a file with the same name in the local disaster recovery machine, it further compares the file size and modification time. If they are the same, it is considered that the file with the same name is consistent with the file of the read file information, and the file is ignored; if the file size or modification time is different, it further reads the content of each data block of the file with the same name in fixed blocks in sequence, calculates the md5 information of each block, and sends it to the working machine. The receiving thread wk_thd2 of the working machine receives the md5 message and saves it to the MD5 list in order. At the same time, the data transmission thread wk_thd3 of the working machine compares the md5 value of the corresponding data block of the corresponding file according to the MD value in the MD5 list. If they are different, it sends the content of the corresponding block to the disaster recovery machine, and the receiving thread bk_thd1 of the disaster recovery machine receives the data and writes it to the same block of the corresponding file to overwrite.

[0066] Figure 4This is the comparison flowchart of file content in the embodiment of the present invention. Taking the existence of the same-name file 1.1 at the disaster recovery machine end in the figure as an example, when the disaster recovery machine end determines that the size or modification time of file 1.1 is inconsistent with that of the working machine file 1.1, it sequentially reads the md5 value of each block of the file at the disaster recovery machine end. Assuming that the fixed size of each block is 64k, it is sent to the working machine. The total length of the data blocks read by the disaster recovery machine for file 1.1 does not exceed the size of the working machine file, because ultimately the size of the file on the disaster recovery machine after backup must be the same as that of the working machine. For example Figure 4 , fsize refers to the size of the working machine file 1.1. If fsize cannot be divided evenly by 64k, the length of the data block used for calculating md5 in the last segment is fsize / 64k. After receiving the md5 string, the working machine saves it in the md5 list sumList at the working machine end. The working machine sequentially reads the data blocks of the corresponding 64k size of the local file, calculates the md5 value of this block, and compares it with the md5 value of the corresponding offset at the disaster recovery machine end in the sumList list. If they are different, it sends the content of the corresponding data block to the disaster recovery machine and writes it to the data block at the same position of the corresponding file at the disaster recovery machine.

[0067] The present invention has the following advantages 1. In the present invention, regardless of which backup method is used, each mirroring only needs to transfer the files that have changed since the last backup to the same target path on the disaster recovery machine. After the transfer is completed, it is determined on the disaster recovery machine whether to perform full storage, incremental storage, or differential storage according to different backup methods. Moreover, the mirroring does not transfer the entire changed file, but the data segments with differences in the changed file. The minimum difference unit is 64k bytes, that is, a differential file only needs to transfer 64k bytes to the disaster recovery machine at least, thus greatly saving the time for transferring data.

[0068] 2. In the present invention, while the working machine transfers the file name, path, and attributes to the disaster recovery machine, it concurrently compares the md5 of the disaster recovery machine file that has been received and transfers the differential file, saving the waiting time for transferring files. As shown in the attached Figure 3 Steps 1 and 2 shown are executed concurrently, saving the time of step 1. The usual method is to start sending the disaster recovery machine md5 to the working machine only after the disaster recovery machine end receives the attribute list information of all files at the working end. In the case of a large number of small files and a complex path structure, traversing and transferring all file directory information of the working machine to the disaster recovery machine takes a long time, so that the actual data transfer needs to wait until all file attributes are traversed and transferred before it can start.

[0069] The following Table 1 shows the comparison results of the remote backup time test for a large number of small files. The time taken for mirroring data transfer is 2 hours and 11 minutes. After using the method of the present invention, the traversal time of 29 minutes is saved, and the total duration of data backup is reduced by nearly 30 minutes compared with traversing and then backing up.

[0070] Table 1

[0071] The above embodiments are only illustrative of the principles and effects of the present invention and are not intended to limit the present invention. Any person skilled in the art can modify and change the above embodiments without departing from the spirit and scope of the present invention. Therefore, the scope of the protection of the present invention shall be as set forth in the claims.

Claims

1. A method for quickly backing up differential images during file traversal, comprising the following steps: Step S1, the traversal thread of the working machine scans all file lists in the original directory of the working machine, obtains the file information of each file to be transferred in turn, and sends it to the disaster recovery machine; Step S2, the disaster recovery machine receiving thread saves the received file information of each file to be transferred into the file information list on the disaster recovery machine; Step S3, while writing the file information into the file information list, the disaster recovery machine simultaneously starts a reading thread to sequentially read the contents of the file information list, and determines the corresponding file difference file according to the read contents. If it is a difference file, the MD5 value of each fixed block of the file on the disaster recovery machine is obtained in sequence according to the fixed block size, and transmitted to the working machine; Step S4, the receiving thread of the working machine receives and saves the MD5 value of each fixed block of the file with the same name sent by the disaster recovery machine, and at the same time starts the data transmission thread to sequentially read the MD5 value of each fixed block of the file with the same name received, and compares it with the MD5 value of the file content of the corresponding block of the corresponding file on the working machine, and transmits the content of the block with different comparison results to the disaster recovery machine; Step S5: The receiving thread of the disaster recovery machine receives the transmission data from the working machine and writes it into the same block of the corresponding file to overwrite it.

2. A method for fast backup of differential images during file traversal as claimed in claim 1, characterized in that: In step S1, the acquired file information at least includes the file path, name and file attributes, and the file attributes include the file size and modification time.

3. The method for quickly backing up differential images during file traversal as claimed in claim 1, characterized in that: Step S3 further comprises: Step S300, while writing the file information into the file information list, starting a reading thread to sequentially read the contents of the file information list; Step S301, each time the file information of a file is read, obtain the file with the same name in the corresponding path of the disaster recovery machine, and determine whether there is a file with the same name. If there is no file with the same name, directly return the corresponding result to the working machine. If there is a file with the same name, go to step S302; Step S302, compare the modification time and size of the file with the same name on the disaster recovery machine with the modification time and size in the read file information to see if they are the same. If they are the same, return the result that the files are the same to the working machine; if they are different, obtain the MD5 value of each fixed block of the file with the same name in the disaster recovery machine in a fixed block length sequence and transmit it to the working machine.

4. A method for fast backup of differential images during file traversal as claimed in claim 3, characterized in that: In step S301, if a file with the same name does not exist, a result indicating that the backup file size is 0 is directly returned to the working machine.

5. The method for quickly backing up differential images during file traversal as claimed in claim 3, characterized in that: In step S302, if the modification time and size of the file with the same name on the disaster recovery machine are different from the modification time and size in the read file information, the content of the file with the same name in the disaster recovery machine is further read in sequence in fixed block length order, the MD5 value of each fixed block is calculated, and the MD5 value of each data block obtained by calculation is transmitted to the working machine.

6. A method for fast backup of differential images during file traversal as claimed in claim 5, characterized in that: In step S4, after the working machine receiving thread receives the MD5 value of the file with the same name on the disaster recovery machine, it saves it in the MD5 list in order, and starts the working machine data transmission thread to read the MD5 list in sequence, and compares each MD5 value with the MD5 value of the file content with the same displacement size of the corresponding file on the working machine. If the MD5 values ​​are different, the data block of the fragment on the working machine is transferred to the disaster recovery machine.

7. A method for fast backup of differential images during file traversal as claimed in claim 6, characterized in that: In step S5, after the working machine transmits data to the disaster recovery machine, the disaster recovery machine first determines whether the corresponding file is open. If not, the file is opened first; if the file is opened, the received data block is directly written at the current offset position; When a file block is written, the working machine sends a completion message, the disaster recovery machine closes the written file, and then processes the next differential file.

8. A method for fast backup of differential images during file traversal as claimed in claim 7, characterized in that: In step S5, if the directory where the file is located does not exist, the parent directories where the file is located are created in sequence, and then the file is opened in a read-write mode.

9. The method for fast backup of differential images during file traversal as claimed in claim 7, characterized in that: When the traversal thread of the working machine completes the traversal, a traversal completion signal is sent to the disaster recovery machine. If the reading thread has finished reading the file information list at this time, the reading task is terminated, otherwise the reading thread continues to execute.

10. A fast backup system for differential images during file traversal, comprising: The working machine is used to start the traversal thread to scan all the file lists in the original directory of the working machine, and obtain the file information of each file to be transferred in turn and send it to the disaster recovery machine; Receive and save the MD5 values ​​of each fixed block of the file with the same name sent by the disaster recovery machine, read the MD5 values ​​of each fixed block of the file with the same name received in sequence, and compare them with the MD5 values ​​of the file contents of the corresponding blocks of the corresponding files on the working machine, and transmit the contents of the blocks with different comparison results to the disaster recovery machine; The disaster recovery machine is used to save the file information of each file received from the working machine into a file information list on the disaster recovery machine by using a receiving thread; While the receiving thread writes the file information into the file information list, the reading thread is started at the same time to read the contents of the file information list in sequence, and the corresponding file difference file is determined based on the read contents. If it is a difference file, the MD5 value of each fixed block of the file on the disaster recovery machine is obtained in sequence according to the fixed block size, and transmitted to the working machine; when the receiving thread receives the data block data transmitted by the working machine, it is written into the same block of the corresponding file and overwritten.

Citation Information

Cited By

  • Data synchronization method and device, electronic equipment, storage medium and program product

    CN120744010A