File backup method and device, computer equipment and readable storage medium
By generating parallel running backup tasks and timed backup set checkpoint files, the problem of backup time growth and repeated backup after interruption is solved, and the efficiency and fault tolerance of file backup are improved.
Patent Information
- Application Number
- CN202411390001.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-08
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2044-10-08
AI Technical Summary
As the amount of backup data continues to grow, the time required to complete backup is also increasing, resulting in an increase in the probability of backup failure. The completed backup data can only be deleted when the backup is interrupted, wasting backup time and resources, resulting in low file backup efficiency.
Provides a file backup method, by obtaining the directory to be backed up, generating multiple backup tasks, running these tasks in parallel, and regularly generating and uploading backup set checkpoint files to restore backup progress when backup is interrupted.
Improve the efficiency of file backup, reduce the operation of repeated backup after backup interruption, ensure the rapid recovery of backup tasks, avoid data loss, and improve the fault tolerance and resource utilization of file backup.
Smart Images

Figure CN120086060A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of computer data storage, and in particular to a file backup method, device, computer device, computer-readable storage medium, and computer program product. Background Art
[0002] Data backup is a key measure to ensure data security and business continuity. In the face of unpredictable risks such as hardware failures, software errors, virus attacks, human operation errors, or natural disasters, backup can prevent data loss and ensure the integrity and availability of critical information. By performing regular backups, enterprises can quickly resume normal operations, reducing potential financial losses and reputational damage.
[0003] Files are the basic form of unstructured data storage, and file backup is an important sub-domain in the field of data backup. With the development of information technology, the total amount of unstructured data has also shown an explosive growth trend. At the same time, with the wide application of many new types of storage such as distributed storage and object storage, the efficiency and performance of file backup, especially the backup of a large number of files and a large number of small files, have become the main challenges faced by the file backup function.
[0004] As the amount of backup data continues to grow, the time required to complete the backup also continues to increase. The longer the backup time, the greater the probability of backup failure due to abnormalities such as network, hardware, software, or human misoperation during the backup process. Once the backup process is abnormally aborted, the backup data that has already been completed often has to be deleted as garbage data and wait for the next re-backup, which wastes a large amount of backup time and backup resources, resulting in low file backup efficiency. Summary of the Invention
[0005] Based on this, in view of the above technical problems, it is necessary to provide a file backup method, device, computer device, computer-readable storage medium, and computer program product that can improve the efficiency of file backup.
[0006] In a first aspect, this application provides a file backup method, the method including:
[0007] Obtain a directory to be backed up; the directory to be backed up includes files sorted in the order of file identifiers and files under subdirectories sorted in the order of directory identifiers;
[0008] For each file in the directory to be backed up, generate multiple backup tasks; the backup tasks are used to back up each file to the server in the order of the file identifiers and the directory identifiers, and generate backup set shards corresponding to the backup tasks; the backup set shards are used to write the data of the file; the backup tasks are associated with context information; the context information includes the backup progress information and backup set shard information of the backup tasks; the backup set shard information includes the amount of data of the file that has been written in the backup set shard.
[0009] Run the multiple backup tasks in parallel, and regularly generate corresponding backup set checkpoint files at the checkpoint time according to the context information of the backup tasks, and upload the backup set checkpoint files to the server; the backup set checkpoint files are used to restore the backup progress of the backup tasks in case the backup tasks are interrupted.
[0010] In one embodiment, the method further includes:
[0011] The method further includes:
[0012] In case the backup task runs and is interrupted, obtain the backup set checkpoint file corresponding to the backup task from the server;
[0013] According to the backup set checkpoint file, determine the checkpoint file identifier and checkpoint directory identifier recorded at the latest checkpoint time; the checkpoint file identifier is the file identifier corresponding to the last backed-up file recorded at the latest checkpoint time; the checkpoint directory identifier is the directory identifier corresponding to the directory being backed up at the latest checkpoint time;
[0014] Traverse the directories where each file in the directory to be backed up is located, compare the directory identifiers corresponding to each file in the directory to be backed up with the checkpoint directory identifier, skip the backed-up directories, and locate the target directory that matches the checkpoint directory identifier;
[0015] Traverse each file in the target directory, compare the file identifiers of each file in the target directory with the checkpoint file identifier, skip the backed-up files, locate the backup interruption position in the directory to be backed up, and resume running the backup task at the backup interruption position; the backup interruption position is the file position that matches both the checkpoint directory identifier and the checkpoint file identifier.
[0016] In one embodiment, the comparing the file identifiers of each file in the target directory with the checkpoint file identifier, skipping the backed-up files, and locating the backup interruption position in the directory to be backed up includes:
[0017] Compare the file identifiers of each file in the target directory with the checkpoint file identifier, determine the files with file identifiers less than or equal to the checkpoint file identifier as the backed-up files, skip the backed-up files, and locate the backup interruption position in the directory to be backed up.
[0018] In one embodiment, traversing the directories where each file in the directory to be backed up is located, comparing the directory identifiers corresponding to each file in the directory to be backed up with the checkpoint directory identifier, skipping the backed-up directories, and locating the target directory that matches the checkpoint directory identifier includes:
[0019] Traverse the directories where each file in the directory to be backed up is located;
[0020] If the target identifier of the currently traversed directory is the parent directory identifier of the checkpoint directory identifier, determine the files in the directory corresponding to the parent directory identifier as the backed-up files, skip the directory corresponding to the parent directory identifier, and determine the target directory that matches the checkpoint directory identifier from the subdirectories included in the directory corresponding to the parent directory identifier.
[0021] In one embodiment, the method further includes:
[0022] In the case where the backup task runs and is interrupted, obtain the backup set checkpoint file corresponding to the backup task from the server;
[0023] According to the backup set checkpoint file, determine the backup set shard information recorded at the latest checkpoint time, discard the target data in the backup set shards corresponding to the backup task, and obtain the backup set shards after truncation operation; the target data is the data written to the backup set shards after the checkpoint time;
[0024] According to the backup set checkpoint file, restore the context information of the backup task to the context information corresponding to the latest checkpoint time;
[0025] According to the context information corresponding to the latest checkpoint time, locate the backup interruption position in the directory to be backed up, and resume running the backup task at the backup interruption position to write the unbacked-up files to the backup set shards after truncation operation.
[0026] In one embodiment, the method further includes:
[0027] Obtain the backup set checkpoint file corresponding to the backup task from the server;
[0028] Discard the target data in the backup set shard corresponding to the backup task according to the backup set shard information corresponding to the checkpoint time recorded in the backup set checkpoint file, so as to obtain the backup set shard after the truncation operation; the target data is the data written into the backup set shard after the checkpoint time.
[0029] Write format data matching the target storage format into the backup set shard to generate a valid storage format file corresponding to the backup set shard; the target storage format is the storage format of the backup set composed of the backup set shards generated according to each backup task; the valid storage format file is used to obtain the files that have been backed up to the server at the checkpoint time in the case where the backup task is interrupted.
[0030] In a second aspect, the present application further provides a file backup device, including:
[0031] A file acquisition module, configured to acquire a directory to be backed up; the directory to be backed up includes files sorted in the order of file identifiers and files in subdirectories sorted in the order of directory identifiers.
[0032] A task generation module, configured to generate a plurality of backup tasks for each file in the directory to be backed up; the backup tasks are used to back up each file to the server in the order of the file identifiers and the directory identifiers, and generate a backup set shard corresponding to the backup task; the backup set shard is used to write the data of the file; the backup task is associated with context information; the context information includes the backup progress information and backup set shard information of the backup task; the backup set shard information includes the amount of data that has been written into the file in the backup set shard.
[0033] An operation module, configured to run the plurality of backup tasks in parallel, regularly generate a corresponding backup set checkpoint file at the checkpoint time according to the context information of the backup task, and upload the backup set checkpoint file to the server; the backup set checkpoint file is used to restore the backup progress of the backup task in the case where the backup task is interrupted.
[0034] In a third aspect, the present application further provides a computer device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the steps of the above method are implemented.
[0035] In a fourth aspect, the present application further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the above method are implemented.
[0036] In a fifth aspect, the present application also provides a computer program product, including a computer program which, when executed by a processor, implements the steps of the above method.
[0037] For the above file backup method, device, computer device, computer-readable storage medium, and computer program product, a directory to be backed up is obtained. The directory to be backed up includes files sorted in the order of file identifiers and files in subdirectories sorted in the order of directory identifiers; multiple backup tasks are generated according to each file in the directory to be backed up. The backup tasks are used to back up each file to a server in the order of file identifiers and directory identifiers, and backup set shards corresponding to the backup tasks are generated. The backup set shards are used for writing data of the files. The backup tasks are associated with context information, which includes backup progress information of the backup tasks and backup set shard information. The backup set shard information includes the amount of data of the files that has been written in the backup set shards; multiple backup tasks are run in parallel, and at regular intervals, corresponding backup set checkpoint files are generated according to the context information of the backup tasks at checkpoint times and uploaded to the server. The backup set checkpoint files are used to restore the backup progress of the backup tasks in case the backup tasks are interrupted. By sorting the files and subdirectories in the directory to be backed up and backing up each file in order, it helps to quickly restore the backup tasks subsequently; generating multiple backup tasks that run in parallel improves the efficiency of file backup. Moreover, associating the backup tasks with context information including backup progress information and backup set shard information can track the execution status of the backup tasks in real time, enabling subsequent accurate restoration of the backup tasks to the state before the backup interruption according to the context information. And according to the context information, backup set checkpoint files are generated at checkpoint times and uploaded to the server regularly. When the backup is interrupted, the backup progress can be restored relying on the backup set checkpoint files, avoiding the operation of repeated backup after the backup interruption, improving the restoration efficiency of the backup tasks when the file backup is interrupted, and thus improving the efficiency of file backup. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required for describing the embodiments of the present application or related technologies. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other related drawings can be obtained based on these drawings.
[0039] Figure 1 It is an application environment diagram of a file backup method in an embodiment;
[0040] Figure 2Schematic flowchart of a file backup method in an embodiment;
[0041] Figure 3 Logic diagram of a file backup method in an embodiment;
[0042] Figure 4 Logic diagram of another file backup method in an embodiment;
[0043] Figure 5 Logic diagram of a method for restoring backup progress from a backup checkpoint in an embodiment;
[0044] Figure 6 Logic diagram of a method for restoring backup progress in an embodiment;
[0045] Figure 7 Schematic flowchart of another file backup method in an embodiment;
[0046] Figure 8 Structural block diagram of a file backup device in an embodiment;
[0047] Figure 9 Internal structure diagram of a computer device in an embodiment. Detailed implementation manners
[0048] In order to make the objectives, technical solutions, and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0049] The file backup method provided by the embodiments of the present application can be applied to, for example Figure 1In the application environment shown. Among them, the terminal 102 communicates with the server 104 through the network. The data storage system can store the data that the server 104 needs to process. The data storage system can be integrated on the server 104, or can be placed on the cloud or other network servers. The terminal 102 obtains the directory to be backed up; the directory to be backed up includes the files sorted in the order of file identifiers and the files in the subdirectories sorted in the order of directory identifiers; the terminal 102 generates multiple backup tasks for each file in the directory to be backed up; the backup tasks are used to back up each file to the server in the order of file identifiers and directory identifiers, and generate backup set shards corresponding to the backup tasks; the backup set shards are used to write the data of the files; the backup tasks are associated with context information; the context information includes the backup progress information of the backup tasks and the backup set shard information; the backup set shard information includes the amount of data of the files that have been written in the backup set shards; the terminal 102 runs multiple backup tasks in parallel, and regularly generates corresponding backup set checkpoint files at the checkpoint time according to the context information of the backup tasks, and uploads the backup set checkpoint files to the server; the backup set checkpoint files are used to restore the backup progress of the backup tasks in case of interruption of the backup tasks. Among them, the terminal 102 can be, but is not limited to, various personal computers, laptop computers, smart phones, tablet computers, Internet of Things devices and portable wearable devices. The Internet of Things devices can be smart speakers, smart TVs, smart air conditioners, smart vehicle-mounted devices, projection devices, etc. The portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc. The head-mounted devices can be virtual reality (VR) devices, augmented reality (AR) devices, smart glasses, etc. The server 104 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services.
[0050] In an exemplary embodiment, as Figure 2 shown, a file backup method is provided. Taking the method applied to Figure 1 the terminal 102 in
[0051] Step S202, obtain the directory to be backed up.
[0052] Among them, the directory to be backed up may include files in the directory to be backed up and files in subdirectories in the directory to be backed up. Exemplarily, "c:\path\to\current" is a directory at the fourth level. Among them, "c:" is the first-level directory, "c:\path" is the second-level directory, "c:\path\to" is the third-level directory, and "c:\path\to\current" is the fourth-level directory. Assume that the directory to be backed up is "c:", then the files in the directory to be backed up may include files in the "c:" directory, as well as files in the subdirectories "c:\path", "c:\path\to", and "c:\path\to\current".
[0053] The terminal can generate a list of files to be backed up based on the files in these directories to be backed up. Each file in the list of files to be backed up corresponds to metadata, such as file size, modification time, permissions, etc. The terminal can use these files in the list of files to be backed up and the corresponding metadata as the input for one or more backup tasks.
[0054] Among them, the directory to be backed up includes files sorted in the order of file identifiers and files in subdirectories sorted in the order of directory identifiers.
[0055] In one embodiment, the order of file identifiers may be the alphabetical order of file names, and the order of target identifiers may be the alphabetical order of directory names. The terminal can sort the files in the alphabetical order of file names and sort the subdirectories in the alphabetical order of directory names. Among them, the alphabetical order may refer to the way in which a computer sorts strings in the order of a dictionary, which is carried out one by one based on the ASCII code values of the strings. Taking the file name (path) of the file as a string and sorting it through the alphabetical order and distributing it to each backup task, so that the file backup can be sorted in the alphabetical order of file names, which can greatly simplify the context information for backup task maintenance, and then improve the speed of generating backup set checkpoint files subsequently and reduce the storage space occupied by the backup set checkpoint files.
[0056] Step S204, generate multiple backup tasks for each file in the directory to be backed up.
[0057] Among them, the backup task is used to back up each file to the server in the order of file identifiers and directory identifiers and generate a backup set shard corresponding to the backup task; the backup set shard is used to write the data of the file; the backup task is associated with context information; the context information includes the backup progress information of the backup task and the backup set shard information; the backup set shard information includes the amount of data of the file that has been written in the backup set shard.
[0058] A backup task can generate a backup set for the files in the directory to be backed up and key information such as corresponding metadata in a specific format (such as tar, mtf, etc.). A backup set shard can be a data block (shard) in the backup set.
[0059] A backup set shard can be a file or a special file structure generated according to the storage format of the backup. Optionally, a backup set shard can be a compressed package in rar format or zip format. Therefore, backing up files to the server can be a process of writing files into the compressed package. Each backup task can generate one or more compressed packages, and these compressed packages are equivalent to backup set shards. The backup set shards generated by each backup task can jointly constitute a complete backup set, which can include the backup of all file data and corresponding metadata in the directory to be backed up.
[0060] In a specific implementation, for the files in the directory to be backed up, generating multiple backup tasks can be distributing the files in the directory to be backed up to multiple backup tasks running in parallel according to certain rules. Each backup task independently completes the backup of part of the files and generates corresponding backup set shards. The backup set shards generated by all backup tasks jointly constitute a complete backup set, which contains the backup of all file data and metadata in the directory to be backed up.
[0061] In a specific implementation, multiple backup tasks run concurrently. Each backup task maintains a context information for saving the backup progress information and the generated backup set shard information of the backup task. Whenever a complete file is backed up, the backup task can refresh the context information of the backup task. The context information is real-time information maintained for each backup task during the backup process, used to track the execution progress and status of the task.
[0062] In an example, the context information can include the ID of the backup task, the current directory being backed up, the last file that has been backed up, the size of the currently generated backup set shard, etc. The backup progress information can include the current directory being backed up, the last file that has been backed up, the percentage of the backup progress, etc.; the backup set shard information can include the size of the currently generated backup set shard, that is, the amount of data of the file that has been written into the backup set shard, such as 10 gigabytes (GB), 20GB, etc.
[0063] For the convenience of understanding by those skilled in the art, Figure 3 An exemplary logical diagram of a file backup method is provided. Figure 1 The terminal 102 of Figure 3 The backup host of Figure 1 The terminal 104 of Figure 3The backup storage server, where the backup host is the host where the backup files are located or the terminal that can access the complete files to be backed up. The backup storage server is a dedicated server that provides storage for backup operations. The backup host communicates with the backup storage server through the backup network to complete the transmission of backup data. The backup network includes, but is not limited to, network systems based on different transmission protocols such as the Internet Protocol (IP network), Storage Area Network (SAN network), etc. The backup host obtains the files in the directory to be backed up, that is, the files to be backed up, and distributes them to different backup tasks. Each backup task independently completes the backup of some files and generates corresponding backup set shards. The backup set shards generated by all backup tasks together constitute a complete backup set, and the backup set contains the backup of the data and metadata of all files to be backed up.
[0064] Step S206, run multiple backup tasks in parallel, and regularly generate corresponding backup set checkpoint files at the checkpoint time according to the context information of the backup tasks, and upload the backup set checkpoint files to the server.
[0065] Among them, the backup set checkpoint file is used to restore the backup progress of the backup task in case the backup task is interrupted.
[0066] In a specific implementation, during the process of running multiple backup tasks in parallel, the generation of the backup set checkpoint file can be triggered regularly. The generation of the backup set checkpoint file may include the following steps: lock the context information of each backup task, and generate the corresponding backup set checkpoint file according to the context information of each backup task, and upload it to the storage server for storage.
[0067] Among them, the checkpoint time is the time when the generation of the backup set checkpoint file is triggered regularly. In an example, the time for triggering the regular generation of the backup set checkpoint file can be a fixed time or one of the configuration options for the running of the backup task, and the format of the backup set checkpoint file is not limited.
[0068] In a specific implementation, the backup set checkpoint file is a file generated regularly based on the context information, which is used to be able to restore to the backup state at the checkpoint time after the backup is interrupted. The backup set checkpoint file may include the corresponding relationship between the backup progress information and the backup set shard information at each checkpoint time. Uploading the backup set checkpoint file to the server can obtain the latest backup set checkpoint file from the server to restore the backup progress of the backup task in case the backup task is interrupted, avoiding discarding the data that has been backed up before and restarting the backup of all files after the backup job is interrupted.
[0069] For the convenience of understanding by those skilled in the art, Figure 4A logic diagram of another file backup method is provided exemplarily. It can be seen that for the files in the list directory, that is, the files in the directory to be backed up, after sorting the file names, they are assigned to multiple backup tasks. Each backup task can generate a corresponding backup set shard, and each backup task maintains a context information, that is, the task context in the figure. Then, a backup checkpoint, that is, a backup checkpoint file, is generated regularly according to the task context.
[0070] The above file backup method, device, computer device, computer-readable storage medium and computer program product obtain a directory to be backed up, where the directory to be backed up includes files sorted in the order of file identifiers and files in subdirectories sorted in the order of directory identifiers; generate multiple backup tasks according to each file in the directory to be backed up. The backup task is used to back up each file to the server in the order of file identifiers and directory identifiers and generate a backup set shard corresponding to the backup task. The backup set shard is used to write the data of the file. The backup task is associated with context information, and the context information includes the backup progress information of the backup task and the backup set shard information. The backup set shard information includes the amount of data that has been written to the file in the backup set shard; run multiple backup tasks in parallel, regularly generate a corresponding backup set checkpoint file at the checkpoint time according to the context information of the backup task, and upload the backup set checkpoint file to the server. The backup set checkpoint file is used to restore the backup progress of the backup task in the case of interruption of the backup task. By sorting the files and subdirectories in the directory to be backed up and backing up each file in order, it helps to quickly restore the backup task subsequently; generating multiple backup tasks running in parallel improves the efficiency of file backup. Moreover, associating the backup task with the context information including the backup progress information and the backup set shard information can track the execution status of the backup task in real time, enabling the subsequent accurate restoration of the backup task to the state before the backup interruption according to the context information, and regularly generating a backup set checkpoint file at the checkpoint time according to the context information and uploading it to the server. When the backup is interrupted, the backup progress can be restored relying on the backup set checkpoint file, avoiding the operation of repeated backup after the backup interruption, improving the restoration efficiency of the backup task when the file backup is interrupted, and thus improving the efficiency of file backup.
[0071] In another embodiment, it further includes: in the case where the backup task runs and is interrupted, obtaining a backup set checkpoint file corresponding to the backup task from the server; determining, according to the backup set checkpoint file, the backup set shard information recorded at the latest checkpoint moment, discarding the target data in the backup set shards corresponding to the backup task to obtain the backup set shards after truncation operation; the target data is the data written to the backup set shards after the checkpoint moment; restoring the context information of the backup task to the context information corresponding to the latest checkpoint moment according to the backup set checkpoint file; positioning to the backup interruption position in the directory to be backed up according to the context information corresponding to the latest checkpoint moment, and resuming the backup task at the backup interruption position to write the unbacked-up files to the backup set shards after truncation operation.
[0072] In one example, the situation where the backup task runs and is interrupted may be caused by the following reasons: network failure, hardware failure, software operation, human operation error, insufficient storage space, power interruption, etc.
[0073] In specific implementation, after the backup task runs and is interrupted and the interruption reason is processed, each backup task can be restarted. At this time, the backup set checkpoint files corresponding to each backup task can be obtained from the server for restoring the backup progress. For example, the latest backup set checkpoint file generated by the backup task during the last run interruption, that is, the backup set checkpoint file generated at the latest checkpoint moment, can be obtained.
[0074] In specific implementation, the backup set checkpoint files of each backup task can record the backup set shard information corresponding to the checkpoint moment, that is, the size of the backup set shards. Through this backup set shard information, a truncation operation can be performed on the backup set shards of the backup task to discard the data written to the backup set shards after the checkpoint moment; the context information of each backup task can be restored through the backup set checkpoint files of each backup task, and the context information of each backup task is restored to the context information recorded at the checkpoint moment. The context information recorded in the backup set checkpoint files of each backup task may include the current directory to be backed up and the last file that has been completed in the current backup. According to these context information, the backup of the files in the directory to be backed up can be restored, so that the files that have not been backed up to the server before the checkpoint moment can continue to be backed up.
[0075] For the convenience of understanding by those skilled in the art, Figure 5 Exemplarily, a logic diagram of a method for restoring the backup progress from a backup checkpoint is provided. It can be seen that the context information and backup progress of the backup task can be restored through the backup set checkpoint file.
[0076] The technical solution of this embodiment obtains the backup set checkpoint file from the server, and truncates the data written to the backup set shard after the checkpoint time according to the backup set shard information recorded at the latest checkpoint time, avoiding re-backing up all files. Only some unfinished shards need to be discarded, minimizing the waste of time and storage resources caused by repeated backups. After the backup task is interrupted, the backup set checkpoint file can be used to quickly restore the context information of the backup task, enabling the backup progress to continue from the breakpoint, avoiding starting the backup from scratch, and improving the efficiency of restoring the backup task. By accurately locating and discarding the invalid data written after the checkpoint time, data consistency can be ensured, preventing data chaos, loss, duplication, and inconsistency caused by interruptions. By using the recovery mechanism of the backup set checkpoint and context information, it is ensured that all file data is completely backed up to the server, achieving efficient recovery of the backup task after interruption, improving the fault tolerance, recovery efficiency, and resource utilization rate of file backup, and thus improving the efficiency of file backup.
[0077] In another embodiment, in the case where the backup task runs and is interrupted, the backup set checkpoint file corresponding to the backup task is obtained from the server; according to the backup set checkpoint file, the checkpoint file identifier and the checkpoint directory identifier recorded at the latest checkpoint time are determined; the checkpoint file identifier is the file identifier corresponding to the last successfully backed up file recorded at the latest checkpoint time; the checkpoint directory identifier is the directory identifier corresponding to the directory being backed up at the latest checkpoint time; traverse the directories where each file in the directory to be backed up is located, compare the directory identifiers corresponding to each file in the directory to be backed up with the checkpoint directory identifier, skip the already backed up directories, and locate the target directory that matches the checkpoint directory identifier; traverse each file in the target directory, compare the file identifiers of each file in the target directory with the checkpoint file identifier, skip the already backed up files, locate the backup interruption position in the directory to be backed up, and resume running the backup task at the backup interruption position; the backup interruption position is the file position that matches both the checkpoint directory identifier and the checkpoint file identifier.
[0078] Among them, the context information recorded by the backup task at the checkpoint time may include the checkpoint file identifier and the checkpoint directory identifier corresponding to the checkpoint time, and these context information can be used to record the precise progress of the backup process.
[0079] Among them, the checkpoint file identifier may be the unique identifier of the last successfully backed up file recorded at the checkpoint time, such as the file path, file name, or metadata of the file. Through this checkpoint file identifier, it can be determined which file has been backed up before the backup is interrupted.
[0080] Among them, the checkpoint directory identifier can be a unique identifier indicating the directory that was being processed when the backup was interrupted. Usually, it is the directory path or the unique number of the directory. This identifier can be used to locate the level of the directory where the backup was interrupted.
[0081] In a specific implementation, traverse the directories where each file in the directory to be backed up is located. This can be to scan the root directory (i.e., the directory to be backed up) and all subdirectories, skip directories that do not match the checkpoint directory identifier, and search for directories that match the checkpoint directory identifier. This process can recursively go deep into multiple subdirectories until the target directory that matches the checkpoint directory identifier is found.
[0082] Among them, the target directory is the directory that was being processed when the backup was interrupted, and all recovery operations of the backup progress can continue within this directory.
[0083] After finding the target directory, compare the file identifiers of each file in the target directory with the checkpoint file identifier, skip files that do not match the checkpoint file identifier, and search for files that match the checkpoint file identifier. The checkpoint file identifier includes the file identifier of the last file that has been successfully backed up. Thus, the backup interruption position in the directory to be backed up can be found in the target directory, and the files after the backup interruption position are the files that have not been backed up. These unbacked-up files refer to the files that failed to be successfully written into the backup set shard at the checkpoint moment. For example, if the file with the checkpoint file identifier is the 10th file in the directory, then the backup will continue from the 11th file.
[0084] After determining the backup interruption position, the backup task can be restarted, and the data of these files can be written into the previous backup set shard. When resuming the operation of the backup task, it will not only ensure that the unbacked-up files are continuously written into the previous backup set shard, but also maintain the sequential backup of new files until the backup of the entire target directory is completed.
[0085] The technical solution of this embodiment can quickly determine and skip the backed-up directories and files by comparing the file identifier and the directory identifier, locate the backup interruption position, greatly reduce the repeated traversal and redundant backup operations, only need to continue from the exact interrupted position, do not need to re-backup the files or directories that have been processed, can quickly resume the progress of the backup task, and improve the efficiency of file backup.
[0086] In another embodiment, comparing the file identifiers of each file in the target directory with the checkpoint file identifier, skipping the files that have been backed up, and locating the backup interruption position in the directory to be backed up, including: comparing the file identifiers of each file in the target directory with the checkpoint file identifier, determining the files with file identifiers less than or equal to the checkpoint file identifier as the files that have been backed up, skipping the files that have been backed up, and locating the backup interruption position in the directory to be backed up.
[0087] Among them, each file has a unique file identifier (file ID). Optionally, the file identifier can be the file name, the path of the file, or a unique value generated through metadata (such as creation time, modification time, etc.). This file identifier is used to determine the order of the files.
[0088] In a specific implementation, the backup task will, according to the sorting result, start backing up files to the server one by one from the file with the smallest file identifier, and generate corresponding backup set shards during the backup process. This ordered backup method can ensure that each file is processed in a predetermined order, facilitating accurate positioning of the backup progress after an interruption occurs.
[0089] In a specific implementation, in the target directory, all files are sorted in ascending order according to the file identifier. The sorting of the file identifiers can follow the principle of the alphabetical order of the computer, that is, the order is determined by comparing the ASCII code values of the characters in the file name or identifier one by one. The smaller file identifier will be ranked in the front, and the larger identifier will be ranked in the back. This lexicographical or alphabetical sorting method can ensure that the files are processed in sequence, guaranteeing the orderliness and consistency of the backup.
[0090] After sorting the files in the target directory according to the file identifier, the file identifiers of each file in the target directory can be compared with the checkpoint file identifier. All files with file identifiers less than or equal to the checkpoint file identifier are files that have been backed up and can be skipped; all files with file identifiers greater than the checkpoint file identifier are files that have not been backed up, so that the files that need to start backing up this time can be found. For example, assuming that the checkpoint file identifier is "file_003", and the files in the target directory are "file_001", "file_002", "file_003", "file_004", "file_005" in sequence, then "file_004" and "file_005" will be determined as the files that have not been backed up, and the backup can start from "file_004".
[0091] The technical solution of this embodiment ensures the orderliness and efficient recovery of the backup task through the sorting of file identifiers and the comparison with the checkpoint file identifier, can quickly locate and recover the unfinished backup task, and improves the efficiency and reliability of file backup.
[0092] In another embodiment, traverse the directories where each file in the directory to be backed up is located, compare the directory identifier corresponding to each file in the directory to be backed up with the checkpoint directory identifier, skip the backed-up directories, and locate the target directory that matches the checkpoint directory identifier, including: traversing the directories where each file in the directory to be backed up is located; if the target identifier of the currently traversed directory is the parent directory identifier of the checkpoint directory identifier, determine that the files in the directory corresponding to the parent directory identifier are backed-up files, skip the directory corresponding to the parent directory identifier, and determine the target directory that matches the checkpoint directory identifier from the subdirectories included in the directory corresponding to the parent directory identifier.
[0093] Among them, if the directory corresponding to the checkpoint directory identifier is the target directory, the parent directory identifier of the checkpoint directory identifier may refer to the directory identifier of the parent directory of the target directory. For example, if the directory corresponding to the checkpoint target identifier is A and the parent directory of A is B, then the parent directory identifier of this checkpoint directory identifier refers to the directory identifier of B.
[0094] Exemplarily, "c:\path\to\current" is a directory at the fourth level. Among them, "c:" is the first-level directory, "c:\path" is the second-level directory, "c:\path\to" is the third-level directory, and "c:\path\to\current" is the fourth-level directory. Since file backup is performed in a hierarchical order, the files in the "c:\path" directory will be backed up before the files in the "c:\path\to" directory are processed. Assume that the target identifier of the directory being processed at the checkpoint time is "c:\path\to\current", that is, the checkpoint directory identifier is "c:\path\to\current", then the files under the parent directory identifier "c:\path\to" of this checkpoint directory identifier have all been backed up and can be skipped directly. It should be noted that "c:" can be the directory to be backed up, that is, the root directory. The root directory can refer to the highest-level directory in the file system and is the starting point for all files and subdirectories; other hierarchical directories are all subdirectories of the root directory; and the parent directory refers to the upper-level directory where the current directory is located.
[0095] For the convenience of understanding by those skilled in the art, Figure 6 An exemplary logic diagram of a method for restoring the backup progress is provided.
[0096] In a specific implementation, the steps may include: S1. The terminal traverses the files in the directory to be backed up (root directory) and the subdirectories in the directory to be backed up; S2. If the currently traversed directory is the target directory that matches the checkpoint directory identifier, perform a sorting operation on the files in the target directory, determine the files with file identifiers less than or equal to the checkpoint file identifier as the already backed-up files and skip them, determine the files with file identifiers greater than the checkpoint file identifier as the unbacked-up files, and then use the file corresponding to the next file identifier after the checkpoint file identifier in the sorting order as the file to start backing up during the backup progress recovery process; S3. If the currently traversed directory is the parent directory of the target directory, perform the following steps: S31. Determine that the files in this parent directory have been backed up and skip them; S32. Perform a sorting operation on the subdirectories in this parent directory according to the directory identifier, and determine the files in the subdirectories with directory identifiers less than or equal to the checkpoint directory identifier as the already backed-up files and skip them, determine the files in the subdirectories with directory identifiers greater than the checkpoint directory identifier as the unbacked-up files, so as to find the file to start backing up during the backup progress recovery process; S32. If the target directory still cannot be found in the subdirectories of this parent directory, the traversal operation can be performed on the parent directory at a deeper level of the target directory and return to the above steps S2 and S3 to search for the unbacked-up files again. For example, if the target directory is "c:\path\to\current", the parent directory at the previous level is "c:\path\to", then the parent directory at a deeper level is "c:\path".
[0097] Exemplarily, there are three subdirectories a, b, and c under the directory "path". If the target directory is "path\c\file1" (the directory of a certain file), when traversing to "path", it is found that "path" is the parent directory of "path\c\file1", then it is necessary to delve into the "path" directory to continue traversing the files and subdirectories in the "path" directory; since the file file1 under "path\c" is already being backed up, the files under "path" and the files under "path\a" and "path\b" can be directly skipped, and the traversal of the subdirectory "path\c" can be continued directly. "path\c" is the directory where the currently backed-up file is located, that is, the target directory, then the files under "path\c" need to be processed. By comparing the file names, the files before file1 are skipped, and the backup starts from the files after file1.
[0098] So far, according to the backup set checkpoint file, the processing of the backup set shards, the restoration of the context information of the backup task, and the restoration of the backup progress of the backup task have been completed, and the backup progress has been restored to the state at the checkpoint moment. Subsequently, the backup can be completed according to the normal backup process.
[0099] In another embodiment, it further includes: obtaining a backup set checkpoint file corresponding to a backup task from a server; discarding target data in a backup set shard corresponding to the backup task according to the backup set shard information corresponding to the checkpoint time recorded in the backup set checkpoint file, to obtain a backup set shard after a truncation operation; the target data is data written to the backup set shard after the checkpoint time; writing format data matching a target storage format into the backup set shard to generate a valid storage format file corresponding to the backup set shard; the target storage format is the storage format of a backup set composed of backup set shards generated according to each backup task; the valid storage format file is used to obtain files that have been backed up to the server at the checkpoint time in case the backup task is interrupted.
[0100] Performing a truncation operation on the backup set shard can truncate the backup set shard to the size of the backup set shard corresponding to the checkpoint time, ensuring that the backup set shard only contains data that has been successfully backed up and is complete at the checkpoint time, and then format data matching the target storage format can be written into the backup set shard.
[0101] The target storage format refers to the format adopted by the backup set finally composed of shards generated by multiple backup tasks, such as tar, mtf, zip, etc. The backup set shards generated by the backup task need to conform to a certain storage format so that the system can identify, manage, and recover data. Writing format data matching the target storage format into the backup set shard makes the backup set shards that have been backed up to the server restored to valid storage format files, which can be regarded as a separate backup set conforming to the target storage format. These files can be directly used for subsequent data recovery or continued backup to ensure data consistency and integrity. Optionally, necessary format data can be supplemented at the end of the backup set shard. For example, the format data can be an end flag or metadata matching the target storage format, etc., to ensure that the backup set shard meets the requirements of the target storage format, so that the data that has been backed up can generate a valid backup set for data recovery through specific processing of the backup set shard.
[0102] The technical solution of this embodiment enables the backup data already generated by a backup-failed job to be fully utilized when necessary by writing specific storage format data. In some extreme cases, it is possible to recover to the most recent checkpoint time from the incomplete backup data, ensuring the maximization of the utility of the backup operation. The generated valid storage format files not only improve the fault tolerance of the backup task but also significantly enhance the data recovery efficiency, avoid redundant operations, optimize resource utilization, and further improve the file backup efficiency.
[0103] In another embodiment, as Figure 7 shown, a file backup method is provided, and this method is applied to Figure 1Taking the terminal 102 in [as an example, the following steps are included:
[0104] S702, obtain the directory to be backed up.
[0105] The directory to be backed up includes files sorted according to the file identifier order and files in subdirectories sorted according to the directory identifier order.
[0106] S704, for each file in the directory to be backed up, generate multiple backup tasks.
[0107] The backup task is used to back up each file to the server according to the file identifier order and the directory identifier order, and generate a backup set shard corresponding to the backup task; the backup set shard is used to write the data of the file; the backup task is associated with context information; the context information includes the backup progress information of the backup task and the backup set shard information; the backup set shard information includes the amount of data of the file that has been written in the backup set shard.
[0108] S706, run multiple backup tasks in parallel, and regularly generate corresponding backup set checkpoint files at the checkpoint moment according to the context information of the backup task, and upload the backup set checkpoint files to the server.
[0109] The backup set checkpoint file is used to restore the backup progress of the backup task in case the backup task is interrupted.
[0110] S708, in case the backup task runs and is interrupted, obtain the backup set checkpoint file corresponding to the backup task from the server.
[0111] S710, according to the backup set checkpoint file, determine the checkpoint file identifier and checkpoint directory identifier recorded at the latest checkpoint moment.
[0112] The checkpoint file identifier is the file identifier corresponding to the last backed-up file recorded at the latest checkpoint moment; the checkpoint directory identifier is the directory identifier corresponding to the directory being backed up at the latest checkpoint moment.
[0113] S712, traverse the directories where each file in the directory to be backed up is located, compare the directory identifier corresponding to each file in the directory to be backed up with the checkpoint directory identifier, skip the backed-up directories, and locate the target directory that matches the checkpoint directory identifier.
[0114] S714, traverse each file in the target directory, compare the file identifier of each file in the target directory with the checkpoint file identifier, skip the backed-up files, locate the backup interruption position in the directory to be backed up, and resume running the backup task at the backup interruption position.
[0115] The backup interruption location is the file location that matches both the checkpoint directory identifier and the checkpoint file identifier.
[0116] In one embodiment, comparing the file identifiers of each file in the target directory with the checkpoint file identifier, skipping the backed-up files, and locating the backup interruption location in the directory to be backed up includes: comparing the file identifiers of each file in the target directory with the checkpoint file identifier, determining the files with file identifiers less than or equal to the checkpoint file identifier as the backed-up files, skipping the backed-up files, and locating the backup interruption location in the directory to be backed up.
[0117] In one embodiment, traversing the directories where each file in the directory to be backed up is located, comparing the directory identifiers corresponding to each file in the directory to be backed up with the checkpoint directory identifier, skipping the backed-up directories, and locating the target directory that matches the checkpoint directory identifier includes: traversing the directories where each file in the directory to be backed up is located; if the target identifier of the currently traversed directory is the parent directory identifier of the checkpoint directory identifier, determining the files in the directory corresponding to the parent directory identifier as the backed-up files, skipping the directory corresponding to the parent directory identifier, and determining the target directory that matches the checkpoint directory identifier from the subdirectories included in the directory corresponding to the parent directory identifier.
[0118] In one embodiment, it further includes: in the case of an interruption in the backup task, obtaining the backup set checkpoint file corresponding to the backup task from the server; determining the backup set shard information recorded at the latest checkpoint moment according to the backup set checkpoint file, discarding the target data in the backup set shards corresponding to the backup task to obtain the backup set shards after truncation operation; the target data is the data written to the backup set shards after the checkpoint moment; restoring the context information of the backup task to the context information corresponding to the latest checkpoint moment according to the backup set checkpoint file; locating the backup interruption location in the directory to be backed up according to the context information corresponding to the latest checkpoint moment, and resuming the execution of the backup task at the backup interruption location to write the unbacked-up files to the backup set shards after truncation operation.
[0119] In one embodiment, it further includes: obtaining a backup set checkpoint file corresponding to a backup task from a server; discarding target data in a backup set shard corresponding to the backup task according to the backup set shard information corresponding to the checkpoint time recorded in the backup set checkpoint file, to obtain a backup set shard after truncation operation; the target data is data written to the backup set shard after the checkpoint time; writing format data matching the target storage format to the backup set shard to generate a valid storage format file corresponding to the backup set shard; the target storage format is the storage format of a backup set composed of backup set shards generated according to each backup task; the valid storage format file is used to obtain files that have been backed up to the server at the checkpoint time in case the backup task is interrupted.
[0120] It can be seen that the above file backup method can continue the backup of an abnormally interrupted file backup job with very low storage and computing consumption, saving backup time, backup and storage resources, and the beneficial effect is particularly obvious for mass file backup. By creating backup checkpoints regularly during the backup process, an aborted file backup job can be restored to the backup state of the nearest backup checkpoint after fault recovery, and continue the previous backup progress in this state until the backup is completed. Compared with the traditional repeated backup strategy, it saves backup time. After fault recovery, the backup data that has been completed before the fault occurs can be continued, and subsequent backups only need to be performed on files that have not been backed up before, saving time and bandwidth and improving the efficiency of file backup; maintaining the integrity of the backup plan, it can restore the previous backup progress at the first time of fault recovery, avoiding missing the backup time window and ensuring the integrity of the backup plan.
[0121] It should be noted that the specific limitations of the above steps can refer to the specific limitations of a file backup method described above.
[0122] It should be understood that although the steps in the flowcharts involved in the above embodiments are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least some of the steps in the flowcharts involved in the above embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily sequential, but can be executed alternately or alternately with at least some of the steps or stages in other steps or other steps.
[0123] Based on the same inventive concept, an embodiment of the present application further provides a file backup device for implementing the file backup method involved above. The solution provided by this device for solving problems is similar to the solution described in the above method. Therefore, the specific limitations in one or more embodiments of the file backup device provided below can refer to the limitations on the file backup method in the above text and will not be repeated here.
[0124] In an exemplary embodiment, as Figure 8 shown, a file backup device is provided, including:
[0125] A file acquisition module 810, configured to acquire a directory to be backed up; the directory to be backed up includes files sorted in the order of file identifiers and files in subdirectories sorted in the order of directory identifiers.
[0126] A task generation module 820, configured to generate multiple backup tasks for each file in the directory to be backed up; the backup tasks are used to back up the respective files to a server in the order of file identifiers and directory identifiers, and generate backup set shards corresponding to the backup tasks; the backup set shards are used to write data of the files; the backup tasks are associated with context information; the context information includes backup progress information and backup set shard information of the backup tasks; the backup set shard information includes the amount of data of the file that has been written in the backup set shard.
[0127] An operation module 830, configured to run the multiple backup tasks in parallel, periodically generate corresponding backup set checkpoint files at checkpoint times according to the context information of the backup tasks, and upload the backup set checkpoint files to the server; the backup set checkpoint files are used to restore the backup progress of the backup tasks in case the backup tasks are interrupted.
[0128] In one embodiment, the running module 830 is specifically configured to, when the backup task running is interrupted, obtain the backup set checkpoint file corresponding to the backup task from the server; determine the checkpoint file identifier and the checkpoint directory identifier recorded at the latest checkpoint time according to the backup set checkpoint file; the checkpoint file identifier is the file identifier corresponding to the last backed-up file recorded at the latest checkpoint time; the checkpoint directory identifier is the directory identifier corresponding to the directory being backed up at the latest checkpoint time; traverse the directories where each file in the directory to be backed up is located, compare the directory identifier corresponding to each file in the directory to be backed up with the checkpoint directory identifier, skip the backed-up directories, and locate the target directory that matches the checkpoint directory identifier; traverse each file in the target directory, compare the file identifier of each file in the target directory with the checkpoint file identifier, skip the backed-up files, locate the backup interruption position in the directory to be backed up, and resume running the backup task at the backup interruption position; the backup interruption position is the file position that matches both the checkpoint directory identifier and the checkpoint file identifier.
[0129] In one embodiment, the running module 830 is specifically configured to compare the file identifier of each file in the target directory with the checkpoint file identifier, determine the files with file identifiers less than or equal to the checkpoint file identifier as the backed-up files, skip the backed-up files, and locate the backup interruption position in the directory to be backed up.
[0130] In one embodiment, the running module 830 is specifically configured to traverse the directories where each file in the directory to be backed up is located; if the target identifier of the currently traversed directory is the parent directory identifier of the checkpoint directory identifier, determine the files in the directory corresponding to the parent directory identifier as the backed-up files, skip the directory corresponding to the parent directory identifier, and determine the target directory that matches the checkpoint directory identifier from the subdirectories included in the directory corresponding to the parent directory identifier.
[0131] In one embodiment, the running module 830 is specifically configured to, when the backup task running is interrupted, obtain the backup set checkpoint file corresponding to the backup task from the server; determine the backup set shard information recorded at the latest checkpoint time according to the backup set checkpoint file, discard the target data in the backup set shards corresponding to the backup task to obtain the backup set shards after truncation operation; the target data is the data written to the backup set shards after the checkpoint time; restore the context information of the backup task to the context information corresponding to the latest checkpoint time according to the backup set checkpoint file; locate the backup interruption position in the directory to be backed up according to the context information corresponding to the latest checkpoint time, and resume running the backup task at the backup interruption position to write the unbacked-up files to the backup set shards after the truncation operation.
[0132] In one embodiment, the file acquisition module 810 is specifically configured to obtain the backup set checkpoint file corresponding to the backup task from the server; discard the target data in the backup set shards corresponding to the backup task according to the backup set shard information corresponding to the checkpoint time recorded in the backup set checkpoint file to obtain the backup set shards after truncation operation; the target data is the data written to the backup set shards after the checkpoint time; write format data matching the target storage format to the backup set shards to generate a valid storage format file corresponding to the backup set shards; the target storage format is the storage format of the backup set composed of the backup set shards generated according to each backup task; the valid storage format file is used to obtain the files that have been backed up to the server at the checkpoint time when the backup task is interrupted.
[0133] Each module in the above file backup device can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor of the computer device in hardware form or be independent of it, or be stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above modules.
[0134] In an exemplary embodiment, a computer device is provided. The computer device can be a terminal, and its internal structure diagram can be as Figure 9As shown in the figure. The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit, and an input device. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface, the display unit, and the input device are connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a mobile cellular network, near field communication (NFC), or other technologies. When the computer program is executed by the processor, it implements a file backup method. The display unit of the computer device is used to form a visually visible picture, which can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the computer device housing, or an external keyboard, touchpad, or mouse, etc.
[0135] Those skilled in the art can understand that Figure 9 the structure shown in the figure is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0136] In an exemplary embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiments are implemented.
[0137] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by the processor, the steps in the above method embodiments are implemented.
[0138] In one embodiment, a computer program product is provided, including a computer program. When the computer program is executed by the processor, the steps in the above method embodiments are implemented.
[0139] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.
[0140] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in this application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in this application can be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, data processing logics based on quantum computing, artificial intelligence (AI) processors, etc., without limitation.
[0141] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this application.
[0142] The above-described embodiments merely represent several implementation manners of this application. The description is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of this application. It should be noted that for those of ordinary skill in the art, without departing from the concept of this application, several deformations and improvements can still be made, and these all belong to the protection scope of this application. Therefore, the protection scope of this application shall be subject to the appended claims.
Claims
1. A file backup method, characterized in that: The method comprises: Obtain a directory to be backed up; the directory to be backed up includes files sorted in file identification order and files under subdirectories sorted in directory identification order; Generate multiple backup tasks for each file in the directory to be backed up; the backup tasks are used to back up each file to a server according to the file identification order and the directory identification order, and generate backup set fragments corresponding to the backup tasks; the backup set fragments are used to write data of the files; the backup tasks are associated with context information; the context information includes backup progress information of the backup tasks and backup set fragment information; the backup set fragment information includes the amount of data of the files that has been written into the backup set fragments; Run the multiple backup tasks in parallel, regularly generate corresponding backup set checkpoint files at checkpoint times according to the context information of the backup tasks, and upload the backup set checkpoint files to the server; the backup set checkpoint files are used to restore the backup progress of the backup tasks in the event that the backup tasks are interrupted.
2. The method according to claim 1, characterized in that The method further comprises: In the case where the backup task is interrupted, obtaining a backup set checkpoint file corresponding to the backup task from the server; According to the backup set checkpoint file, determine the checkpoint file identifier and the checkpoint directory identifier recorded at the latest checkpoint moment; the checkpoint file identifier is the file identifier corresponding to the last backed up file recorded at the latest checkpoint moment; the checkpoint directory identifier is the directory identifier corresponding to the directory being backed up at the latest checkpoint moment; Traversing the directories where each file in the to-be-backed-up directory is located, comparing the directory identifier corresponding to each file in the to-be-backed-up directory with the checkpoint directory identifier, skipping the backed-up directory, and locating the target directory that matches the checkpoint directory identifier; Traverse each file in the target directory, compare the file identifier of each file in the target directory with the checkpoint file identifier, skip the backed up files, locate the backup interruption position in the directory to be backed up, and resume running the backup task at the backup interruption position; the backup interruption position is a file position that matches both the checkpoint directory identifier and the checkpoint file identifier.
3. The method according to claim 2, characterized in that The step of comparing the file identifier of each file in the target directory with the checkpoint file identifier, skipping the backed-up files, and locating the backup interruption position in the directory to be backed up includes: The file identifiers of each file in the target directory are compared with the checkpoint file identifier, and files whose file identifiers are less than or equal to the checkpoint file identifier are determined as the backed-up files, and the backed-up files are skipped to locate the backup interruption position in the directory to be backed up.
4. The method according to claim 2, characterized in that: The traversing the directories where the files in the to-be-backed-up directory are located, comparing the directory identifiers corresponding to the files in the to-be-backed-up directory with the checkpoint directory identifiers, skipping the backed-up directories, and locating the target directory matching the checkpoint directory identifiers, includes: Traverse the directory where each file in the directory to be backed up is located; If the target identifier of the currently traversed directory is the parent directory identifier of the checkpoint directory identifier, the files under the directory corresponding to the parent directory identifier are determined to be backed up files, the directory corresponding to the parent directory identifier is skipped, and the target directory that matches the checkpoint directory identifier is determined from the subdirectories contained in the directory corresponding to the parent directory identifier.
5. The method according to claim 1, characterized in that The method further comprises: In the case where the backup task is interrupted, obtaining a backup set checkpoint file corresponding to the backup task from the server; According to the backup set checkpoint file, the backup set fragment information recorded at the latest checkpoint time is determined, and the target data in the backup set fragment corresponding to the backup task is discarded to obtain the backup set fragment after the truncation operation; the target data is the data written to the backup set fragment after the checkpoint time; According to the backup set checkpoint file, the context information of the backup task is restored to the context information corresponding to the latest checkpoint time; According to the context information corresponding to the latest checkpoint time, the backup interruption position in the directory to be backed up is located, and the backup task is resumed at the backup interruption position to write the unbacked up files to the backup set fragment after the truncation operation.
6. The method according to claim 1, characterized in that The method further comprises: Obtaining a backup set checkpoint file corresponding to the backup task from the server; According to the backup set fragment information corresponding to the checkpoint time recorded in the backup set checkpoint file, the target data in the backup set fragment corresponding to the backup task is discarded to obtain the backup set fragment after the truncation operation; the target data is the data written to the backup set fragment after the checkpoint time; Format data matching the target storage format is written into the backup set fragments to generate a valid storage format file corresponding to the backup set fragments; the target storage format is the storage format of the backup set composed of the backup set fragments generated according to each of the backup tasks; the valid storage format file is used to obtain the files that have been backed up to the server at the checkpoint time in the event that the backup task is interrupted.
7. A file backup device, characterized in that: The device comprises: A file acquisition module is used to acquire files in the directory to be backed up; A task generation module, for generating multiple backup tasks for the files in the directory to be backed up; the backup tasks are used to back up the files to a server and generate backup set fragments corresponding to the backup tasks; the backup set fragments are used to write data of the files; the backup tasks are associated with context information; the context information includes backup progress information of the backup tasks and backup set fragment information; the backup set fragment information includes the amount of data of the files that have been written into the backup set fragments; An operation module is used to run the multiple backup tasks in parallel, regularly generate corresponding backup set checkpoint files at checkpoint times according to the context information of the backup tasks, and upload the backup set checkpoint files to the server; the backup set checkpoint files are used to restore the backup progress of the backup tasks in the event that the backup tasks are interrupted.
8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
File backup recovery method based on sector recombination
CN101477486A
Method, device and equipment for writing in data
CN102999564A
Data back-up method and device
CN106648976A
Backup and recovery method and device for data in database and electronic equipment
CN110874287A
Automatic database backup method and system
CN114138555A
Cited By
Multi-target data backup method and device, computer equipment, readable storage medium and program product
CN121233400A