System for realizing backup of files in virtual machine based on OpenStack platform

Through the backup system based on the OpenStack platform, multi-threaded concurrency and mirroring operations are adopted, the problems of low virtual machine backup efficiency and consistency are solved, efficient and secure virtual machine data backup is achieved, and resource utilization and backup speed are improved.

CN120336085APending Publication Date: 2025-07-18SHANDONG LANGCHAO YUNTOU INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510338179.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing virtual machine backup technology is inefficient, has large resource usage and is difficult to guarantee backup consistency. Especially when the amount of data in the virtualized environment is large and changes frequently, it is difficult to achieve efficient and secure data backup.

Method used

Based on the OpenStack platform, the backup monitoring management module, file system traversal module and data processing module are designed, and a multi-threaded concurrency mechanism is adopted, combining mirroring operations and queue caching mechanisms to ensure the consistency and resource utilization of backup data.

Benefits of technology

It improves backup speed, ensures data consistency, reduces resource consumption, provides high-quality and reliable cloud storage services, solves the problem of data processing rate differences between different storage devices, and realizes an efficient data processing process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336085A_ABST
    Figure CN120336085A_ABST
Patent Text Reader

Abstract

The invention discloses a system for realizing backup of files in a virtual machine based on an OpenStack platform, which relates to the technical field of data storage, and comprises a backup monitoring and management module, which is responsible for scheduling, managing and monitoring backup tasks, including selection of backup targets, issuing of the backup tasks and monitoring of backup progress, meanwhile, after a backup target is selected, the virtual machine is subjected to mirroring operation, the backup state is updated, the backup size is obtained, and temporary resources are cleared up after a backup task is finished; the file system traversal module is responsible for collecting data, distinguishing object types and grouping by traversing a directory structure under a specified path, providing ordered backup type information for the data processing module, and acquiring file metadata information by utilizing a Java standard library; and the data processing module is responsible for concurrently executing data reading, packaging and uploading operations by adopting multiple threads based on a traversal result. According to the invention, the backup speed can be increased, and the problem of rate difference of data processing among different storage devices is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data storage, and more specifically, to a backup system for files in a virtual machine implemented based on the OpenStack platform. Background Art

[0002] With the in-depth development of cloud computing technology in today's society, virtual machines have become an indispensable part of enterprise data centers. However, with the growth of the number of virtual machines, how to efficiently and securely backup the data in virtual machines has become an urgent problem to be solved.

[0003] Due to the dynamic change characteristics of the virtual machine environment, such as frequent file addition, deletion, and modification operations, the backup task has become more complex, and traditional backups are difficult to ensure the consistency and integrity of the backup data. In addition, the resource consumption during the backup process also has an inevitable impact on the virtual machine. Therefore, there is an urgent need for a backup technology architecture that can ensure data consistency and does not cause virtual machine resource consumption to meet the special requirements of the virtualized environment. Summary of the Invention

[0004] Aiming at the problems existing in the existing virtual machine backup technology, such as low efficiency, large resource occupation, and difficulty in ensuring backup consistency, and especially for the characteristics of large data volume and frequent changes in the virtualized environment, the present invention provides a backup system for files in a virtual machine implemented based on the OpenStack platform to achieve highly automated and intelligent backup tasks, and improve backup efficiency and reliability.

[0005] The technical solution adopted by a backup system for files in a virtual machine implemented based on the OpenStack platform of the present invention to solve the above technical problems is as follows:

[0006] A backup system for files in a virtual machine implemented based on the OpenStack platform, which includes:

[0007] A backup monitoring and management module, as the command center of the entire backup system, is responsible for the scheduling, management, and monitoring of backup tasks, including the selection of backup targets, the issuance of backup tasks, the monitoring of backup progress, and at the same time, is responsible for mirroring the virtual machine after selecting the backup target, providing a stable baseline version for the backup work, updating the backup status, obtaining the backup size, and cleaning up temporary resources after the backup task ends;

[0008] A file system traversal module, responsible for data collection, differentiates object types and groups them by traversing the directory structure under the specified path, provides ordered backup type information for the data processing module, and at the same time obtains file metadata information using the Java standard library;

[0009] The data processing module is responsible for concurrently executing data reading, packaging and uploading operations using multiple threads based on the traversal results.

[0010] Optionally, the backup monitoring management module involved is developed in Java language and deployed in an independent management node;

[0011] The file system traversal module involved is developed in Java language and deployed as an agent in the virtual machine that needs to be backed up;

[0012] The data processing modules involved are developed in Java language and deployed as agents in the virtual machines that need to be backed up.

[0013] Optionally, the backup monitoring management module involved specifically includes:

[0014] A target selection unit, used to call the file system traversal module to obtain file and directory information from the virtual machine for backup target selection;

[0015] The mirror operation unit is used to perform a mirror operation on the virtual machine after selecting the backup target, so as to accurately capture and lock all data contents in the virtual machine at the corresponding moment, and provide a reliable and stable baseline version for subsequent backup work;

[0016] The task issuing unit is used to issue a backup command to the file system traversal module and the data processing module deployed in the virtual machine, and send the specific information of the backup task to the virtual machine agent;

[0017] The progress monitoring unit is used to monitor the backup progress in real time until the backup task is completed through a scheduled query mechanism or a corresponding progress reporting mechanism set up on the agent side;

[0018] A status update unit is used to update the backup status after the backup task is completed, mark whether the backup task is successful, and the backup time information, and store it in the status database or configuration file;

[0019] The resource cleanup unit is used to clean up temporary resources generated during the backup process.

[0020] Optionally, the file system traversal modules involved specifically include:

[0021] The traversal acquisition unit is used to traverse the files or directories under the specified path and obtain the metadata information of the files by using the Files.readAttributes method in the Java standard library;

[0022] A parsing and judging unit, used to judge the file type by parsing the metadata information of the file;

[0023] The grouping submission unit is used to group according to the file type and submit the grouping result to the data processing module.

[0024] Optionally, the involved data processing module specifically includes:

[0025] The data packet construction unit is used to construct different types of file data packets for the grouping result;

[0026] The data reading unit is used to read data from different types of file data packets through a data reading thread and put it into a specified position in the queue buffer A;

[0027] The data packing unit is used to take out data from a specified position in the queue buffer A through a data packing thread, perform a packing operation and write it to a specified position in the queue buffer B;

[0028] The data uploading unit is used to take out the packed data from a specified position in the queue buffer B through a data uploading thread and upload it to a backup storage location through a network transmission protocol.

[0029] Further optionally, the involved file types include hard link files, soft link files, directory files and ordinary files;

[0030] The data packet construction unit correspondingly includes a hard link data packet construction subunit, a soft link file data packet construction subunit, a directory data packet construction subunit and an ordinary file data packet construction subunit, where:

[0031] The hard link data packet construction subunit is used to construct a hard link data packet, and the hard link data packet includes a data packet header, a hard link name and a hard link target path;

[0032] The soft link file data packet construction subunit is used to construct a soft link data packet, and the soft link data packet includes a data packet header, a soft link name and a soft link target file name;

[0033] The directory data packet construction subunit is used to construct a directory data packet, and the directory data packet includes a data packet header and a directory name;

[0034] The ordinary file data packet construction subunit is used to construct an ordinary file data packet, and the ordinary file data packet includes a data packet header, a file name and real data.

[0035] Further optionally, the data packet headers of the involved hard link data packets, soft link data packets, directory data packets and ordinary file data packets all include the following contents:

[0036] a) One byte for identifying the data packet version and two bytes for indicating the name length;

[0037] b) Two bytes for identifying the hard link target name;

[0038] c) Two bytes for identifying the soft link target name, and four bytes each for the owner and the group;

[0039] d) Four bytes each for identifying the directory owner and the group, and two bytes for identifying the permissions;

[0040] e) Four bytes each for identifying the file owner and the group, two bytes for identifying the permissions, and 50 bytes for identifying the data length to match a maximum capacity of 1 PiB.

[0041] Further optionally, the involved data packet construction unit constructs hard link data packets, soft link data packets, directory data packets, and ordinary file data packets according to the grouping results;

[0042] For the hard link data packet, the data reading unit first reads the data packet header of the hard link data packet. For the version identification byte, it is used to determine whether the data packet format conforms to the expected version. Then it reads two bytes for indicating the name length, extracts the hard link name from the data packet according to this length information, then reads two bytes for identifying the hard link target name to determine the target name length, extracts the hard link target path, and finally encapsulates the extracted hard link name and hard link target path into a data structure and places it at the specified position in the queue buffer A;

[0043] For the soft link data packet, the data reading unit first reads the version byte of the data packet header in the soft link data packet, then obtains the soft link name length through the next two bytes to extract the soft link name, then reads two bytes for identifying the soft link target name to determine the target name length, extracts the soft link target file name, and at the same time reads four bytes each for the owner and the group for subsequent possible permission processing. Finally, it encapsulates the directory name and permission information into a data structure and places it at the specified position in the queue buffer A;

[0044] For the directory data packet, the data reading unit first reads the version byte of the data packet header in the directory data packet, reads four bytes each for the directory owner and the group and two bytes for identifying the permissions to record the permission information of the directory, then extracts the directory name, and finally encapsulates the directory name and permission information, etc. into a data structure and places it at the specified position in the queue buffer A;

[0045] For ordinary file data packets, the data reading unit first reads the version byte of the data packet header in the ordinary file data packet, then obtains the file name length through the next two bytes, extracts the file name, and then reads four bytes each for identifying the file owner and all groups, two bytes for identifying permissions, and bytes for identifying the data length. According to the data length information, the real data is extracted. Finally, the file name, permission information, and real data are encapsulated into a data structure and placed at a specified position in queue buffer A.

[0046] A virtual machine-internal file backup system implemented based on the OpenStack platform according to the present invention has the following beneficial effects compared with the prior art:

[0047] 1. By introducing a multi-threaded concurrent processing mechanism, the present invention can effectively shorten the backup time required, especially significantly improving the backup speed when dealing with large-scale data sets; during the backup process, by creating a virtual machine image to lock the data state, it ensures data consistency during backup and avoids backup failures or data loss caused by data changes; through a reasonable design of the data processing flow and the adoption of a queue caching mechanism, it can balance the data processing rates between different storage devices, reduce resource waste, and improve the overall system resource utilization rate;

[0048] 2. The present invention is a backup system that can improve the backup speed, ensure the consistency of backup data, and minimize resource consumption as much as possible, providing users with a more high-quality and reliable cloud storage service experience, solving the problem of rate differences in data processing between different storage devices, and thus ensuring the efficient operation of the entire data processing process. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] The appended Figure 1 is the system architecture diagram of Embodiment 1 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0050] To make the technical solutions, technical problems solved, and technical effects of the present invention more clearly understood, the following combines specific embodiments to clearly and completely describe the technical solutions of the present invention.

[0051] Embodiment 1:

[0052] Combined with the appended Figure 1 , this embodiment proposes a virtual machine-internal file backup system implemented based on the OpenStack platform, which includes:

[0053] The backup monitoring management module, as the command center of the entire backup system, is responsible for the scheduling, management, and monitoring of backup tasks, including the selection of backup targets, the issuance of backup tasks, and the monitoring of backup progress. It is also responsible for mirroring virtual machines after selecting backup targets, providing a stable baseline version for backup work, updating backup status, obtaining backup size, and cleaning up temporary resources after the backup task is completed.

[0054] The file system traversal module is responsible for data collection. It traverses the directory structure under the specified path, distinguishes object types and groups them, provides ordered backup type information for the data processing module, and uses the Java standard library to obtain file metadata information;

[0055] The data processing module is responsible for concurrently executing data reading, packaging and uploading operations using multiple threads based on the traversal results.

[0056] In this embodiment, the backup monitoring management module involved is developed in Java language and deployed in an independent management node; the file system traversal module is developed in Java language and deployed as an agent in the virtual machine that needs to be backed up; the data processing module is developed in Java language and deployed as an agent in the virtual machine that needs to be backed up.

[0057] In this embodiment, the backup monitoring management module involved specifically includes:

[0058] A target selection unit, used to call the file system traversal module to obtain file and directory information from the virtual machine for backup target selection;

[0059] The mirror operation unit is used to perform a mirror operation on the virtual machine after selecting the backup target, so as to accurately capture and lock all data contents in the virtual machine at the corresponding moment, and provide a reliable and stable baseline version for subsequent backup work;

[0060] The task issuing unit is used to issue a backup command to the file system traversal module and the data processing module deployed in the virtual machine, and send the specific information of the backup task to the virtual machine agent;

[0061] The progress monitoring unit is used to monitor the backup progress in real time until the backup task is completed through a scheduled query mechanism or a corresponding progress reporting mechanism set up on the agent side;

[0062] A status update unit is used to update the backup status after the backup task is completed, mark whether the backup task is successful, and the backup time information, and store it in the status database or configuration file;

[0063] The resource cleanup unit is used to clean up temporary resources generated during the backup process.

[0064] In this embodiment, the involved file system traversal module specifically includes:

[0065] A traversal acquisition unit, which is used to traverse files or directories under a specified path by using the Files.readAttributes method in the Java standard library to obtain metadata information of the files;

[0066] A parsing and judgment unit, which is used to judge the file type by parsing the metadata information of the files. The file types include hard link files, soft link files, directory files, and ordinary files;

[0067] A grouping and submission unit, which is used to group according to the file type and submit the grouping result to the data processing module.

[0068] In this embodiment, the involved data processing module specifically includes:

[0069] A data packet construction unit, which is used to construct different types of file data packets for the grouping result; a data reading unit, which is used to read data from different types of file data packets through a data reading thread and put it into a specified position in the queue buffer A;

[0070] A data packing unit, which is used to take out data from a specified position in the queue buffer A through a data packing thread, perform a packing operation, and write it to a specified position in the queue buffer B;

[0071] A data uploading unit, which is used to take out the packed data from a specified position in the queue buffer B through a data uploading thread and upload it to a backup storage location through a network transmission protocol.

[0072] Specifically, the data packet construction unit correspondingly includes a hard link data packet construction subunit, a soft link file data packet construction subunit, a directory data packet construction subunit, and an ordinary file data packet construction subunit, where:

[0073] i) The hard link data packet construction subunit is used to construct a hard link data packet, and the hard link data packet includes a data packet header, a hard link name, and a hard link target path;

[0074] ii) The soft link file data packet construction subunit is used to construct a soft link data packet, and the soft link data packet includes a data packet header, a soft link name, and a soft link target file name;

[0075] iii) The directory data packet construction subunit is used to construct a directory data packet, and the directory data packet includes a data packet header and a directory name;

[0076] iv) The ordinary file data packet construction subunit is used to construct an ordinary file data packet, and the ordinary file data packet includes a data packet header, a file name, and real data.

[0077] It should be added that the data headers of hard link data packets, soft link data packets, directory data packets, and ordinary file data packets involved all include the following content:

[0078] a) One byte for identifying the packet version and two bytes for indicating the name length;

[0079] b) Two bytes for identifying the hard link target name;

[0080] c) Two bytes for identifying the soft link target name, four bytes each for the owner and all groups;

[0081] d) Four bytes each for identifying the directory owner and all groups, and two bytes for identifying the permissions;

[0082] e) Four bytes each for identifying the file owner and all groups, two bytes for identifying the permissions, and 50 bytes for identifying the data length to match the maximum 1 PiB capacity of the data.

[0083] That is to say, the packet construction unit constructs hard link data packets, soft link data packets, directory data packets, and ordinary file data packets according to the grouping results.

[0084] For the hard link data packet, the data reading unit first reads the data header of the hard link data packet. For the version identification byte, it is used to determine whether the packet format conforms to the expected version. Then it reads the two bytes for indicating the name length, extracts the hard link name from the packet according to this length information. Then it reads the two bytes for identifying the hard link target name to determine the target name length and extracts the hard link target path. Finally, the extracted hard link name and hard link target path are encapsulated into a data structure and placed in the specified position of the queue buffer A. It should be added that when downloading the hard link data packet, first download the first 5 bytes as the hard link data header, then read the specified length of the data packet again according to the hard link length in the header to obtain the hard link name, and then read the specified length of the data packet according to the target name length in the header to obtain the target name, and execute Files.createLink to create a hard link at the specified path; repeat this step until all hard links are created.

[0085] For soft link data packets, the data reading unit first reads the version byte of the data packet header in the soft link data packet, then obtains the soft link name length through the next two bytes to extract the soft link name, then reads two bytes used to identify the target name of the soft link to determine the target name length, extracts the soft link target file name, and at the same time reads four bytes for each of the owner and all groups for subsequent possible permission processing. Finally, the directory name and permission information are encapsulated into a data structure and placed at a specified position in queue buffer A. It should be added that when downloading the soft link data packet, first download the first 13 bytes as the soft link data packet header, then read the data of the specified length in the data packet again according to the soft link length in the header to obtain the soft link name, and then read the data of the specified length in the data packet according to the target name length in the header to obtain the target name. Execute Files.createSymbolicLink to create a soft link at the specified path and set the owner and all groups of the directory; repeat this step until all soft links are created.

[0086] For directory data packets, the data reading unit first reads the version byte of the data packet header in the directory data packet, reads four bytes for each of the directory owner and all groups and two bytes for identifying permissions to record the permission information of the directory, then extracts the directory name, and finally encapsulates the directory name, permission information, etc. into a data structure and places it at a specified position in queue buffer A. It should be added that when downloading the directory data packet, first download the first 13 bytes as the directory data packet header, then read the data of the specified length in the data packet again according to the directory length in the header to obtain the directory name, execute Files.createDirectory to create a directory at the specified path and set the owner, all groups and permission information of the directory; repeat this step until all directories are created.

[0087] For ordinary file data packets, the data reading unit first reads the version byte of the data packet header in the ordinary file data packet, then obtains the file name length through the next two bytes to extract the file name, then reads four bytes for each of the file owner and all groups, two bytes for identifying permissions, and a byte for identifying the data length, extracts the real data according to the data length information, and finally encapsulates the file name, permission information and real data into a data structure and places it at a specified position in queue buffer A. It should be added that when downloading the ordinary file data packet, first download the first 63 bytes as the file data packet header, then read the data of the specified length in the data packet again according to the file length in the header to obtain the file name, execute Files.write to create a file at the specified path and set the owner, all groups and permission information of the directory, and finally read the corresponding data according to the real data length of the file in the header and write it into the created file; repeat this step until all files are created.

[0088] Based on the backup system described in this embodiment, the specific process of performing a full backup includes the following operations:

[0089] 1) Backup target selection:

[0090] 1.1) The traversal acquisition unit of the file system traversal module uses the Files.readAttributes method in the Java standard library to traverse the files or directories under the specified path and obtain the metadata information of the files; during the traversal process, the traversal acquisition unit stores the metadata information of the traversed files in an appropriate data structure for subsequent processing;

[0091] 1.2) The target selection unit of the backup monitoring and management module selects the backup target according to the metadata information obtained by the traversal acquisition unit;

[0092] 2) Backup task distribution:

[0093] 2.1) After completing the selection of the backup target, the image operation unit of the backup monitoring and management module calls the API of OpenStack to perform an image operation on the virtual machine to accurately capture and lock all the data content in the virtual machine at the corresponding moment, providing a reliable and stable baseline version for subsequent backup work; this operation ensures that the backup is of the consistent state of the virtual machine at a certain moment, preventing data changes during the backup process from affecting the accuracy of the backup;

[0094] 2.2) Based on the selected backup target, the task distribution unit of the backup monitoring and management module sends backup commands to the file system traversal module and the data processing module deployed in the virtual machine, sends the specific information of the backup task to the virtual machine agent, and at the same time, the progress monitoring unit of the backup monitoring and management module starts to monitor the backup progress in real time until the backup task ends.

[0095] 3) Backup task execution:

[0096] 3.1) The file system traversal module adopts a single-threaded traversal method. The parsing and judgment unit parses the metadata information of the file and judges the file type through the Files.readAttributes method. The file types include hard link files, soft link files, directory files, and ordinary files. The specific judgment process is as follows:

[0097] Distinguish hard link files by checking the unix:nlink attribute value: when the attribute value exceeds 1, it is confirmed as a hard link; otherwise, continue to judge;

[0098] Detect soft links using the Files.isSymbolicLink method, and a return value of true indicates that the file is a soft link;

[0099] For non-soft link files, the Files.isDirectory method is used for judgment. If the result is true, it is recognized as a directory file;

[0100] If the above conditions are not met, it is default classified as a normal file;

[0101] 3.2) The grouping submission unit of the file system traversal module groups according to the file type, submits the grouping result to the data processing module. The data processing module processes different types of files in different formats, and writes all the traversed files into the backup file list, which can be used for subsequent data packet construction and processing;

[0102] 3.2.a) For hard link files:

[0103] ① The data processing module checks whether the hard link file exists in the hard link cache (the hard link cache is a key-value structure, the key is the inode number of the hard link, and the value is the remaining hard link count - the hard link target path);

[0104] If it exists, the hard link target path in the hard link cache is obtained, and the remaining hard link count - 1 is updated in the cache. If the remaining hard link count is 0, it means that all the hard links corresponding to this inode have been traversed. To reduce the cache space occupancy, this hard link is deleted from the cache;

[0105] If it does not exist, the hard link information is written into the cache;

[0106] ② The hard link data packet construction subunit constructs a hard link data packet, format: data packet header - hard link name - hard link target path. Among them, the data packet header is 5 bytes. The first byte represents the data packet version, which is used to support subsequent extensions; the second and third bytes represent the length of the hard link name, which can adapt to names with a length not greater than 2 16 length; the fourth and fifth bytes represent the length of the target path, which can adapt to names with a length not greater than 2 16 length;

[0107] ③ The hard link information is packaged and uploaded in a single-threaded manner.

[0108] 3.2.b) For soft link files:

[0109] ① The data processing module obtains the link target of the soft link through the Files.readSymbolicLink method;

[0110] ②Construct a soft link data packet through the soft link file data packet construction subunit, with the format: data packet header - soft link name - soft link target file name. The data packet header is 13 bytes. The first byte represents the data packet version, which is used to support subsequent extensions; the second and third bytes represent the length of the soft link name, which can adapt to names with a length not greater than 2 16 length; the fourth and fifth bytes represent the length of the link target name, which can adapt to names with a length not greater than 2 16 length; the sixth to ninth bytes represent the link owner; the tenth to thirteenth bytes represent the link all group;

[0111] ③Pack and upload the soft link information in a single-threaded manner, and this single-threaded operation ensures the stability of the packing and uploading process of the soft link file.

[0112] 3.2.c) For directory files:

[0113] ①Construct a directory file data packet through the directory data packet construction subunit, with the format: data packet header - directory name. The data packet header is 13 bytes. The first byte represents the data packet version, which is used to support subsequent extensions; the second and third bytes represent the length of the directory name, which can adapt to names with a length not greater than 2 16 length; the fourth and fifth bytes represent the permission information of the directory. The file permission has 9 bits and needs to be stored in 2 bytes; the sixth to ninth bytes represent the directory owner; the tenth to thirteenth bytes represent the directory all group;

[0114] ②Pack and upload the directory information in a single-threaded manner, so as to ensure that the packing and uploading operation of the directory information is carried out in order and avoid potential conflicts in a multi-threaded environment.

[0115] 3.2.d) For ordinary files:

[0116] ①Construct an ordinary file data packet through the ordinary file data packet construction subunit, with the format: data packet header - file name - real data. The data packet header is 63 bytes. The first byte represents the data packet version, which is used to support subsequent extensions; the second and third bytes represent the length of the file name, which can adapt to names with a length not greater than 2 16 length; the fourth and fifth bytes represent the permission information of the file; the sixth to ninth bytes represent the file owner; the tenth to thirteenth bytes represent the file all group; the fourteenth to twenty-first bytes represent the length of the real data. By leveraging the space of these 8 bytes, theoretically, 2 50 bytes of data, that is, 1 PiB of data, can be stored;

[0117] ②Pack and upload the file information and real data concurrently in a multi-threaded manner;

[0118] ③ After the upload is successful, update the total data volume of the corresponding thread data upload to facilitate the update of the backup progress. By recording the amount of data that has been uploaded, the backup progress information can be accurately reflected, providing accurate data for the progress monitoring unit.

[0119] 4) Backup completion and resource cleaning:

[0120] When all data packets are successfully uploaded, update the backup progress to 100% through the status update unit of the backup monitoring and management module; this operation updates the final state of the backup to the corresponding storage (such as a database or configuration file), marking the completion of the backup task;

[0121] When the backup monitoring and management module polls the backup result, call the API of OpenStack through the resource cleaning unit to clean up the temporary resources generated during the backup process in sequence. This operation ensures that resources are released after the backup is completed, avoiding waste and occupation of resources. At the same time, the cleaning process also needs to ensure the stability of the system and avoid affecting other ongoing operations.

[0122] In summary, by using the backup system for files in a virtual machine based on the OpenStack platform of the present invention, a backup system that can improve the backup speed, ensure the consistency of backup data, and minimize resource consumption can provide users with a better quality and more reliable cloud storage service experience, solve the problem of rate differences in data processing between different storage devices, and thus ensure the efficient operation of the entire data processing process.

[0123] The above specific application examples have elaborated in detail the principle and implementation manner of the present invention. These embodiments are only used to help understand the core technical content of the present invention. Based on the above specific embodiments of the present invention, those skilled in the art of this technology field, without departing from the principle of the present invention, any improvements and modifications made to the present invention shall fall within the scope of patent protection of the present invention.

Claims

1. A file backup system for virtual machines implemented based on the OpenStack platform, characterized in that It includes: The backup monitoring management module, as the command center of the entire backup system, is responsible for the scheduling, management, and monitoring of backup tasks, including the selection of backup targets, the issuance of backup tasks, and the monitoring of backup progress. It is also responsible for mirroring virtual machines after selecting backup targets, providing a stable baseline version for backup work, updating backup status, obtaining backup size, and cleaning up temporary resources after the backup task is completed. The file system traversal module is responsible for data collection. It traverses the directory structure under the specified path, distinguishes object types and groups them, provides ordered backup type information for the data processing module, and uses the Java standard library to obtain file metadata information; The data processing module is responsible for concurrently executing data reading, packaging and uploading operations using multiple threads based on the traversal results.

2. The backup system for files in a virtual machine implemented based on the OpenStack platform according to claim 1, wherein, The backup monitoring management module is developed in Java language and deployed in an independent management node; The file system traversal module is developed in Java language and deployed as an agent in the virtual machine that needs to be backed up; The data processing module is developed using Java language and is deployed as an agent in a virtual machine that needs to be backed up.

3. The backup system for files in a virtual machine implemented based on the OpenStack platform according to claim 2, wherein, The backup monitoring management module specifically includes: A target selection unit, used to call the file system traversal module to obtain file and directory information from the virtual machine for backup target selection; The mirror operation unit is used to perform a mirror operation on the virtual machine after selecting the backup target, so as to accurately capture and lock all data contents in the virtual machine at the corresponding moment, and provide a reliable and stable baseline version for subsequent backup work; The task issuing unit is used to issue a backup command to the file system traversal module and the data processing module deployed in the virtual machine, and send the specific information of the backup task to the virtual machine agent; The progress monitoring unit is used to monitor the backup progress in real time until the backup task is completed through a scheduled query mechanism or a corresponding progress reporting mechanism set up on the agent side; A status update unit is used to update the backup status after the backup task is completed, mark whether the backup task is successful, and the backup time information, and store it in the status database or configuration file; The resource cleanup unit is used to clean up temporary resources generated during the backup process.

4. The backup system for files in a virtual machine implemented based on the OpenStack platform according to claim 3, characterized in that, The file system traversal module specifically includes: The traversal acquisition unit is used to traverse the files or directories under the specified path and obtain the metadata information of the files by using the Files.readAttributes method in the Java standard library; A parsing and judging unit, used to judge the file type by parsing the metadata information of the file; The group submission unit is used to group files according to their types and submit the grouping results to the data processing module.

5. A backup system for files within a virtual machine implemented based on the OpenStack platform according to claim 4, wherein, The data processing module specifically includes: A data packet construction unit, used to construct different types of file data packets according to the grouping results; A data reading unit, used to read data from different types of file data packets through a data reading thread and put the data into a designated position of a queue buffer area A; A data packing unit is used to take out data from a specified position of queue buffer area A through a data packing thread, and write the data to a specified position of queue buffer area B after packing operation; A data upload unit, which is used to retrieve the packaged data from a specified location in queue buffer B through a data upload thread and upload it to a backup storage location via a network transmission protocol.

6. The backup system for files in a virtual machine implemented based on the OpenStack platform according to claim 5, characterized in that, The file types include hard link files, soft link files, directory files, and ordinary files; The data packet construction unit correspondingly includes a hard link data packet construction subunit, a soft link file data packet construction subunit, a directory data packet construction subunit, and an ordinary file data packet construction subunit, where: The hard link data packet construction subunit is used to construct a hard link data packet, and the hard link data packet includes a data packet header, a hard link name, and a hard link target path; The soft link file data packet construction subunit is used to construct a soft link data packet, and the soft link data packet includes a data packet header, a soft link name, and a soft link target file name; The directory data packet construction subunit is used to construct a directory data packet, and the directory data packet includes a data packet header and a directory name; The ordinary file data packet construction subunit is used to construct an ordinary file data packet, and the ordinary file data packet includes a data packet header, a file name, and real data.

7. A backup system for files in a virtual machine implemented based on the OpenStack platform according to claim 6, characterized in that, The data packet headers of the hard link data packet, the soft link data packet, the directory data packet, and the ordinary file data packet all include the following contents: a) One byte for identifying the data packet version and two bytes for indicating the name length; b) Two bytes for identifying the hard link target name; c) Two bytes for identifying the soft link target name, and four bytes each for the owner and the group; d) Four bytes each for identifying the directory owner and the group, and two bytes for identifying the permissions; e) Four bytes each for identifying the file owner and the group, two bytes for identifying the permissions, and 50 bytes for identifying the data length to match the maximum 1 PiB capacity of the data.

8. A backup system for files in a virtual machine implemented based on the OpenStack platform according to claim 6, characterized in that, The data packet construction unit constructs hard link data packets, soft link data packets, directory data packets, and ordinary file data packets according to the grouping results; For the hard link data packet, the data reading unit first reads the data packet header of the hard link data packet. For the version identification byte, it is used to determine whether the data packet format conforms to the expected version. Then it reads the two bytes for indicating the name length, extracts the hard link name from the data packet according to this length information. Then it reads the two bytes for identifying the hard link target name to determine the target name length and extracts the hard link target path. Finally, the extracted hard link name and hard link target path are encapsulated into a data structure and placed at a specified location in queue buffer A; For the soft link data packet, the data reading unit first reads the version byte of the data packet header in the soft link data packet, then obtains the soft link name length through the next two bytes to extract the soft link name, then reads two bytes used to identify the target name of the soft link to determine the target name length, extracts the soft link target file name, and at the same time reads four bytes for the owner and all groups respectively for subsequent possible permission processing. Finally, the directory name and permission information are encapsulated into a data structure and placed at the specified position in queue buffer A; For the directory data packet, the data reading unit first reads the version byte of the data packet header in the directory data packet, reads four bytes for the directory owner and all groups respectively and two bytes for identifying the permission to record the permission information of the directory, then extracts the directory name, and finally encapsulates the directory name and permission information, etc. into a data structure and places it at the specified position in queue buffer A; For the ordinary file data packet, the data reading unit first reads the version byte of the data packet header in the ordinary file data packet, then obtains the file name length through the next two bytes to extract the file name, then reads four bytes for identifying the file owner and all groups respectively, two bytes for identifying the permission, and the byte for identifying the data length, extracts the real data according to the data length information, and finally encapsulates the file name, permission information and real data into a data structure and places it at the specified position in queue buffer A.