KVM Data Storage Processing Method, System and Its Continuous Data Point Recovery Method
By combining initialization volume and log volume groups in the data storage of KVM virtual machines and adopting the management method of metadata structures, the tracking and management problems in the data storage are solved, data consistency and integrity are achieved, and a fast solution is provided for data recovery.
Patent Information
- Application Number
- CN202411816798.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-11
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2044-12-11
AI Technical Summary
The prior art is difficult to efficiently track and manage data in the data storage of KVM virtual machines in cloud computing, which makes it difficult to ensure data consistency and integrity, especially in a continuous data protection environment where data growth is fast.
By combining the initialization volume and log volume group, the basic configuration is reasonably organized and the synchronization of data and change data is adopted, and the metadata management method of metadata structure is adopted to achieve efficient data tracking and updates. Dynamic updates to configuration information tables, metadata, initialized volumes, and log volume groups ensure data consistency and integrity.
It realizes efficient tracking and updating data, ensures data consistency and integrity, and provides a basis for rapid location and recovery for data recovery.
Smart Images

Figure CN119645733B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of disaster recovery and backup, and relates to a KVM data storage processing method, system and its continuous data point recovery method. Background Art
[0002] In the field of modern information technology, KVM (Kernel_based Virtual Machine), as an open-source virtualization technology, has become an important choice for building virtualization environments. It allows multiple operating systems to run on the same physical server by simulating hardware, thereby improving resource utilization and flexibility. In this context, the optimization of data storage structures has become an important direction for improving the performance and reliability of virtual machines.
[0003] Currently, like normal hosts, KVM virtual machines in cloud computing also face the risks of data loss and corruption. For this situation, most users adopt backup methods, which are to perform full backups and create snapshots of the data of a KVM virtual machine at a certain time. However, the backup method has a large backup time interval and a large amount of backup data storage. As the disaster recovery system runs for a long time, the data storage volume will continue to increase. Especially in a continuous data protection environment, the data growth rate will be very fast, and it is often impossible to efficiently track data, and the data consistency and integrity cannot be guaranteed.
[0004] Therefore, how to design a data storage structure that can more efficiently track and manage data while ensuring data consistency and integrity is a technical problem that urgently needs to be solved at present. Summary of the Invention
[0005] In order to solve the technical problems in the above background art, the present invention provides a KVM data storage processing method, system and its continuous data point recovery method.
[0006] The technical solutions of the present invention for solving the above technical problems are as follows:
[0007] In a first aspect, a KVM data storage processing method is provided. The method is applied to a server program and includes the steps of:
[0008] A configuration information table creation step of initializing and creating a configuration information table;
[0009] A data volume creation step of initializing and creating an initial volume and a log volume group;
[0010] Metadata structure creation steps: Initialize the creation of the metadata structure, which includes log volume information, mirror table, log metadata file, write operation log, and data block change log. The mirror table stores the latest state of the virtual machine and locates the log metadata file records for each data block; the log metadata file records the versions of all data blocks of the virtual machine at each time point; the data block change log uses a bitmap to store the storage state of all data blocks of the virtual machine at time T.
[0011] Data storage steps: Receive and determine the source data type from the production end. If the source data is basic configuration and synchronization data, first update the configuration information table, and then update the initialization volume and metadata structure; if the source data is data change records, directly update the log volume group and metadata structure.
[0012] In a second aspect, a KVM data storage processing system is provided. The system is applied to a server program and includes:
[0013] Configuration information table creation module, used to initialize the creation of the configuration information table;
[0014] Data volume creation module, used to initialize the creation of the initialization volume and log volume group;
[0015] Metadata structure creation module, used to initialize the creation of the metadata structure, which includes log volume information, mirror table, log metadata file, write operation log, and data block change log. The mirror table stores the latest state of the virtual machine and locates the log metadata file records for each data block, and the log metadata file records the versions of all data blocks of the virtual machine at each time point; the data block change log uses a bitmap to store the storage state of all data blocks of the virtual machine at time T.
[0016] Judgment and update module, used to receive and determine the source data type from the production end. If the source data is basic configuration and synchronization data, first update the configuration information table, and then update the initialization volume and metadata structure; if the source data is data change records, directly update the log volume group and metadata structure.
[0017] In a third aspect, a continuous data point recovery method is provided. Using the above KVM data storage processing method, it further includes the steps of:
[0018] Create recovery volume step: Receive a recovery instruction and create a recovery volume;
[0019] Judgment and execution step: Judge whether the recovery time point M0 is greater than or equal to the latest time point H0. If so, execute the latest time point recovery step; if not, execute the arbitrary time point recovery step;
[0020] Latest time point recovery step: Traverse the mirror table to determine whether the latest content of the data block is in the log volume group or the initialization volume, and then read the data from the log volume group or the initialization volume and write it into the recovery volume. Among them, reading the log volume group is done by reading through the log metadata file;
[0021] Any time point recovery step: Obtain the bitmap at time M0 in the data block change log, then traverse the bitmap to determine whether the latest content of the data block is in the log volume group or the initialization volume, and then read the data from the log volume group or the initialization volume and write it into the recovery volume. Among them, reading the log volume group is performed by combining the mirror table and the log metadata file for positioning and reading, and reading the initialization volume is done by positioning through the mirror table;
[0022] Step-by-step rollback step: When the internal data of the virtual machine after recovery at the latest time point or any time point is not in a logically consistent state, locate the position of the data block before the recovery time point M0 through the write operation log and the mirror table, then reverse-trace the log metadata file to determine the last valid change content of each data block, and then read the corresponding data from the initialization volume or the log volume according to the last valid change content and overwrite it and write it into the recovery volume.
[0023] The beneficial effects of the present invention are:
[0024] The present invention combines the way of the initialization volume and the log volume group, reasonably organizes the basic configuration, synchronized data and changed data, and adopts the metadata management method of the metadata structure, which can efficiently track and update data; at the same time, the dynamic update of the configuration information table, metadata, initialization volume and log volume group ensures data consistency and integrity, and also lays a foundation for quickly locating and recovering the required data during the data recovery process. Description of the Drawings
[0025] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained according to these drawings.
[0026] Figure 1 It is a schematic flowchart of the KVM data storage processing method in Embodiment 1 of the present invention.
[0027] Figure 2 It is a schematic diagram of the data storage structure in Embodiment 1 of the present invention.
[0028] Figure 3 It is a schematic diagram of the log metadata file structure in Embodiment 1 of the present invention.
[0029] Figure 4Schematic flowchart of the KVM data storage processing method in Embodiment 2 of the present invention.
[0030] Figure 5 Schematic structural diagram of the KVM data storage processing system in Embodiment 3 of the present invention.
[0031] Figure 6 Schematic structural diagram of the judgment and update module in Embodiment 3 of the present invention.
[0032] Figure 7 Schematic flowchart of the continuous data point recovery method in Embodiment 4 of the present invention.
[0033] In the attached drawings, the list of components represented by each reference numeral is as follows:
[0034] 3001, configuration information table creation module; 3002, data volume creation module; 3003, metadata structure creation module; 3004, judgment and update module; 3005, write initial volume module; 3006, update bitmap data module; 3007, delete log volume and its log volume information module; 3008, modify metadata module; 30041, receiving unit; 30042, judgment unit; 30043, basic configuration and synchronization data storage unit; 30044, changed data storage unit. Detailed implementation manners
[0035] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the attached drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0036] Embodiment 1
[0037] Currently, like a normal host, the KVM virtual machine in cloud computing also faces the risk of data loss and corruption. For this situation, most users adopt a backup method, which is to perform a full backup and create a snapshot of the data of the KVM virtual machine at a certain time. However, the backup method has a large backup time interval and a large amount of backup data storage. As the disaster recovery system runs for a long time, the data storage amount will continue to increase. Especially in a continuous data protection environment, the data growth rate will be very fast, and it is often impossible to efficiently track the data, and the data consistency and integrity cannot be guaranteed. Therefore, how to design a data storage structure that can more efficiently track and manage data while ensuring data consistency and integrity is a technical problem that needs to be solved urgently at present.
[0038] In view of the above problems, an embodiment of the present invention provides a KVM data storage processing method. Figure 1 Schematic flowchart of the KVM data storage processing method provided by an embodiment of the present invention, combined with Figure 1 andFigure 2 As shown in the figure, the method includes:
[0039] Step S101, initialize and create a configuration information table.
[0040] It can be understood that the configuration information table stores the basic information of all virtual machines. Through this table, it includes the unique identification code of the virtual machine, the total disk size, and the bitmap data path. The configuration information table is arranged according to the unique identification code, and records can be found according to the unique identification code; the initialization volume and the log volume group can be generated according to the total disk size; the bitmap data path stores the absolute path of the sent bitmap data at the local end, and the bitmap data can be located.
[0041] It can also be understood that the configuration information table locates the information position of the target virtual machine, which refers to locating whether the information of the target virtual machine is in the initialization volume, the log volume group, or the metadata structure.
[0042] Step S102, initialize and create an initialization volume and a log volume group.
[0043] It can be understood that since the management mechanism of directly operating on block devices is relatively complex, in this embodiment, files are used to store data, which improves the flexibility of data storage and enables block devices to be operated in a simpler way. The initialization volume stores the basic configuration of the virtual machine and the synchronized data, and the size of the volume can be the same as the total disk size of the virtual machine. The log volume group stores the changed data generated at the production end in sequence, and the size of each log volume can be half of the total disk size of the virtual machine.
[0044] Step S103, initialize and create a metadata structure. The metadata structure includes log volume information, an image table, a log metadata file, a write operation log, and a data block change log. Among them, the image table stores the latest state of the virtual machine and locates the log metadata file records of each data block, and the log metadata file records the versions of all data blocks of the virtual machine at each time point; the data block change log uses a bitmap to store the storage state of all data blocks of the virtual machine at time T.
[0045] It can be understood that the write operation log stores the information of each write operation of the virtual machine. The write operation log dynamically increases as the virtual machine continues to write and is stored in the order of write operations. Each record contains a timestamp, the starting logical block number of the write, and the number of data blocks written.
[0046] It can also be understood that the action mechanism of the data block change log is similar to the snapshot principle. After rolling back the time at any time point, the data volume type where the data block is stored can be directly queried according to the bitmap, minimizing the traversal of the log meta-file. The data block change log record includes the timestamp when the bitmap is created and the absolute path of the bitmap.
[0047] The data block change log dynamically adds records according to the time interval set by the user, and writes the data block flag of the mirror table at that moment into the created bitmap, so that the records can be arranged and searched in the order of timestamps. Each value of the bitmap should be consistent with the data block flag value.
[0048] Optionally, in step S103, if Figure 3 As shown, the log metadata file is a double linked list structure, each log metadata file record includes a head pointer head and a tail pointer last, the head pointer head points to the data block version that was changed for the first time and finds the tail pointer last, the tail pointer last points to the data block version that was changed most recently; the double linked list structure has a head node, a last node and a middle node;
[0049] The mirror table includes a data block flag flag and a log metadata offset log_meta, wherein the log metadata offset log_meta is the offset of the log metadata of the data block recorded in the log metadata file, and the log metadata offset log_meta stores the address of the head pointer head.
[0050] It can be understood that the double linked list structure has a head node, a last node and a middle node. The double linked list supports bidirectional traversal, which can be from the head to the end of the table or from the end of the table to the head of the table.
[0051] It can also be understood that the log metadata file dynamically increases records as the virtual machine continues to write, and each record stores the versions of a data block at each time point. Unlike traditional log metadata files, each log metadata file record uses a bidirectional linked list structure to organize and store the versions of the data block at each time point. Of course, by traversing the log metadata file, it is possible to quickly restore to any continuous data point, thereby providing a guarantee for subsequent accelerated continuous recovery.
[0052] It is also worth pointing out that Figure 3 As shown, each node of the log metadata file contains a data domain, a pre pointer, and a next pointer. The pre pointer and the next pointer point to the previous and next nodes of the node respectively, so the log metadata record can be traversed in positive and reverse order respectively through the pre pointer and the next pointer. The data domain is a structure Data_Info, which contains the microsecond write timestamp of the data block, the log volume file name log_volume where the data block is written, and the offset within the log volume where the data block is written. Therefore, the position of the data block in the log volume can be located through the data domain of the node.
[0053] It should be noted that the order of initializing and creating the configuration information table, initializing the volume, the log volume group, and the metadata structure can be processed in parallel to improve efficiency, but it is necessary to ensure that the dependencies are met, that is, the configuration information table depends on the initialized volume and the log volume group, and the initialized volume and the log volume group depend on the metadata structure; it can also be created sequentially. For example, first initialize and create the configuration information table, then create the initialized volume and the log volume group, and finally create the metadata structure. This embodiment does not make specific limitations on this.
[0054] Step S104, receive and determine the source data type at the production end. If the source data is basic configuration and synchronization data, first update the configuration information table, and then update the initialized volume and the metadata structure; if the source data is a data change record, directly update the log volume group and the metadata structure.
[0055] Optionally, step S102 includes:
[0056] Step S1021, receive the source data at the production end;
[0057] Step S1022, determine the type of the source data. If the source data is configuration information, bitmap data, and non-hole synchronization data, execute the basic configuration and synchronization data storage step; if the source data is a data change record, execute the changed data storage step, where the data change record includes a unique identifier, a timestamp, an offset, a data length, and data content;
[0058] Step S1023, for the configuration information, first generate a configuration information table record according to the received configuration information and add it to the configuration information table, and then generate the data volume and metadata of the target virtual machine according to the generated record; for the bitmap data, first initialize the metadata according to the bitmap data, and then write the absolute path of the bitmap data into the target virtual machine record in the configuration information table; for the non-hole synchronization data, first convert the non-hole synchronization data and the bitmap data to obtain the initialized synchronization data, and then write the initialized synchronization data into the data volume;
[0059] Step S1024, find the metadata and data volume of the virtual machine according to the unique identifier in the data change record, then write the data content in the record into the log volume group, and then update the write operation log, update the mirror table, update the log metadata file, and update the data block change log.
[0060] It can be understood that the configuration information stores virtual machine information, including the unique identifier of the virtual machine and the total size of the virtual disk file. The non-hole synchronization data is a non-hole version of the initialized synchronization data of the virtual machine, and it can be converted into the initialized synchronization data in combination with the bitmap data. The bitmap data is the data block storage state of the initialized synchronization data of the virtual machine, and it can be converted into the initialized synchronization data in combination with the non-hole synchronization data.
[0061] In this embodiment, by combining the initialization of the volume and the log volume group, the basic configuration, synchronization data, and changed data are reasonably organized, and the metadata management method of the metadata structure is adopted, which can efficiently track and update data. At the same time, the dynamic update of the configuration information table, metadata, initialization volume, and log volume group ensures data consistency and integrity, and also lays a foundation for quickly locating and restoring the required data during the data recovery process.
[0062] Embodiment 2
[0063] As Figure 4 shown, a KVM data storage processing method is provided. The method is applied to a server program and includes the steps:
[0064] Step S201, initialize and create a configuration information table;
[0065] Step S202, initialize and create an initialization volume and a log volume group. The log volume group contains multiple log volumes, and each log volume correspondingly contains log volume information in a metadata structure;
[0066] Step S203, create a metadata structure including log volume information, an image table, a log metadata file, a write operation log, and a data block change log. The image table stores the latest state of the virtual machine and locates the log metadata file records of each data block;
[0067] The log metadata file records the versions of all data blocks of the virtual machine at each time point and also stores the storage locations of all data in the log volume group;
[0068] The data block change log uses a bitmap to store the storage status of all data blocks of the virtual machine at time T;
[0069] The log volume information includes a data block offset offset and a data block timestamp time. The data block offset offset is the offset of each data block stored in this log volume in the initialization volume, and the data block timestamp time is the timestamp when the last data block in this log volume was written;
[0070] Step S204, receive and determine the source data type of the production end. If the source data is basic configuration and synchronization data, first update the configuration information table, and then update the initialization volume and the metadata structure. If the source data is a data change record, directly update the log volume group and the metadata structure;
[0071] Step S205, find the old log volume LV0 and its log volume information in the log volume group, traverse the data block offset offset of the log volume information in the log volume LV0, and then overwrite and write the contents of all data blocks in the log volume LV0 into the initialization volume according to the data block offset offset;
[0072] Step S206: Search the configuration information table according to the unique identifier of the target virtual machine to obtain the bitmap data of the target virtual machine, and then synchronously traverse the initialization volume and the bitmap data to update the bitmap data.
[0073] Step S207: Write all the data of the log volume LV0 into the initialization volume, and then delete the log volume LV0 and its corresponding log volume information.
[0074] Step S208: Correspondingly modify the mirror table, the log metadata file, the write operation log, and the data block change log.
[0075] Different from the above embodiments, this embodiment provides a storage space recycling processing strategy. When the data storage reaches the limit, the data content of the virtual machine log volume is merged into the initialization volume, and the relevant metadata is modified and deleted. This not only ensures the long-term stable operation of the system, but also improves the utilization efficiency of storage, adapts to the growing demand for data volume, and enhances the scalability and sustainability of the system.
[0076] Embodiment 3
[0077] As Figure 5 shown, a KVM data storage processing system is provided. The system is applied to a server program and includes:
[0078] A configuration information table creation module 3001, which is used to initialize and create a configuration information table.
[0079] A data volume creation module 3002, which is used to initialize and create an initialization volume and a log volume group.
[0080] A metadata structure creation module 3003, which is used to initialize and create a metadata structure. The metadata structure includes log volume information, a mirror table, a log metadata file, a write operation log, and a data block change log. The mirror table stores the latest state of the virtual machine and locates the log metadata file records of each data block. The log metadata file records the versions of all data blocks of the virtual machine at each time point. The data block change log uses a bitmap to store the storage state of all data blocks of the virtual machine at time T.
[0081] A judgment and update module 3004, which is used to receive and judge the source data type of the production end. If the source data is basic configuration and synchronization data, the configuration information table is updated first, and then the initialization volume and the metadata structure are updated. If the source data is a data change record, the log volume group and the metadata structure are directly updated.
[0082] Optionally, in the metadata structure creation module 3003, the log metadata file has a doubly linked list structure. Each log metadata file record contains a head pointer "head" and a tail pointer "last". The head pointer "head" points to the version of the data block that was changed first and locates the tail pointer "last", and the tail pointer "last" points to the version of the data block that was changed most recently;
[0083] The mirror table includes a data block flag "flag" and a log metadata offset "log_meta". The log metadata offset "log_meta" is the offset of the log metadata record of the data block in the log metadata file, and the log metadata offset "log_meta" stores the address of the head pointer "head".
[0084] Optionally, in the data volume creation module 3002, the log volume group contains multiple log volumes, and each log volume correspondingly contains the log volume information in a metadata structure;
[0085] In the metadata structure creation module 3003, the log metadata file stores the storage locations of all data in the log volume group; the log volume information includes a data block offset "offset" and a data block timestamp "time", where the data block offset "offset" is the offset of each data block stored in this log volume in the initialized volume, and the data block timestamp "time" is the timestamp when the last data block stored in this log volume was written;
[0086] Similarly, as Figure 5 shown, the system further includes:
[0087] The initial volume writing module 3005 is used to find the old log volume LV0 and its log volume information in the log volume group, traverse the data block offsets "offset" of the log volume information in the log volume LV0, and then overwrite and write the contents of all data blocks in the log volume LV0 into the initialized volume according to the data block offset "offset";
[0088] The bitmap data update module 3006 is used to find the configuration information table according to the unique identifier of the target virtual machine, obtain the bitmap data of the target virtual machine, and then synchronously traverse the initialized volume and the bitmap data to update the bitmap data;
[0089] The log volume and its log volume information deletion module 3007 is used to delete the log volume LV0 and its corresponding log volume information after writing all the data of the log volume LV0 into the initialized volume;
[0090] The metadata modification module 3008 is used to correspondingly modify the mirror table, the log metadata file, the write operation log, and the data block change log
[0091] As Figure 6As shown, optionally, the determination and update module 3004 includes:
[0092] A receiving unit 30041, configured to receive source data from the production end;
[0093] A determination unit 30042, configured to determine the type of the source data. If the source data is configuration information, bitmap data, and hole-free synchronization data, then execute the basic configuration and synchronization data storage steps; if the source data is a data change record, then execute the change data storage steps, where the data change record includes a unique identification code, a timestamp, an offset, a data length, and data content;
[0094] A basic configuration and synchronization data storage unit 30043, for configuration information, first generate a configuration information table record according to the received configuration information and add it to the configuration information table, and then generate a data volume and metadata of the target virtual machine according to the generated record; for bitmap data, first initialize the metadata according to the bitmap data, and then write the absolute path of the bitmap data into the target virtual machine record in the configuration information table; for hole-free synchronization data, first convert the hole-free synchronization data and the bitmap data to obtain initialized synchronization data, and then write the initialized synchronization data into the data volume;
[0095] A change data storage unit 30044, configured to find the metadata and data volume of the virtual machine according to the unique identification code in the data change record, then write the data content in the record into the log volume group, and then update the write operation log, update the mirror table, update the log metadata file, and update the data block change log.
[0096] In this embodiment, by integrating the initialization volume and the log volume group and relying on the metadata structure to implement data management, an efficient data tracking and update framework is constructed, the basic configuration and data synchronization processes are optimized, and at the same time, the consistency and integrity of the data are ensured through the dynamic update mechanism.
[0097] Embodiment 4
[0098] As Figure 7 shown, a continuous data point recovery method is provided. Using the KVM data storage processing method described in Embodiment 1 or 2, it includes the steps of:
[0099] Step S401, initialize and create a configuration information table;
[0100] Step S402, initialize and create an initialization volume and a log volume group;
[0101] Step S403: Initialize and create a metadata structure, which includes log volume information, an image table, log metadata files, write operation logs, and data block change logs. The image table stores the latest state of the virtual machine and locates the log metadata file records for each data block. The log metadata files record the versions of all data blocks of the virtual machine at each time point. The data block change logs use a bitmap to store the storage status of all data blocks of the virtual machine at time T.
[0102] Step S404: Receive and determine the source data type at the production end. If the source data is basic configuration and synchronization data, update the configuration information table first, and then update the initialization volume and metadata structure. If the source data is a data change record, directly update the log volume group and metadata structure.
[0103] Step S405: Receive a recovery instruction and create a recovery volume.
[0104] Step S406: Determine whether the recovery time point M0 is greater than or equal to the latest time point H0. If so, execute Step S407; if not, execute Step S108.
[0105] Step S407: Traverse the image table to determine whether the latest content of the data block is in the log volume group or the initialization volume, and then read the data from the log volume group or the initialization volume and write it into the recovery volume. Reading the log volume group is done through the log metadata files.
[0106] Step S408: Obtain the bitmap at time M0 in the data block change logs, then traverse the bitmap to determine whether the latest content of the data block is in the log volume group or the initialization volume, and then read the data from the log volume group or the initialization volume and write it into the recovery volume. Reading the log volume group is performed by combining the image table and the log metadata file for positioning, and reading the initialization volume is done through the image table for positioning.
[0107] Step S409: When the internal data of the virtual machine after recovery at the latest time point or any time point is not in a logically consistent state, locate the position of the data block before the recovery time point M0 through the write operation logs and the image table, then trace back the log metadata files in reverse to determine the last valid change content of each data block, and then read the corresponding data from the initialization volume or the log volume according to the last valid change content and overwrite it into the recovery volume.
[0108] It can be understood that before recovery, the server needs to create a recovery volume. Generally, the data of the recovery volume is set to zero, and the size of the recovery volume during recovery is the same as that of the initialization volume.
[0109] The latest time point H0 refers to the time point when the system last successfully recorded data changes. It is the snapshot time point of the latest data state in the system and represents the latest version of the current data.
[0110] The recovery time point M0 refers to the target time point to which the data is desired to be restored. This time point can be any past instant. Due to reasons such as data loss, corruption, or errors, it is desired to restore the data to this specific historical state.
[0111] In this embodiment, if the recovery time point M0 is greater than or equal to the latest time point H0, it means that the user desires to restore to a state containing the latest data. Therefore, the latest time point recovery step is executed. If the recovery time point M0 is less than the latest time point H0, the user desires to restore to an earlier state. Therefore, the any time point recovery step is executed.
[0112] It can be understood that the data block flag flag of the mirror table has three values. 0 indicates that the latest content of the data block is in the log volume group. 1 indicates that the latest content of the data block is in the initialization volume. 2 indicates that the data block is invalid data, that is, the block is an unallocated block or the data is zero. The log metadata offset log_meta is the offset of the log metadata record of the data block in the log metadata file, and the log metadata offset log_meta stores the address of the head pointer head.
[0113] In this embodiment, for the latest time point recovery, this embodiment traverses the mirror table and checks the data block flag flag field of each record to determine whether the latest content of the data block is in the log volume group or the initialization volume. If the retrieved data block flag flag is 2, it can be directly skipped. If the retrieved mirror table data block record has a data block flag flag of 1, since the offset of this record in the mirror table is actually the offset of the data block in the initialization volume, the data can be directly read from the initialization volume according to the offset of this record in the mirror table. If the retrieved data block flag flag is 0, it is necessary to locate the log metadata file according to the log metadata offset log_meta. The log metadata file records the versions of all data blocks of the virtual machine at each time point. Then, the log volume is read through the log metadata file. More specifically, the head node in the log metadata file is located according to the log metadata offset log_meta, the last node is found through the pre pointer of the head node, and the log volume file name (log_volume) and the offset within the volume in the Data_Info structure are obtained, and the data is read from the log volume according to the log volume file name and the offset within the volume.
[0114] In this embodiment, for recovery at any time point, the data block change log uses a bitmap to store the storage status of all data blocks of the virtual machine at a certain moment. After the time of rollback at any time point, the data volume type where the data block is stored can be directly queried according to the bitmap. The record closest to the user-specified time point M0 can be quickly located through the binary search method. After determining the record closest to M0, the status of each data block can be quickly checked through the bitmap. If the bit value is 2, it means the data block is invalid and is skipped; if the bit value is 1, it means the latest data of the data block is in the initialization volume; if the bit value is 0, it means the data may be in the log volume group. If it is in the initialization volume, the mirror table record is located according to the offset in the bitmap, and the data is read from the initialization volume; if it may be in the log volume group, the mirror table record is located according to the offset in the bitmap, and then the head node in the log metadata file is located according to the log metadata offset log_meta. The log metadata records are traversed in reverse order to find the node with a timestamp less than or equal to M0, and the log volume file name and offset are obtained, and the data is read from the log volume. If no qualified node is found, it means the data is in the initialization volume.
[0115] In addition, step-by-step rollback recovery is a method to ensure that the data is restored to the specified historical state. By gradually rolling back to the nearest consistent state, the consistency and integrity of the data are ensured.
[0116] In this embodiment, the write operation log records all write operations, and the mirror table provides the current status and location information of the data blocks. Combining these two parts, the last changed content of the data blocks can be determined; the log metadata file records the versions of all data blocks of the virtual machine at each time point, and the data status at the specified time point can be found through reverse traversal; and by gradually rolling back and overwriting the write recovery volume, it can be ensured that the restored data is consistent with the status at the specified time point, and even in the case of data inconsistency, it can be handled through rollback.
[0117] More specifically, use the binary search method to locate the write operation record with a timestamp less than or equal to M0 and closest to the recovery time point M0. According to the starting block number and the number of blocks in the write operation log, find the mirror table record to determine the head node of each data block in the log metadata file. Traverse the log metadata records in reverse order to find the node with a timestamp less than or equal to M0, and determine the last changed content of the data block.
[0118] If the previous node of the target node is null, the data is stored in the initialization volume; if not, the data is stored in the log volume. If the data is stored in the initialization volume, the data is directly read from the initialization volume according to the offset in the mirror table; if the data is stored in the log volume, the data is read from the log volume according to the log volume file name and offset in the Data_Info structure of the previous node. Then, the data read from the initialization volume or the log volume is overwritten and written to the recovery volume according to the offset in the mirror table. If the internal data of the virtual machine after recovery is inconsistent, continue to roll back and repeat the above operations.
[0119] This embodiment supports multiple recovery methods, including recovery at the latest time point, recovery at any time point, and step-by-step rollback recovery, which can meet different business requirements. When data loss or errors occur, users can select the most suitable recovery plan according to the actual situation. By traversing the mirror table and the log metadata file, the required data can be accurately recovered from the initialization volume and the log volume group, ensuring fast recovery in different failure situations.
[0120] The computer storage medium for storing the embodiments of the present invention may adopt any combination of one or more computer-readable media. The computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the computer-readable storage medium include: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this document, the computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0121] The computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries the computer-readable program code. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium may also be any computer-readable medium other than the computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device.
[0122] The program code contained on a computer-readable medium can be transmitted with any suitable medium, including but not limited to wireless, wire, optical fiber cable, RF, etc., or any suitable combination of the above.
[0123] The computer program code for performing the operations of the present invention can be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, Ruby, Go, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).
[0124] The above-described embodiments merely represent several implementation manners of the present invention. The description is relatively specific and detailed, but should not be construed as a limitation on the scope of the patent for the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the patent for the present invention shall be subject to the appended claims.
Claims
1. A KVM data storage and processing method, characterized in that: The method is applied to a server program and comprises the steps of: Configuration information table creation step, initialization creation of configuration information table; Data volume creation steps: Initialize and create initialization volume and log volume group; The metadata structure creation step is to initialize and create the metadata structure. The metadata structure includes log volume information, mirror table, log metadata file, write operation log and data block change log. The mirror table stores the latest state of the virtual machine and locates the log metadata file record of each data block; the log metadata file records the versions of all data blocks of the virtual machine at each time point. The data block change log uses bitmap to store the storage status of all data blocks of the virtual machine at time T; Data storage step: receiving and determining the source data type of the production end. If the source data is basic configuration and synchronization data, the configuration information table is updated first, and then the initialization volume and metadata structure are updated; If the source data is a data change record, the log volume group and metadata structure are directly updated; The data storage step includes receiving and determining the source data type of the production end: Receiving step, receiving source data from the production end; A judgment step, judging the type of source data, if the source data is configuration information, bitmap data and synchronization data without holes, then executing the basic configuration and synchronization data storage steps; if the source data is a data change record, then executing the change data storage step, wherein the data change record includes a unique identification code, a timestamp, an offset, a data length and a data content; Basic configuration and synchronization data storage steps: for configuration information, first generate a configuration information table record according to the received configuration information and add it to the configuration information table, and then generate the data volume and metadata of the target virtual machine according to the generated record; for bitmap data, first initialize the metadata according to the bitmap data, and then write the absolute path of the bitmap data into the target virtual machine record in the configuration information table; for synchronization data without holes, first convert the synchronization data without holes and the bitmap data to obtain the initialization synchronization data, and then write the initialization synchronization data into the data volume; The step of storing changed data is to search for the metadata and data volume of the virtual machine according to the unique identification code in the data change record, write the data content in the record into the log volume group, and then update the write operation log, update the mirror table, update the log metadata file and update the data block change log.
2. The KVM data storage and processing method according to claim 1, characterized in that: In the metadata structure creation step, the log metadata file is a double linked list structure, and each log metadata file record includes a head pointer head and a tail pointer last, the head pointer head points to the data block version that was changed for the first time and finds the tail pointer last, and the tail pointer last points to the data block version that was changed most recently; The mirror table includes a data block flag flag and a log metadata offset log_meta, wherein the log metadata offset log_meta is the offset of the log metadata of the data block recorded in the log metadata file, and the log metadata offset log_meta stores the address of the head pointer head.
3. The KVM data storage and processing method according to claim 1, characterized in that: In the data volume creation step, the log volume group includes a plurality of log volumes, and each log volume corresponds to log volume information in a metadata structure; In the metadata structure creation step, the log metadata file stores the storage location of all data in the log volume group; the log volume information includes a data block offset and a data block timestamp, wherein the data block offset is the offset of each data block stored in the log volume in the initialization volume, and the data block timestamp is the timestamp when the last data block stored in the log volume is written.
4. The KVM data storage and processing method according to claim 3, characterized in that: The method further comprises the steps of: In the step of writing the initial volume, find the old log volume LV0 and its log volume information in the log volume group, traverse the data block offset offset of the log volume information in the log volume LV0, and then overwrite all the data block contents in the log volume LV0 with the initialization volume according to the data block offset offset; The step of updating the bitmap data is to search the configuration information table according to the unique identification code of the target virtual machine, obtain the bitmap data of the target virtual machine, and then synchronously traverse the initialization volume and the bitmap data to update the bitmap data; The step of deleting the log volume and its log volume information is to write all the data of the log volume LV0 into the initialization volume and then delete the log volume LV0 and its corresponding log volume information; The metadata modification step corresponds to modifying the mirror table, log metadata file, writing operation log and data block change log.
5. A KVM data storage and processing system, characterized in that: The system is applied to a server program, including: A configuration information table creation module is used to initialize and create a configuration information table; The data volume creation module is used to initialize and create the initialization volume and log volume group; The metadata structure creation module is used to initialize and create the metadata structure. The metadata structure includes log volume information, mirror table, log metadata file, write operation log and data block change log. The mirror table stores the latest state of the virtual machine and locates the log metadata file record of each data block. The log metadata file records the versions of all data blocks of the virtual machine at each time point. The data block change log uses bitmap to store the storage state of all data blocks of the virtual machine at time T. The judgment update module is used to receive and judge the source data type of the production end. If the source data is basic configuration and synchronization data, the configuration information table is updated first, and then the initialization volume and metadata structure are updated; if the source data is a data change record, the log volume group and metadata structure are directly updated; Determine the update module, including: A receiving unit, used for receiving source data from the production end; A judgment unit, used to judge the type of source data, if the source data is configuration information, bitmap data and synchronization data without holes, then execute the basic configuration and synchronization data storage steps; if the source data is a data change record, then execute the change data storage step, wherein the data change record includes a unique identification code, a timestamp, an offset, a data length and a data content; The basic configuration and synchronization data storage unit is used to generate a configuration information table record according to the received configuration information and add it to the configuration information table, and then generate the data volume and metadata of the target virtual machine according to the generated record; for bitmap data, first initialize the metadata according to the bitmap data, and then write the absolute path of the bitmap data into the target virtual machine record in the configuration information table; for synchronization data without holes, first convert the synchronization data without holes and the bitmap data to obtain the initialization synchronization data, and then write the initialization synchronization data into the data volume; The change data storage unit is used to find the metadata and data volume of the virtual machine according to the unique identification code in the data change record, write the data content in the record into the log volume group, and then update the write operation log, update the mirror table, update the log metadata file and update the data block change log.
6. The KVM data storage and processing system according to claim 5, characterized in that: In the metadata structure creation module, the log metadata file is a double-linked list structure, and each log metadata file record includes a head pointer head and a tail pointer last, the head pointer head points to the data block version that was changed for the first time and finds the tail pointer last, and the tail pointer last points to the data block version that was changed most recently; The mirror table includes a data block flag flag and a log metadata offset log_meta, wherein the log metadata offset log_meta is the offset of the log metadata of the data block recorded in the log metadata file, and the log metadata offset log_meta stores the address of the head pointer head.
7. The KVM data storage and processing system according to claim 5, characterized in that: In the data volume creation module, the log volume group includes multiple log volumes, and each log volume corresponds to log volume information in a metadata structure; In the metadata structure creation module, the log metadata file stores the storage location of all data in the log volume group; the log volume information includes a data block offset and a data block timestamp, wherein the data block offset is the offset of each data block stored in the log volume in the initialization volume, and the data block timestamp is the timestamp of the last data block stored in the log volume. The system further comprises: The write initial volume module is used to find the old log volume LV0 and its log volume information in the log volume group, traverse the data block offset offset of the log volume information in the log volume LV0, and then overwrite all data block contents in the log volume LV0 into the initialization volume according to the data block offset offset; The bitmap data updating module is used to search the configuration information table according to the unique identification code of the target virtual machine, obtain the bitmap data of the target virtual machine, and then synchronously traverse the initialization volume and the bitmap data to update the bitmap data; The module for deleting the log volume and its log volume information is used to delete the log volume LV0 and its corresponding log volume information after writing all the data of the log volume LV0 into the initialization volume; The metadata modification module is used to modify the image table, log metadata file, write operation log and data block change log.
8. A method for recovering continuous data points, characterized in that: The KVM data storage processing method according to any one of claims 1 to 4 further comprises the steps of: A recovery volume creation step is to receive a recovery instruction and create a recovery volume; Determine the execution step, determine whether the recovery time point M0 is greater than or equal to the latest time point H0, if so, execute the latest time point recovery step, if not, execute the arbitrary time point recovery step; The latest point-in-time recovery step traverses the mirror table to determine whether the latest content of the data block is in the log volume group or the initialization volume, and then reads the data from the log volume group or the initial volume and writes it to the recovery volume, wherein the log volume group is read through the log metadata file; The recovery step at any time point is to obtain the bitmap at time M0 in the data block change log, and then traverse the bitmap to determine whether the latest content of the data block is in the log volume group or the initialization volume, and then read the data from the log volume group or the initial volume and write it to the recovery volume. The reading of the log volume group is done by combining the mirror table and the log metadata file positioning, and the reading of the initialization volume is done by positioning the mirror table; In the step-by-step rollback procedure, when the internal data of the virtual machine after restoration at the latest time point or any time point is not in a logically consistent state, the data block position before the recovery time point M0 is located by writing the operation log and the mirror table, and then the log metadata file is traced back to determine the last valid change content of each data block, and then the corresponding data is read from the initialization volume or log volume according to the last valid change content, and overwritten and written into the recovery volume.
Citation Information
Patent Citations
Continuous data protection method
CN104866435A
Backup method and system based on continuous writing CDP, storage medium and recovery method
CN114461456A