A data repair method, device and medium based on a virtual disk
The method addresses data loss in qcow2 virtual disks by traversing and replacing lost data addresses, ensuring data recovery during system failures or anomalies.
Patent Information
- Application Number
- CN202210461441.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-28
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-04-28
AI Technical Summary
In the prior art, when the system is powered off or the storage device is abnormal, the data in the secondary index table is easily lost, resulting in the data being unable to be repaired.
By obtaining the cluster size, the initial address of the L1 table and the L1 table size, determine the L1 table entry and the L2 table entry, traverse the virtual disk data, judge and replace the damaged user data address, and use the index number of the L2 table and the index number of the L2 table entry for data repair.
It realizes accurate acquisition of user data addresses and repairing corrupt data in the event of system abnormalities, avoiding data loss and improving the efficiency and accuracy of data repair.
Smart Images

Figure CN114840358B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of virtualization technology, and particularly to a data repair method, apparatus and medium based on a virtual disk. Background Art
[0002] In current virtualization applications, a mainstream virtual disk is in the qcow2 format. The qcow2 virtual disk splits data into clusters, and then organizes these clusters with a secondary index table. The first level is called the L1 Table (L1 Table), and each entry is used to store the address index of the L2 Table. The second level is called the L2 Table (L2 Table), and each entry is used to store the address of the user data (Data Cluster). This secondary index table can be called the metadata of qcow2. Among them, the starting address of the L1 Table is located in the header of the qcow2 virtual disk, and the L1 Table can be found through this starting address; the starting address of the L2 Table is stored in the L1 Table, and the L2 Table can be found through the content of the L1 Table; the address of the Data Cluster is stored in the L2 Table, and the Data Cluster can be found using the content of the L2 Table.
[0003] During the addressing process of the qcow2 virtual disk, in order to improve the efficiency of data search, the L2 cache (Cache) in the memory is often used to record the Data Clusters that have been accessed. When the system power is off or the storage device is abnormal, it may cause the L2 Cache not to be flushed to the virtual disk, resulting in partial loss of data in the secondary index table of the qcow2 virtual disk, that is, the Data Cluster address is damaged. For this problem, the currently commonly used method is to verify the data consistency. When the data consistency fails the verification, it means that the data has been lost.
[0004] The currently used method for verifying data consistency can only determine whether the data has been lost, and directly loses the data when the data consistency verification fails, and cannot repair the lost data.
[0005] Therefore, those skilled in the art now urgently need a data repair method based on a virtual disk to solve the problem that the lost data cannot be repaired currently. Summary of the Invention
[0006] The purpose of this application is to provide a data repair method, apparatus and medium based on a virtual disk to solve the problem that the lost data cannot be repaired currently.
[0007] To solve the above technical problems, the present application provides a data repair method based on a virtual disk, including:
[0008] Obtain the size of the cluster, the initial address of the L1 table, and the size of the L1 table;
[0009] Determine each L1 table entry according to the initial address and size of the L1 table, determine each L2 table entry according to the L1 table entry, and obtain each user data address according to each L2 table entry as the first result;
[0010] Traverse the data stored in the virtual disk with the size of the cluster as the offset to obtain each user data address as the second result;
[0011] When each user data address in the first result is damaged, obtain the damaged user data address, as well as the index number of the corresponding L2 table and the index number of the L2 table entry, as the third result;
[0012] Compare the first result with the second result, remove the same user data addresses, and use the remaining user data addresses in the second result as the fourth result;
[0013] Replace the damaged user data address according to the third result and the fourth result.
[0014] Preferably, determining whether each user data address in the first result is damaged includes:
[0015] Respectively determine whether each user data address can be divided evenly by the size of the cluster. If not, the current user data address is damaged;
[0016] If so, determine whether there is data in the user data address. If there is no data, it is determined that the current user data address is damaged.
[0017] Preferably, the L1 table entry and the L2 table entry are N bits in size. Traversing the data stored in the virtual disk to obtain each user data address as the second result includes:
[0018] Skip the header of the virtual disk, obtain N bits of data starting from the current address as the data to be judged, and obtain N bits of data starting from the current address as the data to be judged every time an offset is made;
[0019] Judge whether the data to be judged is equal to the qcow2 keyword, the refcount Table keyword, the refcount Block keyword, or the L1 table keyword. If not, use the data to be judged as the user data address;
[0020] Use all user data addresses as the second result.
[0021] Preferably, traversing the data stored in the virtual disk to obtain each user data address as the second result includes:
[0022] Opening the virtual disk in read-only mode and traversing the data stored therein to obtain each user data address as the second result.
[0023] Preferably, obtaining the size of the cluster, the initial address of the L1 table, and the size of the L1 table includes:
[0024] Obtaining the header of the virtual disk and parsing the header to obtain the size of the cluster, the initial address of the L1 table, and the size of the L1 table.
[0025] Preferably, the first result, the second result, the third result, and the fourth result are stored in different databases.
[0026] Preferably, after determining whether each user data address is damaged and taking the damaged user data addresses as the third result, it further includes:
[0027] Returning a prompt message; wherein, the prompt message includes the third result.
[0028] To solve the above technical problems, the present application further provides a data repair device based on a virtual disk, including:
[0029] An acquisition module for obtaining the size of the cluster, the initial address of the L1 table, and the size of the L1 table;
[0030] A first result determination module for determining each L1 table entry according to the initial address of the L1 table and the size of the L1 table, determining each L2 table entry according to the L1 table entry, and traversing each L2 table entry to obtain each user data address as the first result;
[0031] A second result determination module for traversing the data stored in the virtual disk with the size of the cluster as an offset to obtain each user data address as the second result;
[0032] A third result determination module for determining whether each user data address in the first result is damaged, obtaining the index number of the L2 table and the index number of the L2 table entry corresponding to the damaged user data address, and taking each damaged user data address, as well as the index number of the L2 table and the index number of the L2 table entry corresponding thereto, as the third result;
[0033] A fourth result determination module for comparing the first result and the second result, removing the same user data addresses, and taking the remaining user data addresses in the second result as the fourth result; wherein, the number of user data addresses in the fourth result is the same as the number of user data addresses in the third result;
[0034] A repair module, configured to replace the damaged user data address according to the third result and the fourth result.
[0035] Preferably, it further includes:
[0036] A prompt module, configured to return a prompt message; wherein, the prompt message includes the third result.
[0037] To solve the above technical problem, the present application further provides a data repair device based on a virtual disk, including:
[0038] A memory, configured to store a computer program;
[0039] A processor, configured to implement the steps of the data repair method based on a virtual disk as described above when executing the computer program.
[0040] To solve the above technical problem, the present application further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the data repair method based on a virtual disk as described above are implemented.
[0041] The data repair method based on a virtual disk provided by the present application first obtains all user data addresses obtained by the addressing method based on L1 to L2 by traversing the secondary index table as the first result; then traverses the data stored in the virtual disk by offset to obtain the actual storage addresses of each user data in the virtual disk. The user data addresses obtained by this method will not result in data loss due to system power-off or storage device anomalies, that is, traverses to obtain accurate user data addresses as the second result; then, determines whether each user data address in the secondary index table is damaged, finds the damaged user data address, and obtains the index number of the L2 table where it is located and the index number of the L2 table entry where it is located as the third result; then, based on the comparison between the first result and the second result, determines the accurate user data address that needs to be replaced as the fourth result; according to the L2 table and the index numbers of the L2 table entries in the third result, finds the damaged user data address and replaces it with the accurate user data address in the fourth result, thus realizing data repair. It solves the problem that currently it can only determine whether data is lost but cannot be repaired.
[0042] The data repair device based on a virtual disk and the computer-readable storage medium provided by the present application correspond to the above method and have the same effect. Description of the Drawings
[0043] To more clearly illustrate the embodiments of the present application, the following will briefly introduce the accompanying drawings required for the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0044] Figure 1 Flowchart of a data repair method based on a virtual disk provided by the present invention;
[0045] Figure 2 Schematic diagram of a secondary index with data loss;
[0046] Figure 3 For the present invention Figure 2 Schematic diagram of intermediate results for data repair;
[0047] Figure 4 For the present invention Figure 2 Schematic diagram of the secondary index after data repair;
[0048] Figure 5 Structural diagram of a data repair device based on a virtual disk provided by the present invention;
[0049] Figure 6 Structural diagram of another data repair device based on a virtual disk provided by the present invention. Detailed implementation manners
[0050] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present application.
[0051] The core of the present application is to provide a data repair method, device and medium based on a virtual disk.
[0052] To enable those skilled in the art to better understand the solution of the present application, the following will further elaborate on the present application with reference to the accompanying drawings and specific implementation manners.
[0053] Virtual disk technology is a technology that virtualizes one or more disks through memory. Relying on the advantage that the speed of memory is much faster than that of a hard disk, it speeds up the data exchange speed of the disk, thereby improving the running speed of the computer.
[0054] That is to say, a terminal (computer, server, workstation) can virtualize multiple disks in the memory and store data in such virtual disks, which can obtain a faster data exchange speed than storing data in a hard disk.
[0055] However, it is easy to understand that the data in the memory used by the terminal will be lost after a power failure, or when the storage device used as the terminal memory is abnormal, it may also cause the loss of data in the virtual disk.
[0056] Currently, it is possible to detect whether data is lost by verifying the data consistency. However, this method has drawbacks. When the data consistency check fails, in order to avoid the influence of incorrect data, the entire data is directly lost and the lost data cannot be repaired. Therefore, to solve the above problems, as Figure 1 shown, the present application provides a data repair method based on a virtual disk, including:
[0057] S11: Obtain the size of the cluster, the initial address of the L1 table, and the size of the L1 table.
[0058] The size of the cluster represents the capacity of data that a cluster can store, which is set when the virtual disk is created; the size of the L1 table is also set when the virtual disk is created, representing the capacity of the L1 table. L1 can occupy only one cluster or multiple consecutive clusters. Usually, the size of the L1 table is adapted to the size of the virtual disk. That is, when the size of the cluster remains unchanged, the larger the size of the virtual disk, the larger the size of the L1 table should be; while the L2 table fixedly occupies one cluster, and different L2 tables do not need to be placed on adjacent clusters.
[0059] In addition, since the size of the cluster, the initial address of the L1 table, and the size of the L1 table are all preset when the virtual disk is created, the setting information can be stored and obtained when needed; and since the above information will be saved in the header of the virtual disk when the virtual disk is created, it can also be obtained by parsing the header of the virtual disk. The present application does not limit this, but considering the ease of implementation of the method for obtaining the size of the cluster, the initial address of the L1 table, and the size of the L1 table, it is preferably obtained by parsing the header of the virtual disk. Correspondingly, the specific implementation method of this preferred solution includes:
[0060] Obtain the header of the virtual disk and parse the header to obtain the size of the cluster, the initial address of the L1 table, and the size of the L1 table.
[0061] S12: Determine each L1 table entry according to the initial address and size of the L1 table, determine each L2 table entry according to the L1 table entry, and obtain each user data address as the first result according to each L2 table entry.
[0062] As described above, based on the initial address of the L1 table, the L1 table can be found; and based on the size of the L1 table, it is possible to determine how many L1 table entries the L1 table contains, and then determine the address of each L1 table entry; after obtaining the address of each L1 table entry, the index address of the L2 table stored in the L1 table entry can be read, and then each L2 table can be determined; finally, based on each L2 table, the user data address stored therein is determined, realizing the address retrieval according to the L1-L2-Data Cluster to obtain all the user data addresses saved in the secondary index table.
[0063] S13: Traverse the data stored in the virtual disk with the size of the cluster as the offset to obtain each user data address as the second result.
[0064] The purpose of this step is to obtain accurate user data addresses. Since when the L2 Cache cannot be flushed to the virtual disk due to system power failure or storage device abnormality as described above, what is lost is the user data address stored in the secondary index table, that is, the user data addresses in each L2 table entry in the secondary index table will be damaged.
[0065] Therefore, by traversing and offsetting the data stored in the virtual disk with the cluster size as the offset, the accurate addresses of each user data address can be obtained for use in subsequent data repair steps.
[0066] S14: When the user data addresses in the first result are damaged, obtain the damaged user data addresses, as well as the index number of the corresponding L2 table and the index number of the L2 table entry, as the third result.
[0067] When the user data addresses stored in the secondary index table are damaged, it is necessary to first find out the damaged user data addresses, that is, the damaged data blocks in each L2 table. The specific method for determining whether a data block is damaged is not limited in this application. It can be by comparing the first result and the second result, and the part in the first result that is different from the second result is the user data address where the data is damaged; it can also be by verifying the data consistency to determine whether the data is damaged. However, it should be noted that currently, the data consistency check defaults to losing the entire data when the check fails. In this solution, the damaged data is also used in subsequent data patching steps. Therefore, if the method of verifying data consistency is used to determine whether data loss occurs, care should be taken not to lose the damaged data; other methods can also be used to determine whether data is damaged, as shown in a preferred solution provided in subsequent embodiments of this application, which will not be elaborated here.
[0068] After determining the addresses of the damaged user data (i.e., the damaged data blocks in the secondary index table), it is also necessary to obtain the index number of the L2 table corresponding to the damaged data block and the index number of the L2 table entry, so as to find the position of the damaged data block in the secondary index table, so that accurate data can be used to replace the corresponding position during subsequent data repair.
[0069] S15: Compare the first result and the second result, remove the same user data addresses, and use the remaining user data addresses in the second result as the fourth result.
[0070] As can be seen from the above, the secondary index table stores the addresses of all user data saved by the virtual disk. Therefore, the number of user data addresses included in the first result obtained through the L1-L2-Data Cluster retrieval method should be the same as that in the second result obtained by traversing the virtual disk data. And if there is no damage, the first result should be equal to the second result. Therefore, by comparing the first result and the second result and removing the same parts of the two, the remaining in the first result are the addresses of the damaged user data, and the remaining in the second result are the accurate user data addresses corresponding to the damaged data blocks.
[0071] S16: Replace the damaged user data addresses according to the third result and the fourth result.
[0072] The following further illustrates a data repair method based on a virtual disk provided by the present application, which is described in combination with examples:
[0073] As Figure 2 shown, Figure 2 An exemplary secondary index table with data loss is shown. Each piece of data framed by a solid line box or a dotted line box is a user data address; among them, the data framed by a dotted line box is the damaged data block, that is, there are damaged data blocks: 24AAAA, 6EE0BBBB, and 1CCCCC.
[0074] According to the above steps, obtain the first result 21, the second result 22, the third result 23, and the fourth result 24 corresponding to Figure 2 , specifically as Figure 3 shown, including: the first result 21, the second result 22, the third result 23, and the fourth result 24.
[0075] Figure 3 The first result 21 shown in only shows all the user data addresses of the damaged L2 table. There is no data damage in other L2 tables in this example, so they are not shown. However, it is easy to understand that when actually implementing this method, the first result 21 should have data user addresses of other L2 tables that are not shown. The same is true for the second result 22, the third result 23, and the fourth result 24.
[0076] The second result 22 contains all the user data addresses obtained by traversing the data stored in the virtual disk with an offset. By comparing the user data addresses in the first result 21 and the second result 22, the user data addresses shown in the third result 23 and the fourth result 24 can be obtained. Among them, the user data addresses in the third result 23 are damaged, and the user data addresses in the fourth result 24 are accurate.
[0077] In addition, the third result 23 should also include the index number of the L2 table and the index number of the L2 table entry of the damaged user data address. For example, for the data block 24AAAA, it is the 3rd table entry of the L2 table with the index number of 100000; for the data block 6EE0BBBB, it is the 3rd table entry of the L2 table with the index number of F40000; for the data block 1CCCCC, it is the 2nd table entry of the L2 table with the index number of 180000.
[0078] For the method of replacing the user data addresses in the third result 23 with the user data addresses in the fourth result 24 respectively, taking the repair of the data block 24AAAA as an example, the specific replacement process is as follows:
[0079] Determine the data block in the third result 23 that is closest to the damaged user data address and the data block 240000, as Figure 3 shown, which is 24AAAA; as can be seen from the above, the data block 24AAAA is the 3rd table entry of the L2 table with the index number of 100000. Therefore, replacing the 3rd table entry of the L2 table with the index number of 100000 in the secondary index table with the data block 240000 can achieve the repair of the data block 24AAAA.
[0080] The repair of other data blocks in the third result 23 is the same, which will not be elaborated here. Finally, the repaired disk data as shown in Figure 4 is obtained.
[0081] A data repair method based on a virtual disk provided by the present application obtains accurate user data addresses through offset traversal, and then replaces the obtained accurate user data addresses according to the positions of the damaged user data addresses in the index table, thereby completing the repair of the lost data, so that the metadata in the virtual disk can be ensured not to be lost when the system has an abnormal power-off or a storage failure.
[0082] As can be seen from the above, there are many methods for judging whether the user data addresses in the secondary index table are damaged, and the present application does not limit the specific implementation manner. For this embodiment, a preferred implementation scheme is provided, including:
[0083] Judge whether each user data address can be divided evenly by the size of the cluster. If not, the current user data address is damaged.
[0084] If so, judge whether there is data in the user data address. If there is no data, judge that the current user data address is damaged.
[0085] That is, when judging whether each user data address is damaged, it can be judged whether the current user data address is an integer multiple of the cluster size. If not, it means that the user data address is damaged. If so, enter the next judgment step.
[0086] If the current user data address can be divided evenly by the cluster size, it is also necessary to judge whether there is data stored at the position pointed to by the user data address. If there is data stored, it means that the user data address is not damaged. If there is no data stored, it means that the user data address is damaged.
[0087] Using the preferred solution provided in this embodiment to judge whether the user data address is damaged is simple to implement. Compared with the method of verifying data consistency, the entire data will not be lost when it is judged that the user data address is damaged; and compared with the method of obtaining by comparing the above first result and second result, since this method judges each user data address separately, when it is judged that there is damage, the index number of the L2 table and the index number of the L2 table entry corresponding to the current user data address can be directly saved to obtain the third result. However, the method of comparing the first result and the second result can only obtain the damaged user data address, but for the corresponding index numbers of the L2 table and the L2 table entry, it is necessary to obtain them again. In summary, a preferred solution for judging whether the user data address is damaged provided in this embodiment is simple to implement and will not lose the entire data, further improving the efficiency of a data repair method provided in this application.
[0088] Regarding traversing the data stored in the virtual disk to obtain the second result, this embodiment provides a specific preferred solution, including:
[0089] Skip the header of the virtual disk, obtain N bits of data starting from the current address as the data to be judged, and obtain N bits of data starting from the current address as the data to be judged every time an offset is made.
[0090] Judge whether the data to be judged is equal to the qcow2 keyword, refcount Table keyword, refcount Block keyword, or L1 table keyword. If not, use the data to be judged as the user data address.
[0091] Take all user data addresses as the second result.
[0092] Considering that in practical applications, with the existing disk capacity, the sizes of the L1 table entries and L2 table entries being 64 bits can fully meet the addressing requirements, so the value of N is preferably 64 here.
[0093] N is 64, that is, the number of bits of the data to be judged is 64 bits. Therefore, if the subsequent qcow2 keyword, refcount Table keyword, refcount Block keyword, or L1 table keyword is less than 64 bits, it can be supplemented with 0s to 64 bits for subsequent comparison.
[0094] Among them, since the qcow2 file supports creating disk snapshots (snapshots) inside the file, it is necessary to record the usage of each cluster, that is, to record the number of times each cluster is referenced. In this way, when a certain snapshot is deleted, it can be known which clusters can be released for continued use and which clusters are still in use, and the host cannot modify them actively. To maintain the record of the reference count of the clusters, a two-level index table is also used in the qcow2 file to record the reference count of all clusters. The first level of this two-level index table is called the recfound Table, and the second level is called the refcount Block, and its structure is similar to the above-mentioned two-level index table of L1-L2.
[0095] In addition, since the data stored in the above-mentioned qcow2 virtual disk contains the qcow2 keyword OX514649FB, if the data to be judged is equal to the above-mentioned qcow2 keyword, it means that the data to be judged is not the user data address.
[0096] Similarly, it is also necessary to compare whether the data to be judged is the L1 table keyword 0X000C0000. If the data to be judged is equal to the L1 table keyword, it means that the data to be judged is not the user data address.
[0097] The size of the refcount Table is variable and needs to occupy continuous clusters. In the case of the same cluster size, the larger the capacity of the virtual machine disk, the larger the size of the refcount Table, similar to the L1 Table. Each refcount Block occupies one cluster, and different refcount Blocks do not need to occupy continuous clusters. In the refcount Block, the size of each entry can be set, not necessarily 64 bits similar to the L2 Table, because each entry is just a simple count and there is no need to use 64 bits, which is a waste of space.
[0098] Therefore, after supplementing the refcount Table keyword and the refcount Block keyword to 64 bits, they should be 0X00040000 and 0X00080000. Compare the data to be judged with them. If they are the same, it means that the data to be judged is not the user data address.
[0099] That is, if the data to be judged is the same as any one of the above four keywords, it means that the data to be judged is not the user data address; if the data to be judged is different from all of the above four keywords, it means that the data to be judged is the user data address.
[0100] A preferred solution provided in this embodiment skips the data of the next N bits after the current address of the header as the data to be judged under the premise that the sizes of the L1 table entry and the L2 table entry are N bits. Since the data to be judged is only a kind of data that may be the user data address, it also includes keywords such as the qcow2 keyword, the refcount Table keyword, the refcount Block keyword, and the L1 table keyword. Therefore, this embodiment also compares the data to be judged with the qcow2 keyword, the refcount Table keyword, the refcount Block keyword, or the L1 table keyword. If the data to be judged is not consistent with any of the above keywords, it means that the data to be judged is the user data address. Through the screening of the four keywords, the obtained user data address is more accurate, which further ensures that the data repair method provided in this application can correctly repair the lost data.
[0101] As can be seen from the above embodiments, in order to obtain an accurate user data address, this application adopts the method of offset traversing the data stored in the virtual disk. Therefore, operations such as directly reading the data in the virtual disk are performed. However, in the process of directly processing the data in the virtual disk, the situation of data modification may occur, thus affecting the accuracy of the data. Therefore, to solve the above problems, this embodiment provides a preferred implementation scheme. The above traversing of the data stored in the virtual disk to obtain each user data address as the second result is specifically as follows:
[0102] Open the virtual disk in read-only mode, traverse the data stored therein, and obtain each user data address as the second result.
[0103] Opening in read-only mode can only perform read operations on the data, which can meet the requirement of the above offset traversal to obtain an accurate user data address. At the same time, the read-only mode cannot perform operations such as modifying and deleting the data, and will not affect the data stored in the virtual disk, further ensuring the accuracy of the obtained user data address.
[0104] As described above, the present application repairs the lost data by using the obtained first result, second result, third result, and fourth result, and the data stored in the above four results are all different. To avoid data chaos, this embodiment provides a preferred implementation scheme, including: storing the first result, second result, third result, and fourth result in different databases.
[0105] Similarly, as described above, this method determines whether the user data addresses stored in each L2 entry in the secondary index table are damaged, and uses the damaged user data addresses, the index number of their L2 table, and the index number of the L2 entry as the third result for subsequent data repair.
[0106] Therefore, when there is a user data address in the third result, it indicates that data loss has occurred in the virtual disk. At this time, this embodiment also provides another preferred implementation scheme, and the above method further includes: returning a prompt message; where the prompt message includes the third result.
[0107] Since data loss is usually caused by system power-off and storage device anomalies, this embodiment returns a prompt message when data loss occurs to inform relevant personnel to promptly check the above problems, further ensuring the stable operation of the system where the virtual disk is located. In addition, the prompt message returned by this embodiment also includes the third result, enabling relevant personnel to know the specific location of the data block with data loss in the secondary index table, and further enabling manual data repair to further ensure the accuracy of the data stored in the virtual disk.
[0108] In the above embodiment, a data repair method based on a virtual disk is described in detail. The present application also provides an embodiment corresponding to a data repair device based on a virtual disk. It should be noted that the present application describes the embodiment of the device part from two perspectives, one is from the perspective of functional modules, and the other is from the perspective of hardware.
[0109] From the perspective of functional modules, as Figure 5 shown, this embodiment provides a data repair device based on a virtual disk, including:
[0110] An acquisition module 31, configured to acquire the size of the cluster, the initial address of the L1 table, and the size of the L1 table;
[0111] A first result determination module 32, configured to determine each L1 entry according to the initial address of the L1 table and the size of the L1 table, determine each L2 entry according to the L1 entry, and traverse each L2 entry to obtain each user data address as the first result;
[0112] The second result determination module 33 is configured to traverse the data stored in the virtual disk with the size of the cluster as an offset to obtain each user data address as the second result;
[0113] The third result determination module 34 is configured to determine whether each user data address in the first result is damaged, and obtain the index number of the L2 table and the index number of the L2 table entry corresponding to the damaged user data address, and use each damaged user data address, and its corresponding index number of the L2 table and the index number of the L2 table entry as the third result;
[0114] The fourth result determination module 35 is configured to compare the first result and the second result, remove the same user data addresses, and use the remaining user data addresses in the second result as the fourth result; wherein, the number of user data addresses in the fourth result is the same as the number of user data addresses in the third result;
[0115] The repair module 36 is configured to replace the damaged user data addresses according to the third result and the fourth result.
[0116] Preferably, it further includes:
[0117] The prompt module is configured to return a prompt message when data loss occurs.
[0118] Wherein, the prompt message includes the third result.
[0119] Since the embodiments of the apparatus part correspond to the embodiments of the method part, for the embodiments of the apparatus part, please refer to the description of the embodiments of the method part, and will not be elaborated here.
[0120] The data repair apparatus based on a virtual disk provided by the present application obtains accurate user data addresses through the second result determination module, and then determines the positions of the damaged user data addresses obtained by the third result determination module in the index table, and replaces the obtained accurate user data addresses through the repair module, thereby completing the repair of the lost data, so that the metadata in the virtual disk can be ensured not to be lost when the system has an abnormal power failure or a storage failure.
[0121] Figure 6 Shown in Figure 6 is a structural diagram of a data repair apparatus based on a virtual disk provided by another embodiment of the present application. As
[0122] shown, a data repair apparatus based on a virtual disk includes: a memory 40 for storing a computer program;
[0123] A data repair device based on a virtual disk provided in this embodiment may include, but is not limited to, a smart phone, a tablet computer, a notebook computer, a desktop computer, etc.
[0124] Among them, the processor 41 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 41 may be implemented in at least one hardware form of a digital signal processor (DSP), a field-programmable gate array (FPGA), and a programmable logic array (PLA). The processor 41 may also include a main processor and a coprocessor. The main processor is a processor used to process data in the wake state, also known as a central processing unit (CPU); the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, the processor 41 may be integrated with a graphics processing unit (GPU), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 41 may further include an artificial intelligence (AI) processor, and the AI processor is used to process computational operations related to machine learning.
[0125] The memory 40 may include one or more computer-readable storage media, and the computer-readable storage media may be non-transitory. The memory 40 may further include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash storage devices. In this embodiment, the memory 40 is at least used to store the following computer program 401. After the computer program is loaded and executed by the processor 41, it can implement the relevant steps of a data repair method based on a virtual disk disclosed in any of the foregoing embodiments. In addition, the resources stored in the memory 40 may further include an operating system 402 and data 403, etc., and the storage method may be temporary storage or permanent storage. Among them, the operating system 402 may include Windows, Unix, Linux, etc. The data 403 may include, but is not limited to, a data repair method based on a virtual disk, etc.
[0126] In some embodiments, a data repair device based on a virtual disk may further include a display screen 42, an input / output interface 43, a communication interface 44, a power supply 45, and a communication bus 46.
[0127] Those skilled in the art can understand, Figure 6The structure shown does not constitute a limitation on a data repair device based on a virtual disk, and may include more or fewer components than those shown in the figure.
[0128] A data repair device based on a virtual disk provided by an embodiment of the present application includes a memory and a processor. When the processor executes a program stored in the memory, the following method can be implemented: A data repair method based on a virtual disk.
[0129] The data repair device based on a virtual disk provided by the present application enables the processor to execute a program stored in the memory, so as to obtain an accurate user data address by traversing the data in the virtual disk through an offset, and then replace the accurate user data address according to the position of the damaged user data address in the index table, so as to repair the lost data, ensuring that the metadata in the virtual disk will not be lost when the system has an abnormal power outage or a storage failure.
[0130] Finally, the present application also provides an embodiment corresponding to a computer-readable storage medium. A computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, the steps recorded in the method embodiment as described above are implemented.
[0131] It can be understood that if the method in the above embodiment is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and executes all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.
[0132] The computer-readable storage medium provided by the present application, when the computer program stored therein is executed, can obtain an accurate user data address by traversing the data in the virtual disk through an offset, and then replace the data therein with the accurate user data address according to the position of the damaged user data address in the index table, thereby completing the repair of the lost data, ensuring that the metadata in the virtual disk will not be lost when the system has an abnormal power outage or a storage failure.
[0133] The above has introduced in detail a data repair method, device, and medium based on a virtual disk provided by this application. Each embodiment in the specification is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. For the same or similar parts among the embodiments, reference can be made to each other. For the device disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple. For the relevant parts, reference can be made to the description in the method section. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the principle of this application, several improvements and modifications can still be made to this application, and these improvements and modifications also fall within the protection scope of the claims of this application.
[0134] It should also be noted that in this specification, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including", or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, article, or device including a series of elements not only includes those elements but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article, or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, article, or device including the said element.
Claims
1. A data repair method based on a virtual disk, characterized in that, including: obtaining the size of the cluster, the initial address of the L1 table, and the size of the L1 table; determining each L1 table entry according to the initial address and the size of the L1 table, determining each L2 table entry according to the L1 table entry, and obtaining each user data address according to each L2 table entry as a first result; traversing the data stored in the virtual disk with the size of the cluster as an offset to obtain each user data address as a second result; when each user data address in the first result is damaged, obtaining the damaged user data address, as well as the index number of the corresponding L2 table and the index number of the L2 table entry as a third result; comparing the first result with the second result, removing the same user data addresses, and taking the remaining user data addresses in the second result as a fourth result; replacing the damaged user data address according to the third result and the fourth result.
2. The data repair method based on a virtual disk according to claim 1, wherein Determining whether each user data address in the first result is damaged includes: respectively determining whether each user data address can be divided evenly by the size of the cluster, if not, then the current user data address is damaged; if so, determining whether there is data in the user data address, if there is no data, then determining that the current user data address is damaged.
3. The data repair method based on a virtual disk according to claim 1, wherein The size of the L1 table entry and the L2 table entry is N bits. The traversing the data stored in the virtual disk to obtain each user data address as a second result includes: skipping the header of the virtual disk, obtaining N bits of data starting from the current address as the data to be judged, and obtaining N bits of data starting from the current address every time an offset is made as the data to be judged; judging whether the data to be judged is equal to the qcow2 keyword, the refcount Table keyword, the refcount Block keyword, or the L1 table keyword, if not, then taking the data to be judged as the user data address; taking all the user data addresses as the second result.
4. The data repair method based on a virtual disk according to claim 1, wherein The traversing the data stored in the virtual disk to obtain each user data address as a second result includes: opening the virtual disk in a read-only manner and traversing the data stored therein to obtain each user data address as a second result.
5. The data repair method based on a virtual disk according to claim 1, wherein The obtaining the size of the cluster, the initial address of the L1 table, and the size of the L1 table includes: obtaining the header of the virtual disk and parsing the header to obtain the size of the cluster, the initial address of the L1 table, and the size of the L1 table.
6. The data repair method based on a virtual disk according to claim 1, wherein The first result, the second result, the third result, and the fourth result are stored in different databases.
7. The data repair method based on a virtual disk according to any one of claims 1 to 6, characterized in that, After obtaining the damaged user data address as the third result when each user data address in the first result is damaged, it further includes: returning a prompt message; wherein, the prompt message includes the third result.
8. A data repair device based on a virtual disk, characterized in that including: an obtaining module, configured to obtain the size of the cluster, the initial address of the L1 table, and the size of the L1 table; The first result determination module is configured to determine each L1 table entry according to the initial address and the size of the L1 table, determine each L2 table entry according to the L1 table entry, and traverse each L2 table entry to obtain each user data address as the first result; The second result determination module is configured to traverse the data stored in the virtual disk with the size of the cluster as an offset to obtain each user data address as the second result; The third result determination module is configured to determine whether each user data address in the first result is damaged, obtain the index number of the L2 table and the index number of the L2 table entry corresponding to the damaged user data address, and use each damaged user data address, and its corresponding index number of the L2 table and the index number of the L2 table entry as the third result; The fourth result determination module is configured to compare the first result and the second result, remove the same user data addresses, and use the remaining user data addresses in the second result as the fourth result; wherein, the number of user data addresses in the fourth result is the same as the number of user data addresses in the third result; The repair module is configured to replace the damaged user data address according to the third result and the fourth result.
9. A data repair device based on a virtual disk, characterized in that, Comprising: A memory for storing computer programs; A processor, when executing the computer program, is configured to implement the steps of the virtual disk-based data repair method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, the steps of the virtual disk-based data repair method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Metadata recovery method and device and medium
CN107704208A
Addressing method of virtual disk format and computer readable storage medium
CN109388524A