Virtual machine recovery method and electronic device
By writing the data from the failed hard drive to the updated hard drive and updating the storage pool in the hyperconverged system, the problem of data loss caused by hard drive failure is solved, local data recovery and real-time protection are achieved, operating costs are reduced and system reliability is improved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-26
- Publication Date
- 2026-03-27
AI Technical Summary
In existing technologies, hyperconverged systems are prone to data loss in distributed storage when hard drives fail simultaneously across nodes. Furthermore, relying on third-party data backup platforms increases operating costs, makes it difficult to achieve real-time protection, and the backup process consumes a large amount of network resources, affecting performance.
By writing available and faulty data to an updated disk after detecting a faulty hard drive, and updating the storage pool of the distributed storage system based on the updated data, virtual machine recovery is performed using storage status and configuration status data, thus achieving local data recovery and avoiding dependence on external backup platforms.
It enables local data recovery without the need for an external backup platform, reduces operating costs, achieves high-granularity real-time data protection, reduces network resource consumption, and ensures the business continuity and reliability of the system.
Smart Images

Figure CN121433811B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of cloud computing, and in particular to a virtual machine recovery method and an electronic device. BACKGROUND
[0002] The distributed storage of the hyper-converged system can be a cluster composed of multiple distributed nodes, adopts a node copy redundancy mechanism, and virtual machine service data can be stored on different distributed nodes, so that a single node failure does not affect the normal operation of the service application. However, when data hard disks are damaged across nodes at the same time, the redundancy failure domain of the distributed storage node copy is exceeded, which can easily lead to loss of distributed storage data. In related technologies, a third-party data backup platform is connected to the hyper-converged system, periodic backup of virtual machine data is set, preventive protection is performed, and after the virtual machine fails, the virtual machine data is recovered through the backup platform.
[0003] However, connecting to the third-party data backup platform increases the operating cost, especially for data-intensive applications; the backup granularity is large, real-time protection is difficult to achieve, and there is a risk of data loss; the backup and recovery process occupies a large amount of network resources and storage resources, affecting the performance of other services of the hyper-converged platform; and the failure of the backup platform itself will make the backup data unavailable, making it difficult to cope with the situation of superimposed failures in actual applications. SUMMARY
[0004] In view of the above problems, the present application provides a virtual machine recovery method and an electronic device.
[0005] According to a first aspect of the present application, a virtual machine recovery method is provided, comprising: in response to detecting a fault hard disk mounted by an initial virtual machine, sequentially writing available data and fault data in the fault hard disk into an update hard disk as update available data and update fault data respectively; based on the update available data and the update fault data, updating a storage pool corresponding to multiple nodes in a distributed storage system to obtain an update storage pool; based on storage state data of the update storage pool and configuration state data of the initial virtual machine, recovering the initial virtual machine to obtain an intermediate virtual machine; and in a case where a data detection result and an application detection result of the intermediate virtual machine meet a preset condition, determining the intermediate virtual machine as a target virtual machine, so as to execute service processing by using the target virtual machine, wherein the preset condition includes a data integrity condition of the data detection result and an application program test condition of the application detection result.
[0006] The second aspect of the present application provides a virtual machine recovery device, comprising: a writing module configured to write available data and failure data in a failure hard disk into an update hard disk in sequence as update available data and update failure data, in response to detecting that the initial virtual machine is mounted on the failure hard disk; an updating module configured to update storage pools corresponding to a plurality of nodes in a distributed storage system based on the update available data and the update failure data, to obtain updated storage pools; a recovery module configured to recover the initial virtual machine based on storage state data of the updated storage pools and configuration state data of the initial virtual machine, to obtain an intermediate virtual machine; and a determination module configured to determine the intermediate virtual machine as a target virtual machine in a case where data detection results of the intermediate virtual machine and application detection results meet preset conditions, to execute business processing by using the target virtual machine, wherein the preset conditions include a data integrity condition of the data detection results and an application program test condition of the application detection results.
[0007] The third aspect of the present application provides an electronic device, comprising: one or more processors; a memory for storing one or more computer programs, wherein the one or more processors execute the one or more computer programs to implement the steps of the method.
[0008] The fourth aspect of the present application further provides a computer-readable storage medium having stored thereon a computer program or instructions, wherein the computer program or instructions are executed by a processor to implement the steps of the method.
[0009] The fifth aspect of the present application further provides a computer program product comprising a computer program or instructions, wherein the computer program or instructions are executed by a processor to implement the steps of the method. BRIEF DESCRIPTION OF DRAWINGS
[0010] The above and other objects, features and advantages of the present application will become more apparent from the following description when taken in conjunction with the accompanying drawings, in which:
[0011] Figure 1 An application scenario diagram of the virtual machine recovery method and the electronic device according to an embodiment of the present application is shown;
[0012] Figure 2 A flowchart of the virtual machine recovery method according to an embodiment of the present application is shown;
[0013] Figure 3A An example schematic diagram of a failure hard disk data reading process according to an embodiment of the present application is shown;
[0014] Figure 3B A flowchart of a method of writing available data in a failure hard disk into an update hard disk according to an embodiment of the present application is shown;
[0015] Figure 3C A flow chart of a method for writing and updating a hard disk with a failed sector address in a failed hard disk according to an embodiment of the present application is shown;
[0016] Figure 4A A flow chart of a method for writing and updating a failed sector with failed data according to an embodiment of the present application is shown;
[0017] Figure 4B A flow chart of a method for repairing a storage pool according to an embodiment of the present application is shown;
[0018] Figure 5 A block diagram of a virtual machine recovery apparatus according to an embodiment of the present application is shown;
[0019] Figure 6 A block diagram of an electronic device adapted to implement a virtual machine recovery method according to an embodiment of the present application is shown. DETAILED DESCRIPTION
[0020] Embodiments of the present application will be described herein below with reference to the accompanying drawings. It is to be understood, however, that the description is merely exemplary and is not intended to limit the scope of the present application. In the following detailed description of embodiments of the present application, numerous specific details are set forth in order to provide a thorough understanding of the embodiments. However, it will be apparent to one skilled in the art that one or more embodiments can be practiced without these specific details. In other instances, well-known structures and functions have not been described in detail in order to avoid obscuring aspects of the application.
[0021] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the present application. As used herein, the term "includes" and tautological equivalents thereof, means that the named feature, step, operation, and / or component is present, but not excluding the presence or addition of one or more other features, steps, operations, or components.
[0022] All terms used herein including technical and scientific terms have the same meanings as commonly understood by one of ordinary skill in the art unless otherwise defined herein. It should be noted that the terms used herein are merely specific embodiments proposed as examples in order to describe the present application and the present application is not limited thereto. Accordingly, it should be understood that the terms used herein have meanings consistent with those of the corresponding technical terms in the art.
[0023] In the case of using expressions similar to "at least one of A, B, and C", generally, it should be interpreted to include at least one of each item enumerated, unless the context clearly indicates otherwise. For example, "a system having at least one of A, B, and C" should be interpreted to include a system having at least one of A, a system having at least one of B, a system having at least one of C, a system having at least one of A and B, a system having at least one of A and C, a system having at least one of B and C, and / or a system having at least one of A, B, and C, etc.
[0024] The following describes and explains technical terms related to the technical solutions of the present application.
[0025] The hyper-converged system is an information technology infrastructure architecture deeply integrated with computing, storage and network. Through software-defined technology, the local storage of multiple general-purpose server nodes, such as solid state disks and hard disk drives, is aggregated into a unified distributed storage resource pool, and computing virtualization is deployed on each node to achieve hardware generalization, functional software, and centralized management, breaking the traditional computing, storage and network deployment mode.
[0026] The core of the hyper-converged system is distributed storage, which can directly provide storage resources for virtual machines without relying on external storage devices.
[0027] Hard disks are the basic components of distributed storage. In the case of a large amount of data read and write, frequent sector erasing will greatly increase the failure rate of hard disks. At the same time, hard disk firmware version defects are prone to cause batch hard disk failures, and the environment of the computer room with high temperature, dust and vibration is the inducement of hard disk damage. Hard disk damage will cause failure of distributed storage. The data blocks of the virtual machines of the hyper-converged system are scattered on each hard disk of the distributed storage. The hard disk damage in the super fault domain will cause multiple virtual machine data to be incomplete, the operating system and business application program to fail, and the virtual machine to be unavailable.
[0028] Non-destructive recovery after hard disk failure, adding data to the distributed storage to recover data, repairing damaged data storage pools, recovering virtual machine hard disk data, enabling virtual machines and business applications, and ensuring that business data is not lost are the urgent needs in current practical applications.
[0029] In some examples, the hyper-converged system interfaces with a third-party data backup platform, sets up periodic backup of virtual machine data, and performs preventive protection. After the virtual machine fails, the virtual machine data is recovered through the backup platform.
[0030] However, interfacing with a third-party data backup platform requires users to purchase additional backup servers and backup software, increasing the cost of additional equipment procurement, and the backup software is charged based on capacity. For a large amount of training model data in the artificial intelligence scenario, high-capacity authorization needs to be purchased, greatly increasing the operating cost. At the same time, the data backup granularity is large, and the third-party data backup platform cannot achieve real-time full backup data, but can only backup periodically according to time nodes, and some data cannot be backed up in time, causing full recovery.
[0031] During virtual machine backup and data recovery, the performance of the hyper-converged system is reduced, the data backup and recovery transmit data through the network, which occupies a large amount of network bandwidth, and the data read and write generate a large number of read and write tasks on the distributed storage, which has a great performance impact on the read and write of other business virtual machines; further, the failure of the third-party data backup platform will cause all backup data to be unavailable or lost, and at this time, if the super fault domain of the hard disk is damaged, there will be no available backup data for data recovery.
[0032] Therefore, embodiments of the present application provide a virtual machine recovery method and an electronic device, the method comprising: in response to detecting that the initial virtual machine is mounted on a fault hard disk, sequentially writing available data and fault data in the fault hard disk into an update hard disk as update available data and update fault data respectively; updating a storage pool corresponding to a plurality of nodes in a distributed storage system based on the update available data and the update fault data, to obtain an update storage pool; recovering the initial virtual machine based on storage state data of the update storage pool and configuration state data of the initial virtual machine, to obtain an intermediate virtual machine; and in a case where a data detection result and an application detection result of the intermediate virtual machine satisfy a preset condition, determining the intermediate virtual machine as a target virtual machine, so as to execute business processing by using the target virtual machine, wherein the preset condition comprises a data integrity condition of the data detection result and an application program test condition of the application detection result.
[0033] According to embodiments of the present application, by writing the available data and the fault data in the fault hard disk into the update hard disk, and updating the storage pool of the distributed storage system in real time based thereon, local data recovery without relying on an external backup platform is realized, the operating cost and the compatibility requirement for data-intensive applications are reduced. Through the multiple recovery strategies based on the storage state data and the configuration state data, high-granularity real-time data protection is realized, and the data loss risk caused by the large backup granularity of the traditional backup scheme is avoided. During the recovery process, the occupation of network resources and storage resources is minimized through localized data operation, and the performance of other businesses in the hyper-converged system is guaranteed. In addition, the storage pool update and virtual machine recovery based on the built-in redundancy mechanism of the distributed storage system eliminate the dependence on the data backup platform, effectively cope with the superimposed fault scenario through the system automatic recovery capability, and significantly improve the business continuity and system reliability.
[0034] Figure 1 An application scenario diagram of the virtual machine recovery method and the electronic device according to embodiments of the present application is shown.
[0035] As Figure 1As shown, the application scenario 100 according to the embodiment can include a server 101, a network 102, a failed hard disk 103, and a virtual machine 104. The network 102 is a medium to provide a communication link between the failed hard disk 103, the virtual machine 104, and the server 101. The failed hard disk 103 is mounted on the virtual machine 104. The network 102 can include various connection types, such as wired, wireless communication links, or fiber optic cables, and the like.
[0036] The server 101 can be a processor integrated in a hyper-converged system for data recovery of the virtual machine 104. The server 101 can provide both computing resources (e.g., running the virtual machine 104) and storage resources (whose local hard disk is part of a distributed storage pool), providing processing, memory, and the like resources for the recovery of the virtual machine 104.
[0037] For example, data migration, storage pool update, virtual machine recovery operation, can be performed by different functional modules (e.g., distributed storage management module and virtualization management module) running in the server 101 in coordination.
[0038] For example, the server 101 can communicate with other server nodes in the hyper-converged system through the network 102. The server 101 can detect a failed hard disk 103 on itself or another node in the hyper-converged system, manage the configuration of the virtual machine 104 that needs to be recovered. The server 101 can also coordinate writing data to other healthy hard disks ("update hard disks") in the distributed node cluster, responsible for performing life cycle operations (e.g., shutdown, rebuild, start) of the virtual machine recovery.
[0039] The failed hard disk 103 can be a failed component in the distributed storage pool relied on by the virtual machine 104. The server 101 detects the failed hard disk 103 and takes the failed hard disk 103 as a processing object to perform data extraction and classification. The virtual machine 104 internally runs a plurality of application programs and services.
[0040] It should be noted that the virtual machine recovery method provided by the embodiment of the application can generally be executed by the server 101. Accordingly, the virtual machine recovery apparatus provided by the embodiment of the application can generally be arranged in the server 101. The virtual machine recovery method provided by the embodiment of the application can also be executed by a server or a server cluster different from the server 101 and capable of communicating with the server 101. Accordingly, the virtual machine recovery apparatus provided by the embodiment of the application can also be arranged in a server or a server cluster different from the server 101 and capable of communicating with the server 101.
[0041] It should be understood that, Figure 1The number of servers, networks, failed hard disks, and virtual machines in the system is merely illustrative. Any number of servers, networks, failed hard disks, and virtual machines can be provided as needed for implementation.
[0042] Figure 2 A flowchart of a virtual machine recovery method according to an embodiment of the present application is shown.
[0043] As shown in Figure 2 The virtual machine recovery method of this embodiment includes operations S210-S240.
[0044] In operation S210, in response to detecting a failed hard disk on which an initial virtual machine is mounted, available data and failed data in the failed hard disk are sequentially written to an update hard disk as update available data and update failed data, respectively.
[0045] In an embodiment of the present application, the initial virtual machine can be an original virtual machine instance running in a hyper-converged system and requiring recovery due to a hard disk failure, containing a complete operating system, application programs, and business data. The failed hard disk can be a storage device that has physical or logical damage, resulting in data that cannot be normally accessed, including a hard disk drive (HDD) or a solid state drive (SSD). The available data can be complete, undamaged data blocks in the failed hard disk that can be read through normal or recovery means, having complete logical structure and readability. The failed data can be damaged data blocks in the failed hard disk that cannot be normally read due to physical damage, logical errors, or storage medium degradation.
[0046] In an embodiment of the present application, the update hard disk can be a local redundant storage device built into a distributed storage system for replacing the failed hard disk, having a healthy state and sufficient capacity for receiving and storing data extracted from the failed hard disk. Sequential writing can refer to a serialization operation process of transferring data from the failed hard disk to the update hard disk in a specific order and strategy. The update available data can be data blocks successfully extracted from the failed hard disk, the integrity of which is verified in the update hard disk, and can be directly used as valid data for storage pool reconstruction. The update failed data can be damaged data blocks and metadata information thereof identified from the failed hard disk, and the data stored in the update hard disk through writing or recovery, including damaged locations, damage degrees, and repair status markers.
[0047] For example, the degree of failure can be evaluated based on hard disk key parameter information (such as read error rate, remapping sector count) to develop a differentiated data extraction scheme. For available data, high-speed continuous reading is used to maximize extraction efficiency; for failed data, low-speed safe reading is used in combination with error correction codes to attempt to recover data in the failed hard disk to the update hard disk.
[0048] At operation S220, based on the updated available data and the updated failure data, the storage pool corresponding to the plurality of nodes in the distributed storage system is updated to obtain an updated storage pool.
[0049] In an embodiment of the present application, the distributed storage system can be a software-defined storage architecture for data dispersion storage composed of a plurality of storage nodes connected through a network. The plurality of nodes can be physical or virtual server units independently providing storage services in the distributed storage system. The storage pool can be an abstracted logical storage resource set in the distributed storage system, and the physical storage of the plurality of nodes can be integrated into a unified storage space. The updated storage pool can be a healthy storage pool with integrity and consistency restored by integrating the updated available data and the updated failure data information.
[0050] For example, the storage pool can be updated in a progressive rolling manner. The node with the highest data integrity in the storage pool is selected as a temporary master node, and the updated available data is synchronized to the master node. The storage pool data is divided into update areas according to shards, and each shard is sequentially synchronized from the master node to the slave node. After each node completes synchronization, local data verification is performed immediately, and the verification result is reported to the coordinator. After all nodes are synchronized, the global metadata version is updated to obtain the updated storage pool.
[0051] At operation S230, based on the storage state data of the updated storage pool and the configuration state data of the initial virtual machine, the initial virtual machine is recovered to obtain an intermediate virtual machine.
[0052] At operation S240, in a case where the data detection result and the application detection result of the intermediate virtual machine satisfy a preset condition, the intermediate virtual machine is determined as a target virtual machine to execute business processing using the target virtual machine. The preset condition includes a data integrity condition of the data detection result and an application program test condition of the application detection result.
[0053] In an embodiment of the present application, the storage state data can be real-time running state information of the repaired distributed storage pool, including data integrity indicators, performance indicators, availability states, and load conditions. The configuration state data can be a configuration snapshot of the initial virtual machine before the failure, containing virtual hardware configuration data, system configuration data, network setting data, and application dependency relationships. The intermediate virtual machine can be a transitional virtual machine instance generated during the recovery process, used to verify the recovery effect, and not yet put into business use.
[0054] In an embodiment of the present application, the data detection result can be an output result of verifying the integrity and consistency of the business data in the intermediate virtual machine. The application detection result can be a verification result of testing the running state of the business application program on the intermediate virtual machine. The preset condition can be a quantitative standard condition for the virtual machine to be put into business use.
[0055] For example, allocating the storage space required to restore the virtual machine from the update storage pool, creating a virtual machine hardware template based on configuration status data; deploying the operating system and basic operating environment, restoring system configuration and user data; performing integrity verification on business data, checking data consistency and correctness. Under the condition that the data is complete, installing and configuring the business application, testing the application's functional integrity, and determining that the virtual machine recovery is complete when both data integrity and test results meet the corresponding conditions.
[0056] In one feasible implementation, virtual machines can be recovered incrementally. Priority is given to restoring the system's core functions and critical data to ensure basic operational capabilities; system, middleware, and application functions are verified layer by layer, with each layer being restored only after successful verification; the recovery content is gradually improved based on the verification results, and recovery strategies and resource configurations are dynamically adjusted.
[0057] For example, data can be extracted from the faulty hard drive before writing the available data and faulty data from the faulty hard drive to the update hard drive.
[0058] Figure 3A A schematic diagram illustrating an example of a faulty hard disk data reading process according to an embodiment of the present invention is shown.
[0059] After detecting a faulty hard drive, the faulty hard drive is safely removed from the physical server to avoid further physical damage, and a full backup of the hard drive is performed.
[0060] like Figure 3A As shown, a high-speed serial hard drive cable can be used to connect the faulty hard drive 103. A data reading tool is then used to obtain the configuration information and minimum storage unit of the disk platter of the faulty hard drive 103, resulting in fault data 31. Based on the hard drive configuration data in the fault data 31, the faulty hard drive type 32 is confirmed. The data reading tool automatically analyzes and reads data based on the faulty hard drive type 32. If the faulty hard drive is a mechanical hard drive (HDD), the disk data reading tool first reads the entire disk sector 33 and reads the available data 34 on the sector. If a sector is unreadable, it is marked as a faulty sector, and the faulty sector address 35 is recorded for later insertion of the damaged sector location number into the replacement hard drive. If the faulty hard drive is a solid-state drive (SSD), the non-volatile storage medium flash memory chips 36 of the disk can be analyzed first, and the data on the flash memory chips 36 can be extracted. Locations where data cannot be read are marked as bad block addresses 37.
[0061] For example, the disk data extraction tool can first traverse the disk sectors in the minimum storage unit of the disk slice, and the sectors internally include an address area, a data area, a synchronization area, and an error correction code (ECC) area. The data area is read, and a normally readable data area is recorded, and a damaged data sector is recorded if the data area is read abnormally or cannot be accessed.
[0062] According to the embodiment of the application, by writing the available data and the fault data in the fault hard disk into the update hard disk, and updating the storage pool of the distributed storage system in real time based on this, local data recovery without relying on an external backup platform is realized, the operation cost and the compatibility requirement for data-intensive applications are reduced. Through the multiple recovery strategies based on the storage state data and the configuration state data, high-granularity real-time data protection is realized, and the data loss risk caused by the large backup granularity of the traditional backup scheme is avoided. In the recovery process, the occupation of network resources and storage resources is minimized through localized data operation, and the performance of other businesses in the hyper-converged system is ensured not to be affected. In addition, the storage pool update and the virtual machine recovery based on the built-in redundancy mechanism of the distributed storage system eliminate the dependence on the data backup platform, effectively cope with the superimposed failure scenario through the system automatic recovery capability, and significantly improve the business continuity and system reliability.
[0063] It can be understood that the above has described how to recover the virtual machine from the whole technical scheme, and the following will specifically describe how to write the data in the fault hard disk into the update hard disk.
[0064] According to the embodiment of the application, the available data and the fault data in the fault hard disk are sequentially written into the update hard disk as update available data and update fault data, including: based on the mapping relationship between the available address in the fault hard disk for storing the available data and the update available address in the update hard disk, the available data is written into the update hard disk as the update available data; based on the error correction code of the available data, the comparison result between the available data and the update available data is determined, and in the case that the comparison result meets a preset comparison condition, the fault data is written into the update hard disk as the update fault data.
[0065] In the embodiment of the application, the mapping relationship can be an exact correspondence between the logical address of the intact data block in the fault hard disk and the corresponding storage position in the update hard disk. The error correction code can be the check information attached to the data block, which is used to detect and correct errors generated in the data transmission or storage process. The preset comparison condition can be a quantitative standard for data integrity verification, including a bit error rate threshold, a checksum matching degree, and other measurable indicators.
[0066] For example, the available data in the failed hard disk can be scanned, an address mapping table between the available addresses and the updated available addresses is established, and a corresponding address mapping structure is created in the updated hard disk; the available data is read in the order of data blocks, and is written in the corresponding positions of the updated hard disk according to the mapping table, and the transmission progress and the verification result are recorded in real time; the error correction code of each written data block is verified, the position of the data block with a failed verification is recorded, and the data block is retried or marked; and the overall verification pass rate is counted, and when the pass rate reaches a threshold value, it is indicated that the available data is completed to be written in the updated hard disk.
[0067] Figure 3B A flowchart of a method for writing available data in a failed hard disk into an updated hard disk according to an embodiment of the present application is shown.
[0068] As shown in Figure 3B The method for writing available data in a failed hard disk into an updated hard disk can include operations S301-S306.
[0069] In operation S301, the available data and the error correction code are written into the updated hard disk. After the available data is read from the failed hard disk, the available data can be written into the updated hard disk at corresponding updated available addresses according to the available addresses of the available data; and the error correction code in the failed hard disk is written into the updated error correction code area of the updated hard disk to obtain updated error correction code.
[0070] In operation S302, the completion of writing state of the updated hard disk is sent to a data verification tool
[0071] In operation S303, the data verification tool is called to detect the data consistency. Whether the starting sector address and the ending sector address in the failed hard disk and the updated hard disk are consistent is compared, and is double-confirmed in combination with the error correction code. If yes, operation S304 is performed. If no, operation S305 is performed.
[0072] Operation S304 ends the detection.
[0073] Operation S305 reads the data again according to the inconsistent sector address from the failed hard disk according to the sector address.
[0074] Operation S306 writes the read data again into the updated sector address of the updated hard disk.
[0075] According to the embodiments of the present application, the integrity of the intact data recovery can be ensured by the accurate address mapping, the validity of the repair can be ensured based on the conditional judgment of the verification result, and the writing speed is improved while the data writing quality is ensured by the hierarchical processing strategy. It can be understood that the efficiency, reliability and integrity of the data recovery of the failed hard disk are realized by the systematic data classification, accurate address mapping and scientific verification mechanism, and the reliability of the data recovery is further improved.
[0076] According to the embodiment of the present application, the method for writing the fault sector address in the fault hard disk into the update hard disk comprises the following steps: writing the fault sector address in the fault data into a check area; writing the fault sector address in the check area into the update hard disk as an update fault sector address when the fault sector address in the check area is the same as the fault sector address in the fault hard disk; and rewriting the data in the update fault sector in the update hard disk based on the update fault sector address.
[0077] In the embodiment of the present application, the fault sector address can be a logical block address number identifying the physical location of the damaged sector, which is used to accurately locate the fault sector. The check area can be a special memory area for temporarily storing fault metadata, which is used to verify the accuracy and consistency of the fault address. The update fault sector address can be the fault location identifier written into the update hard disk after verification, which provides target positioning for subsequent data rewriting. The data rewriting can be an operation of writing the repaired update fault data into the update fault sector in the update hard disk based on error correction code, redundant copy or data reconstruction algorithm.
[0078] The fault sector address in the fault hard disk can be inserted into the corresponding sector address of the update hard disk to form a complete disk sector address chain, which facilitates subsequent recovery of the entire disk data.
[0079] Figure 3C A flowchart of the method for writing the fault sector address in the fault hard disk into the update hard disk according to the embodiment of the present application is shown.
[0080] As shown in Figure 3C , the method comprises operations S310-S360.
[0081] In operation S310, the fault sector address in the fault hard disk is read and written into a check area in the virtual machine recovery device. The check area can be a special memory area in the virtual machine recovery device for temporarily storing fault metadata.
[0082] To avoid inserting the damaged fault sector address into the wrong update hard disk location, causing errors in the available data, a program for secondary checking of the fault sector address can be added.
[0083] In operation S320, the fault sector address in the check area is compared with the fault sector address in the fault hard disk, and the secondary-confirmed fault sector address is transmitted to a confirmation area in the cache area.
[0084] In operation S330, the confirmed fault sector address is inserted into the corresponding update sector address of the update hard disk.
[0085] In operation S340, it is detected whether the full backup data of the update hard disk and the fault hard disk are consistent. If yes, operation S350 is performed, and if no, operation S360 is performed.
[0086] In operation S350, the end is detected.
[0087] In operation S360, the failure data is read from the failure hard disk again and written to the update hard disk.
[0088] It should be noted that the data consistency detection can be checked multiple times during the entire data reading process to ensure the consistency of the data in the update hard disk and the data in the failure hard disk.
[0089] For example, the data rewriting can be performed through sequential verification. The failure hard disk can be scanned to identify the logical block addresses of all damaged sectors, record the error type and severity of each failure sector, generate a failure address mapping table and store it in the verification area. Thus, through the first verification, the consistency of the failure sector addresses in the verification area and the failure sector addresses in the failure hard disk is compared; and the second verification is performed, the adjacent sectors of the failure sector addresses are sampled and read to confirm the boundary accuracy, and a signature confirmation file is generated after the verification passes; the failure sector addresses that pass the verification are written in batches to the special metadata area of the update hard disk, each failure sector address is allocated a rewriting priority and a strategy identifier, and a mapping relationship between the failure sector address and the rewriting strategy is established, the slightly damaged sectors can be processed preferentially, the standard error correction code is used for rewriting, the multiple retry strategy can be used for the moderately damaged sectors, and the data reconstruction scheme can be used for the severely damaged sectors.
[0090] According to the embodiments of the present application, through the intermediate verification link of the verification area, the error address information pollution of the update hard disk is effectively prevented, the address consistency verification mechanism can ensure the accuracy of the failure positioning, avoid the false repair of normal sectors, realize the quality control through the data rewriting process, and ensure the availability and integrity of the repaired data.
[0091] According to the embodiments of the present application, based on the update failure sector address, the data rewriting of the update failure sector in the update hard disk includes: classifying the update failure sector based on the damage degree of the failure data in the failure hard disk to obtain at least one classification result; in the case that the classification result represents that the damage degree of the failure data is less than a first damage threshold, the failure data is written to the update failure sector once; in the case that the classification result represents that the damage degree of the failure data is greater than or equal to the first damage threshold and less than a second damage threshold, the failure data is written to the update failure sector in batches; and in the case that the classification result represents that the damage degree of the failure data is greater than or equal to the second damage threshold, the failure data is backed up before being written to the update failure sector in batches.
[0092] In embodiments of the present application, the damage degree can be a sector data damage level quantification index based on a comprehensive evaluation of indicators such as read error rate, number of error correction code error correction failures, and read delay. The classification result can be a grouping result of dividing the faulty sectors into different processing priorities according to the damage degree, for differential rewriting strategy. The first damage threshold can be a critical value that distinguishes between mild damage and moderate damage, usually corresponding to a damage degree level that can be repaired at one time. The second damage threshold distinguishes between moderate damage and severe damage, usually corresponding to a damage degree level that needs to be repaired after backup.
[0093] Single write can be a direct write strategy adopted for mild damage sectors, suitable for low error rate and one-time repairable cases. Batched write can be a step-by-step write strategy for moderate damage sectors, which improves the success rate of repair through multiple attempts. Backup and write can be a strategy of trying to repair after backup for severely damaged sectors, to ensure that the original data is not further damaged.
[0094] For example, data rewriting of updated faulty sectors in a hard disk can be based on threshold values. First, damage degree quantification evaluation can be performed by multi-dimensional detection of each faulty sector, including: read error rate statistics, number of error correction code error correction requirements, average read delay time, adjacent sector influence evaluation, and weighted algorithm to calculate the comprehensive damage degree score (0-100 points). Dynamic threshold setting can set the first damage threshold to 30 points: below this value, single write can be used, the second damage threshold is set to 70 points: above this value, backup and write is required, and batched write strategy is used between 30-70 points. Further classification and differential write, for scores < 30 points, direct single write can be used, and verification is passed to mark completion; for 30 points ≤ score < 70 points, write in 3-5 batches, adjust parameters after each batch verification; for scores ≥ 70 points: backup to a safe area first, then try to write in 5-8 batches.
[0095] Figure 4A A method flowchart for writing fault data to update faulty sectors according to an embodiment of the present application is shown.
[0096] As shown in Figure 4A The method of writing fault data to update faulty sectors can include operations S41-S43.
[0097] In operation S41, sector scanning of the entire disk of the updated hard disk is performed to confirm the existence of faulty sector data with data damage.
[0098] In operation S42, a fault sector data analysis program is started. Verify the address, start address and end address, error correction code and other information of the faulty sector.
[0099] In operation S43, the data recovery tool is invoked to rewrite data in the failed sectors. By continuously scanning the failed sectors until it is confirmed that all the failed sectors have completed data rewriting, the data of the disk is complete, and the sector scanning of the hard disk data analysis tool on the hard disk is ended.
[0100] It can be understood that, through accurate classification according to damage degrees, targeted repair solutions are provided for different levels of damage, so that the repair success rate is improved. The single-write strategy realizes rapid repair for lightly damaged sectors, and reduces unnecessary retry times; the batch-write strategy gradually overcomes moderately damaged sectors through parameter optimization, and improves repair completion degree; and the backup-then-write strategy ensures that severely damaged sectors do not cause further data loss during the repair process.
[0101] According to the embodiments of the present application, the classification processing avoids using a uniform high-cost repair solution for all sectors, shortens the overall repair time, the priority management ensures that resources are preferentially allocated to sectors with high repair success rates and short time consumption, and the parallel processing mechanism allows corresponding repair strategies to be simultaneously used for sectors of different categories, thereby significantly improving the repair success rate and efficiency of the failed sectors under the premise of ensuring data security.
[0102] According to the embodiments of the present application, based on updating available data and updating failed data, the storage pools corresponding to the plurality of nodes in the distributed storage system are updated to obtain an updated storage pool, including: in the case that the updating available data and the updating failed data complete data rewriting in the updated hard disk, mounting the updated hard disk after data rewriting to the distributed storage system to obtain an intermediate storage system; performing data consistency detection on the storage pools corresponding to the plurality of intermediate nodes in the intermediate storage system to obtain a detection result; in the case that the detection result represents that the data between the plurality of intermediate nodes in the storage pool is the same, taking the intermediate storage system as an updated storage system, and taking the storage pool of the updated storage system as an updated storage pool.
[0103] In the embodiments of the present application, the intermediate storage system can be a transitional storage environment after data rewriting is completed and before consistency verification is passed, used for safety testing and verification. The data consistency detection can be a verification process for the integrity and synchronization of the data stored by the plurality of nodes in the distributed storage system. The intermediate node can be a physical or logical unit participating in data storage and processing in the distributed storage system. The updated storage system can be a healthy storage environment formally put into use after all verifications are passed.
[0104] For example, the manner of determining the updated storage system can include that the updated hard disk can be mounted to the distributed storage cluster, the storage pool parameters and access strategy are configured, the storage service is started, and the basic functions are verified; the checksum hash value of each data block is calculated by starting the local data scanning of all nodes in parallel, and the scanning results of the nodes are collected for centralized analysis; the data differences are identified and the difference report is generated by comparing the checksums of the same data block on different nodes and verifying the consistency and synchronization state between the node replicas, and then the health degree of the intermediate storage system is evaluated according to the consistency degree, and the intermediate storage system is taken as the updated storage system in the case that the health degree reaches the health standard.
[0105] For example, after the updated hard disks of all distributed storage nodes are mounted, the storage pool repair tool can be started, the data of the updated hard disk is restored to the distributed storage cluster, and the damaged distributed storage is repaired. The repaired distributed storage data can be checked by using the storage data checking program, and if the storage data is consistent, the repair work is completed, and the updated storage system is determined.
[0106] According to the embodiments of the present application, the complete synchronization of the data node replicas in the distributed system can be ensured by the multi-node cross verification, and the business logic error caused by the inconsistent data between nodes can be prevented; the hierarchical detection mechanism is adopted to comprehensively verify from the basic metadata to the actual content, and the potential risk of the inconsistent data is eliminated, and the consistency detection process can identify and locate the specific position of the data difference, and provide accurate guidance for targeted repair. The introduction of the intermediate storage system establishes a safe buffer layer, prevents the data with potential problems from directly affecting the production environment, the multi-node collaborative verification mechanism avoids the limitation of single-point detection, improves the probability and accuracy of problem discovery, the automatic retry and repair mechanism when the verification fails reduces the demand for manual intervention, and improves the automatic recovery ability of the system.
[0107] According to the embodiments of the present application, the above method further includes: in the case that the detection result represents that the data of the plurality of intermediate nodes in the storage pool exists data loss or misplacement, the intermediate node with data loss or misplacement is detected and data repair is performed to obtain the updated storage pool.
[0108] In the embodiments of the present application, the data loss can be a state that part of the data blocks in the storage node are completely lost or inaccessible, which is manifested as that there is no corresponding physical data content for the logical address. The data misplacement can be a state that the data blocks are stored in the wrong physical location or the logical address mapping is wrong, which is manifested as that the address pointer does not match the actual content. The detection result can be a node data state analysis report result obtained by consistency comparison, hash checking and metadata verification. The data repair can be a data reconstruction, address remapping or content correction operation taken for the data loss or misplacement problem.
[0109] For example, if the data between multiple intermediate nodes is not the same, data inconsistency analysis can be performed, and the analysis result is whether the data is misaligned or the data is missing, and a data repair tool is called to restore the data from the updated hard disk to the distributed storage cluster until the storage data verification is consistent, the distributed storage data repair program is exited, and the updated storage pool is obtained.
[0110] For example, for data missing repair, it can be checked whether the local node copy is available, a healthy copy is obtained from the adjacent node, and data reconstruction is performed based on the checksum. For data misalignment correction, the correct data location can be verified, the data can be migrated to the correct address, and all related mapping tables can be updated. During the repair process, the repair progress and success rate can be monitored in real time, abnormal conditions during the repair process can be recorded, and repair parameters and strategies can be dynamically adjusted. After verifying the integrity of the repaired data, the correctness of the address mapping, and the stability of the data access, the updated storage pool is obtained.
[0111] It should be noted that after the hard disk data recovery is completed, the updated hard disk with complete data can be inserted into the server slot, and the updated hard disk is not manually mounted in the hyper-converged system to avoid formatting the hard disk to erase data. The user needs to start the distributed storage repair program on the distributed storage node where the updated hard disk is inserted, and the repair program mounts all updated hard disks to the distributed storage system.
[0112] In a feasible embodiment, after the updated hard disks of all distributed storage nodes are mounted, the storage pool repair tool is started, the updated data of the updated hard disks is restored to the distributed storage cluster, and the damaged distributed storage is repaired. The repaired distributed storage data can be checked, and if the storage data is consistent, the repair work is completed, and if the data is inconsistent, data inconsistency analysis is performed, and the inconsistent reason is that the data is misaligned or the storage data is missing, and a data repair tool is called to extract the updated data from the updated hard disk to the distributed storage cluster until the storage data verification is consistent, and the distributed storage data repair program is exited.
[0113] According to the embodiments of the present application, through accurate problem detection and classification, appropriate repair solutions can be adopted for different types of data problems, improving the accuracy and integrity of the repair; the hierarchical repair strategy ensures the adaptability to different degrees of data problems, gradually solving from simple to complex, avoiding over-repair or insufficient repair.
[0114] According to an embodiment of the present application, the intermediate node with data loss or misplacement is detected and data repair is performed to obtain an updated storage pool, including: performing non-destructive detection on the intermediate node with data loss or misplacement to obtain entry type and entry address of inconsistent state, wherein the entry type includes at least one of metadata inconsistency, data block inconsistency and index node inconsistency; in the case of metadata inconsistency, reconstructing the damaged metadata based on the metadata of the available intermediate node in the storage pool; in the case of data block inconsistency, synchronizing the data block from the available intermediate node; and in the case of index node inconsistency, repairing or reconstructing the damaged index node.
[0115] In an embodiment of the present application, the non-destructive detection can be a read-only scanning and verification of the intermediate node without modifying the data in the intermediate node, and a technical means for identifying the inconsistent state of the data. The entry type of the inconsistent state can be a data abnormality classification found by detection, including three inconsistent types of metadata, data block and index node. The metadata inconsistency can refer to abnormal file system structure information, including directory structure, file attribute, permission setting and other description information errors. The data block inconsistency can refer to abnormal actual stored user data content, including data damage, loss or version inconsistency. The index node inconsistency can refer to abnormal file index information, including node damage, pointer error, link failure and other problems.
[0116] The entry address can be a location identifier of the inconsistent data on the storage medium, used for accurately positioning the location where the problem occurs. The available intermediate node can be a node with complete data and passed consistency verification in the distributed storage system, serving as a reference source for data repair.
[0117] For example, data repair can be performed in parallel on the intermediate node with data loss or misplacement. Multiple detection threads can be started simultaneously to process different data areas, the detection intensity can be dynamically adjusted according to the intermediate node load, and the detection results can be analyzed in real time; transactional update is used to ensure consistency, incremental synchronization is used to reduce data transmission, reference counting verification is used to ensure correctness, a repair task scheduling system is established, multiple repair tasks are executed in parallel, and a resource conflict resolution mechanism is implemented. It can be understood that the parallel processing mechanism can fully utilize system resources and shorten the overall repair time.
[0118] Figure 4B A flowchart of a storage pool repair method according to an embodiment of the present application is shown.
[0119] The storage hard disk of the virtual machine data in the hyper-converged system is stored in a storage pool in a file form, and the data storage pool needs to be repaired first, then the virtual machine hard disk is repaired, and the virtual machine repair is completed. The storage pool can be constructed by a cluster file system, and the inconsistent state is confirmed by analyzing the cluster file system of the storage pool.
[0120] As shown in Figure 4B The storage pool repair method can include operations S401-S406.
[0121] In operation S401, the inconsistent state of the storage pool is detected by a detection instruction (fsck.ocfs2 -nf / dev / mapper / XX). If consistent, operation S402 is performed. If inconsistent, operation S403 is performed.
[0122] In operation S402, the storage pool repair is completed.
[0123] In operation S403, the entry type and entry address of the inconsistent state are recorded.
[0124] In operation S404, the cluster file system repair program is started. For example, the cluster file system with inconsistent state is repaired by a repair instruction (fsck.ocfs2 -fy / dev / mapper / XX).
[0125] In operation S405, after the repair is completed, the consistency of the repaired cluster file system is scanned again. If consistent, operation S402 is performed. If inconsistent, operation S406 is performed.
[0126] In operation S406, the cluster file system repair program is called again until all cluster file system states are repaired and consistent.
[0127] According to the embodiments of the present application, the problem type is accurately identified by non-destructive detection, the false operation on healthy data is avoided, the targeted repair strategy is adopted based on the entry type, the accuracy and effectiveness of the repair operation are improved, the entry address can realize accurate positioning to reduce the repair range, and the influence on normal data is reduced.
[0128] According to an embodiment of the present application, based on the storage state data of the update storage pool and the configuration state data of the initial virtual machine, the initial virtual machine is recovered to obtain an intermediate virtual machine, comprising: determining the recovery mode of the initial virtual machine based on the integrity of the storage pool and the configuration state data, wherein the integrity of the storage pool represents the data integrity and functional integrity of the plurality of nodes in the storage pool, and the recovery mode comprises an initial virtual machine recovery mode and an update virtual machine mounting recovery mode; in the case of the initial virtual machine recovery mode, registering the initial virtual machine disk and starting the initial virtual machine, and taking the normally started initial virtual machine as the intermediate virtual machine; in the case of the update virtual machine mounting recovery mode, creating an update virtual machine and mounting the initial virtual machine disk, and taking the normally started update virtual machine as the intermediate virtual machine.
[0129] In an embodiment of the present application, the integrity of the storage pool can be used to represent the degree of normal access to data and functional integrity in the storage pool, including two dimensions of data integrity and service functional integrity. The initial virtual machine recovery mode can be a recovery method of directly recovering the system and data on the original virtual machine framework. The update virtual machine mounting recovery mode can be an innovative recovery method of creating a new virtual machine instance and mounting the original virtual disk.
[0130] For example, the adaptive recovery virtual machine process according to the integrity score of the storage pool can include: based on the integrity quantitative evaluation, setting the data integrity score (0-100 points), based on the storage pool data block verification success rate, the functional integrity score (0-100 points), based on the storage service availability test result, the comprehensive score = data score × 60% + functional score × 40%.
[0131] According to the comprehensive score, the recovery strategy is determined, in the case of the comprehensive score ≥ 80 points, the initial virtual machine recovery mode is adopted; in the case of the comprehensive score < 80 points but ≥ 50 points, the update virtual machine mounting recovery mode is adopted; in the case of the comprehensive score < 50 points, a hybrid recovery mode (preferably trying the mounting mode) can be adopted.
[0132] Performing the recovery operation, the initial virtual machine recovery mode can include verifying the virtual machine configuration file integrity, registering the virtual disk to the virtualization platform, starting the virtual machine and monitoring the starting process. The update virtual machine mounting recovery mode can include: creating a new virtual machine instance (same configuration), mounting the original virtual disk to the new instance, adjusting the hardware compatibility settings, and then taking the normally started update virtual machine as the intermediate virtual machine.
[0133] In a feasible embodiment, after the data storage pool is repaired, the virtual machine can be recovered in two ways. Way one is to recover on the initial virtual machine, and way two is to recover in the update virtual machine mounting recovery disk mode.
[0134] In the case that the hyper-converged system virtual machine configuration information is completed, in the first mode, the initial virtual machine disk can be used to restore the virtual machine by registering the virtual machine, the virtual machine is started normally, and the virtual machine restoration is completed. In the second mode, the virtual machine can be restored by updating the initial disk mounted by the initial virtual machine, the virtual machine is started normally, and the virtual machine restoration is completed.
[0135] If the virtual machine fails to start in the above-mentioned modes, the virtual machine recovery program can automatically perform a virtual machine system startup failure process analysis, detect the virtual machine operating system, and determine whether the startup failure reason is a configuration problem or a data problem. If it is a system configuration problem, the operating system is entered into a safe mode, the system configuration is repaired, the virtual machine operating system is restarted, and the virtual machine startup is completed. If it is a data problem, a disk data detection tool is started to detect disk data, data of a missing or incorrect virtual machine disk is repaired, the virtual machine is started again through the above-mentioned two virtual machine restoration modes, the virtual machine is started normally, and the virtual machine recovery program is completed.
[0136] According to the embodiments of the present application, the most suitable recovery mode is selected by storage state evaluation to avoid forced recovery under inappropriate storage conditions. The dual recovery path ensures automatic attempt of an alternative solution when one mode fails, improves the overall recovery success rate, and the intelligent decision mechanism dynamically adjusts the recovery strategy based on real-time state data to adapt to various complex failure scenarios. The optimal recovery path is quickly selected according to the storage pool integrity, which can reduce the time waste of trial and error; and the parallel operation synchronously verifies and optimizes in the recovery process, which shortens the overall recovery time.
[0137] According to the embodiments of the present application, the above-mentioned method further comprises: in response to the normal startup of the intermediate virtual machine into the operating system, verifying the business data integrity in the operating system, and generating a data detection result; in the case that the data detection result indicates that the business data satisfies the integrity, starting the application program, testing the startup state and running state of the application program, and obtaining an application detection result.
[0138] In the embodiments of the present application, the normal startup into the operating system can represent that the virtual machine successfully completes the booting process, the operating system kernel is loaded, the system service startup is completed, and the state of being acceptable to user operation is reached. The business data integrity can be the accuracy and consistency of a data set relied on by a business application program, including database files, configuration files, transaction logs, user data, etc. The data detection result can be a quantitative evaluation result generated after verifying the business data integrity, including an integrity score, an error type classification, and problem positioning information. The application program can be a business software system running on the virtual machine, providing specific business functions, such as a database service, a Web application, etc.
[0139] The startup state can represent the success degree of the application initialization process, including the normal completion of service startup, resource allocation, dependent component loading, etc. The running state can be the ability of the application to process business requests after startup, including function normality, performance stability, and reasonable resource usage. The application detection result can be a comprehensive evaluation result of the application startup and running state test, including availability indicators and performance indicators.
[0140] For example, data detection and application detection can be performed through hierarchical verification. Environment preparation can be performed first, and the intermediate virtual machine is fully started, network connectivity is confirmed, and data verification tools and test frameworks are deployed and configured with verification required parameters and thresholds; thereby performing data integrity verification, the contents of which include infrastructure verification and business logic verification. Infrastructure verification can include checking database file integrity, verifying configuration file format and content correctness, and checking log file continuity and integrity. Business logic verification can include performing database consistency checks, verifying business data reference integrity, and checking transaction log continuity and recoverability. On this basis, a data detection report can be generated, including integrity score (e.g. 98.5% pass rate), detailed error list and positioning information, and repair suggestions and priority assessment.
[0141] Further, application program testing can be performed. This includes startup process testing, running state testing, and performance benchmarking. Among them, the startup process test can include monitoring the application startup time and resource occupation, verifying the correctness of service dependencies, and checking error and warning information in system logs. The running state test can include: executing core function test cases, verifying interface availability and response time, and testing business process end-to-end smoothness. Performance benchmarking can include comparing performance indicators before and after recovery, verifying system load bearing capacity, and confirming that resource usage efficiency is within the normal range. Finally, result analysis and decision making can be performed, by comprehensively analyzing the data detection result and the application detection result, to determine whether the recovery meets the business availability standard, and to decide whether to put into production or need to be repaired again.
[0142] In a feasible embodiment, virtual machine data integrity detection can be implemented in the following way. After the virtual machine is normally started, the virtual machine business data integrity is the final purpose of the recovery of the virtual machine, and the business data is complete and the application program can be normally started, which is the completion of the recovery work.
[0143] After the virtual machine starts to enter the operating system, the data checking tool automatically verifies the business data integrity, if the data is verified to be complete, the user starts the business application program, if the application program is started normally, the virtual machine recovery program is completed and exited. If the business data is not complete, the program calls the business data repair tool to start the data synchronization program, and performs the virtual machine disk and data storage pool verification comparison, if the virtual machine disk has inconsistent data, it is synchronized and repaired, after the data repair is completed, the virtual machine performs business data verification again, and after the verification is complete, it enters the business program starting stage. If the business application program of the virtual machine fails to start, the process of the business program starting failure is automatically analyzed, if there is a configuration problem, the user adjusts the configuration and restarts to try. If there is data error, the business data repair tool is called to check the data, and the data synchronization program is started, the virtual machine disk and the data storage pool are verified and compared, if the inconsistent data of the virtual machine disk is synchronized and repaired, after the data repair is completed, the virtual machine performs business data verification again, and after the verification is passed, it enters the business program starting stage again.
[0144] The virtual machine can be normally started, and the business application program can be normally used, which indicates that the virtual machine data recovery work is completed, and the user further verifies the application program. If the customer encounters data error problem in the application program use process, the business data repair tool can be started again to perform business data synchronization again.
[0145] According to the embodiment of the application, through the systematic business data integrity verification, the actual effect of data recovery can be objectively evaluated, and the situation of inconsistent data after surface recovery is avoided; the application program starting and running state test ensures that the business function is truly available, rather than only system level recovery, the hierarchical verification mechanism distinguishes the verification strength according to the business importance, and optimizes the verification efficiency and quality balance.
[0146] In an embodiment, the virtual machine recovery method can be implemented by a virtual machine recovery device, which can include a hard disk data recovery module, a distributed storage repair module, and a virtual machine recovery module, to realize full-chain repair from bottom-layer data to upper-layer service application, and achieve the purpose of lossless recovery of virtual machine service data. The hard disk data recovery module can realize full-amount recovery of fault hard disk data to an updated hard disk, extract available data and damaged sector data from the fault hard disk, repair the data of the damaged sector, and achieve lossless recovery of hard disk data. The module can be extended to be applied to recovery of fault hard disks of servers, storages, and other devices. The distributed storage repair module can realize distributed storage repair and data storage pool data recovery, automatically mount the updated hard disk executed by a program, synchronize the data in the updated hard disk to the distributed storage system, repair the distributed storage, and repair the file system of the data storage pool through verification, to realize repair at the level of the storage file system. The module can be integrated into distributed storage products and hyper-converged products, and applied to repair of storage platforms and file systems.
[0147] The virtual machine recovery module can realize repair of virtual machine disks, recovery of service data, verification and repair of service applications, and achieve lossless recovery of virtual machine service applications. The module can be integrated into hyper-converged, virtualized, and other cloud computing platforms as a high-level function supplement for virtual machine recovery. It should be noted that the above function modules can be applied to multiple platforms and devices, and can be integrated into hyper-converged, distributed storage, and other platforms as function modules.
[0148] Based on the virtual machine recovery method, the application further provides a virtual machine recovery device. The following will be described in detail in combination with Figure 5 the device.
[0149] Figure 5 A structure block diagram of the virtual machine recovery device according to an embodiment of the application is shown.
[0150] As Figure 5 shown, the virtual machine recovery device 500 of this embodiment includes a write module 510, an update module 520, a recovery module 530, and a determination module 540.
[0151] The write module 510 is configured to write available data and fault data in the fault hard disk to an updated hard disk in sequence as updated available data and updated fault data, in response to detection of a fault hard disk mounted by an initial virtual machine. In an embodiment, the write module 510 can be configured to perform the operation S210 described above, and thus details are not repeated here.
[0152] The updating module 520 is configured to update the storage pool corresponding to the plurality of nodes in the distributed storage system based on the update available data and the update failure data, to obtain an updated storage pool. In an embodiment, the updating module 520 can be configured to perform operation S220 described above, and details are not repeated here.
[0153] The recovery module 530 is configured to recover the initial virtual machine based on the storage state data of the updated storage pool and the configuration state data of the initial virtual machine, to obtain an intermediate virtual machine. In an embodiment, the recovery module 530 can be configured to perform operation S230 described above, and details are not repeated here.
[0154] The determining module 540 is configured to determine the intermediate virtual machine as a target virtual machine to perform the service processing by using the target virtual machine, in a case where the data detection result and the application detection result of the intermediate virtual machine satisfy a preset condition, wherein the preset condition includes a data integrity condition of the data detection result and an application program test condition of the application detection result. In an embodiment, the determining module 540 can be configured to perform operation S240 described above, and details are not repeated here.
[0155] According to an embodiment of the present application, based on the writing module 510, the updating module 520, the recovery module 530 and the determining module 540 in the virtual machine recovery apparatus 500, the available data and the failure data in the failure hard disk are written into the update hard disk, and the storage pool of the distributed storage system is updated in real time based on this, which realizes local data recovery without relying on an external backup platform, reduces operation cost and compatibility requirements for data-intensive applications. Through the multiple recovery strategies based on the storage state data and the configuration state data, high-granularity real-time data protection is realized, and the risk of data loss caused by large backup granularity in the traditional backup scheme is avoided. In the recovery process, the occupation of network resources and storage resources is minimized through localized data operation, and the performance of other services in the hyper-converged system is ensured not to be affected. In addition, the storage pool updating and virtual machine recovery based on the built-in redundancy mechanism of the distributed storage system eliminate the dependence on the data backup platform, effectively cope with the superimposed failure scenario through the system automatic recovery capability, and significantly improve the business continuity and system reliability.
[0156] According to an embodiment of the present application, the writing module 510 includes a writing submodule and a determining submodule. The writing submodule is configured to write the available data as the update available data into the update hard disk based on a mapping relationship between the available address in the failure hard disk for storing the available data and the update available address in the update hard disk. The determining submodule is configured to determine a comparison result between the available data and the update available data based on the error correction code of the available data, and write the failure data as the update failure data into the update hard disk in a case where the comparison result satisfies a preset comparison condition.
[0157] According to the embodiment of the present application, the determining submodule comprises a writing unit and a rewriting unit. The writing unit is configured to write the failed sector address in the failure data into the check area, and write the failed sector address in the check area into the update hard disk as an update failed sector address in the case that the failed sector address in the check area is the same as the failed sector address of the failed hard disk. The rewriting unit is configured to perform data rewriting on the update failed sector in the update hard disk based on the update failed sector address.
[0158] According to the embodiment of the present application, the rewriting unit comprises a classification subunit, a single write subunit, a batch write subunit and a backup subunit. The classification subunit is configured to classify the update failed sector based on the damage degree of the failure data in the failed hard disk, and obtain at least one classification result. The single write subunit is configured to perform single write of the failure data to the update failed sector in the case that the classification result represents that the damage degree of the failure data is less than a first damage threshold. The batch write subunit is configured to perform batch write of the failure data to the update failed sector in the case that the classification result represents that the damage degree of the failure data is greater than or equal to the first damage threshold and less than a second damage threshold. The backup subunit is configured to backup the failure data before performing batch write to the update failed sector in the case that the classification result represents that the damage degree of the failure data is greater than or equal to the second damage threshold.
[0159] According to the embodiment of the present application, the update module 520 comprises a mounting submodule, a detecting submodule and a replacing submodule. The mounting submodule is configured to mount the update hard disk after data rewriting to the distributed storage system to obtain an intermediate storage system in the case that the update available data and the update failure data complete data rewriting in the update hard disk. The detecting submodule is configured to perform data consistency detection on the storage pools corresponding to the plurality of intermediate nodes in the intermediate storage system to obtain a detection result. The replacing submodule is configured to replace the intermediate storage system with an update storage system and replace the storage pool of the update storage system with an update storage pool in the case that the detection result represents that the data between the plurality of intermediate nodes in the storage pool is the same.
[0160] According to the embodiment of the present application, the apparatus further comprises a node detecting module configured to detect the intermediate node with data loss or misplacement and perform data repair to obtain the update storage pool in the case that the detection result represents that the data of the plurality of intermediate nodes in the storage pool has data loss or misplacement.
[0161] According to an embodiment of the present application, the node detecting module comprises a detecting submodule and a reconstructing submodule. The detecting submodule is configured to non-destructively detect an intermediate node with data loss or misplacement to obtain an entry type and an entry address of an inconsistency state, wherein the entry type comprises at least one of metadata inconsistency, data block inconsistency and index node inconsistency. The reconstructing submodule is configured to, in the case that the entry type is metadata inconsistency, reconstruct the damaged metadata based on the metadata of the available intermediate node in the storage pool; in the case that the entry type is data block inconsistency, synchronize the data block from the available intermediate node; and in the case that the entry type is index node inconsistency, repair or reconstruct the damaged index node.
[0162] According to an embodiment of the present application, the recovery module 530 comprises a mode determining submodule, a registering submodule and a creating submodule. The mode determining submodule is configured to determine a recovery mode of the initial virtual machine based on the integrity of the storage pool and the configuration state data, wherein the integrity of the storage pool represents the data integrity and the functional integrity of the plurality of nodes in the storage pool, and the recovery mode comprises an initial virtual machine recovery mode and an updated virtual machine mounting recovery mode. The registering submodule is configured to, in the case that the recovery mode is the initial virtual machine recovery mode, register the initial virtual machine disk and start the initial virtual machine, and take the normally started initial virtual machine as the intermediate virtual machine. The creating submodule is configured to, in the case that the recovery mode is the updated virtual machine mounting recovery mode, create the updated virtual machine and mount the initial virtual machine disk, and take the normally started updated virtual machine as the intermediate virtual machine.
[0163] According to an embodiment of the present application, the above device further comprises a verifying module and a starting module. The verifying module is configured to, in response to the intermediate virtual machine normally starting into the operating system, verify the service data integrity in the operating system to generate a data detection result. The starting module is configured to, in the case that the data detection result indicates that the service data satisfies the integrity, start the application program, test the starting state and the running state of the application program to obtain an application detection result.
[0164] According to an embodiment of the present application, any of the write module 510, the update module 520, the recovery module 530 and the determination module 540 can be combined in one module, or any of them can be split into multiple modules. Alternatively, at least part of the function of one or more of these modules can be combined with at least part of the function of the other modules, and implemented in one module. According to an embodiment of the present application, at least one of the write module 510, the update module 520, the recovery module 530 and the determination module 540 can be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on chip, a system on board, a system on package, an application specific integrated circuit (ASIC), or any other reasonable way of integrating or packaging a circuit, etc. in hardware or firmware, or implemented in any one of software, hardware and firmware or in a proper combination of any of them. Alternatively, at least one of the write module 510, the update module 520, the recovery module 530 and the determination module 540 can be at least partially implemented as a computer program module which, when executed, can perform the corresponding function.
[0165] Figure 6 A block diagram of an electronic device suitable for implementing the virtual machine recovery method according to an embodiment of the present application is shown.
[0166] As shown in Figure 6 , the electronic device 600 according to an embodiment of the present application includes a processor 601 which can perform various appropriate actions and processes according to a program stored in a read only memory (ROM) 602 or a program loaded from a storage portion 608 into a random access memory (RAM) 603. The processor 601 can include, for example, a general purpose microprocessor (such as a CPU), an instruction set processor and / or a special purpose microprocessor (such as an application specific integrated circuit (ASIC)), etc. The processor 601 can also include an on-board memory for cache use. The processor 601 can include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present application.
[0167] In the RAM 603, various programs and data required for the operation of the electronic device 600 are stored. The processor 601, the ROM 602, and the RAM 603 are connected to each other via the bus 604. The processor 601 performs various operations of the method flow according to the embodiments of the present application by executing the programs in the ROM 602 and / or the RAM 603. It should be noted that the programs can also be stored in one or more memories other than the ROM 602 and the RAM 603. The processor 601 can also perform various operations of the method flow according to the embodiments of the present application by executing the programs stored in the one or more memories.
[0168] According to the embodiments of the present application, the electronic device 600 can further include an input / output (I / O) interface 605, which is also connected to the bus 604. The electronic device 600 can further include one or more of the following components connected to the input / output (I / O) interface 605: an input part 606 including a keyboard, a mouse, etc.; an output part 607 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage part 608 including a hard disk, etc.; and a communication part 609 including a network interface card such as a LAN card, a modem, etc. The communication part 609 performs communication processing via a network such as the Internet. A drive 610 is also connected to the input / output (I / O) interface 605 as necessary. A removable medium 611 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc. is mounted on the drive 610 as necessary, so that a computer program read out therefrom is installed in the storage part 608 as necessary.
[0169] The present application also provides a computer readable storage medium, which can be included in the device / apparatus / system described in the above embodiments; or can exist separately without being assembled into the device / apparatus / system. The above computer readable storage medium carries one or more programs, when the one or more programs are executed, the method according to the embodiments of the present application is implemented.
[0170] According to an embodiment of the present application, the computer readable storage medium can be a non-transitory computer readable storage medium, for example, can include but not limited to: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In the present application, the computer readable storage medium can be any tangible medium that contains or stores a program that can be used by or in connection with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present application, the computer readable storage medium can include one or more memories of the ROM 602 and / or the RAM 603 described above and / or one or more memories other than the ROM 602 and the RAM 603.
[0171] Embodiments of the present application also include a computer program product, which includes a computer program containing program codes for executing the methods shown in the flowcharts. When the computer program product is run in a computer system, the program codes are used to make the computer system implement the virtual machine recovery method provided by the embodiments of the present application.
[0172] The above functions defined in the system / device of the embodiments of the present application are performed when the computer program is executed by the processor 601. According to an embodiment of the present application, the system, device, module, unit, etc. described above can be implemented by computer program modules.
[0173] In one embodiment, the computer program can rely on tangible storage media such as optical storage media, magnetic storage media, etc. In another embodiment, the computer program can also be transmitted, distributed, and downloaded in the form of signals on network media, and be downloaded and installed through the communication part 609, and / or be installed from the detachable medium 611. The program codes contained in the computer program can be transmitted by any appropriate network media, including but not limited to: wireless, wired, etc., or any suitable combination of the foregoing.
[0174] In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 609, and / or be installed from the detachable medium 611. When the computer program is executed by the processor 601, the above functions defined in the system of the embodiments of the present application are performed. According to an embodiment of the present application, the system, device, apparatus, module, unit, etc. described above can be implemented by computer program modules.
[0175] According to embodiments of the present application, program code for implementing the computer programs provided by embodiments of the present application can be written in any combination of one or more programming languages, and can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. The programming language can include, but is not limited to, Java, C++, python, "C" language, or similar programming languages. The program code can execute entirely on the user's computing device, partly on the user's device, as a stand-alone software package, partly on the remote computing device, or entirely on the remote computing device or server. In the latter scenario, the remote computing device can be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computing device, such as through the Internet using an Internet Service Provider.
[0176] The computer program instructions can also be loaded onto a computer or other programmable information processing apparatus to cause a series of operations to be performed on the computer or other programmable information processing apparatus to produce a computer implemented process such that the instructions which execute on the computer or other programmable information processing apparatus implement the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0177] Those skilled in the art will appreciate that the features recited in the various embodiments of the present application can be combined and / or integrated in a variety of ways, even if such combinations or integrations are not expressly noted in the present application. In particular, the features recited in the various embodiments of the present application can be combined and / or integrated in a variety of ways without departing from the spirit and scope of the present application. All such combinations and / or integrations are within the scope of the present application.
[0178] The embodiments of the present application have been described above. However, these embodiments are merely for the purpose of illustration and are not intended to limit the scope of the present application. Although the respective embodiments are described above separately, this does not mean that the measures in the respective embodiments cannot be used advantageously in combination. Various alternatives and modifications can be made to the embodiments of the present application by those skilled in the art without departing from the scope of the present application, and all such alternatives and modifications are within the scope of the present application.
Claims
1. A virtual machine recovery method characterized by comprising: The method comprises: in response to detecting a fault hard disk mounted by an initial virtual machine, sequentially writing available data and fault data in the fault hard disk into an update hard disk as update available data and update fault data respectively; based on the update available data and the update fault data, updating storage pools corresponding to a plurality of nodes in a distributed storage system to obtain updated storage pools; based on storage state data of the updated storage pools and configuration state data of the initial virtual machine, recovering the initial virtual machine to obtain an intermediate virtual machine; in a case where data detection results and application detection results of the intermediate virtual machine satisfy preset conditions, determining the intermediate virtual machine as a target virtual machine to execute service processing by using the target virtual machine, wherein the preset conditions include a data integrity condition of the data detection results and an application program test condition of the application detection results; wherein sequentially writing the available data and the fault data in the fault hard disk into the update hard disk as the update available data and the update fault data comprises: based on a mapping relationship between an available address in the fault hard disk for storing the available data and an update available address in the update hard disk, writing the available data as the update available data into the update hard disk, wherein writing the available data as the update available data into the update hard disk comprises: writing a fault sector address in the available data into a verification area, in a case where the fault sector address in the verification area is same as a fault sector address of the fault hard disk, writing the fault sector address in the verification area into the update hard disk as an update fault sector address; based on the update fault sector address, rewriting data in an update fault sector in the update hard disk; based on an error correction code of the available data, determining a comparison result between the available data and the update available data, and in a case where the comparison result satisfies a preset comparison condition, writing the fault data as update fault data into the update hard disk.
2. The method of claim 1, wherein, based on the update fault sector address, rewriting data in an update fault sector in the update hard disk comprises: based on a damage degree of the fault data in the fault hard disk, classifying the update fault sector to obtain at least one classification result; in a case where the classification result represents that the damage degree of the fault data is less than a first damage threshold, writing the fault data into the update fault sector for a single time; in a case where the classification result represents that the damage degree of the fault data is greater than or equal to the first damage threshold and less than a second damage threshold, writing the fault data into the update fault sector in batches; in a case where the classification result represents that the damage degree of the fault data is greater than or equal to the second damage threshold, backing up the fault data before writing the fault data into the update fault sector in batches.
3. The method of claim 1, wherein, based on the update available data and the update fault data, updating storage pools corresponding to a plurality of nodes in a distributed storage system to obtain updated storage pools comprises: In the case that the updated available data and the updated failure data complete data rewriting in the updated hard disk, the updated hard disk after data rewriting is mounted to the distributed storage system to obtain an intermediate storage system; Data consistency detection is performed on storage pools corresponding to a plurality of intermediate nodes in the intermediate storage system to obtain a detection result; In the case that the detection result represents that data among the plurality of intermediate nodes in the storage pool is the same, the intermediate storage system is taken as an updated storage system, and a storage pool of the updated storage system is taken as the updated storage pool.
4. The method of claim 3, wherein, The method further comprises: In the case that the detection result represents that data of the plurality of intermediate nodes in the storage pool has data loss or misplacement, an intermediate node having data loss or misplacement is detected and data repair is performed to obtain the updated storage pool.
5. The method of claim 4, wherein, The intermediate node having data loss or misplacement is detected and data repair is performed to obtain the updated storage pool, comprising: Non-destructive detection is performed on the intermediate node having data loss or misplacement to obtain an entry type and an entry address in an inconsistent state, wherein the entry type comprises at least one of metadata inconsistency, data block inconsistency, and index node inconsistency; In the case that the entry type is metadata inconsistency, damaged metadata is reconstructed based on metadata of an available intermediate node in the storage pool; in the case that the entry type is data block inconsistency, data blocks are synchronized from the available intermediate node; and in the case that the entry type is index node inconsistency, damaged index nodes are repaired or reconstructed.
6. The method of claim 1, wherein, Based on storage state data of the updated storage pool and configuration state data of an initial virtual machine, the initial virtual machine is recovered to obtain an intermediate virtual machine, comprising: The recovery mode of the initial virtual machine is determined based on integrity of the storage pool and the configuration state data, wherein the integrity of the storage pool represents data integrity and functional integrity of a plurality of nodes in the storage pool, and the recovery mode comprises an initial virtual machine recovery mode and an updated virtual machine mounting recovery mode; In the case that the recovery mode is the initial virtual machine recovery mode, an initial virtual machine disk is registered and the initial virtual machine is started, and the initial virtual machine normally started is taken as the intermediate virtual machine; In the case that the recovery mode is the updated virtual machine mounting recovery mode, an updated virtual machine is created and an initial virtual machine disk is mounted, and the updated virtual machine normally started is taken as the intermediate virtual machine.
7. The method of claim 1, wherein, The method further comprises: In response to the intermediate virtual machine normally starting into an operating system, service data integrity in the operating system is verified to generate the data detection result; In the case that the data detection result indicates that the service data satisfies integrity, an application program is started, a start state and a running state of the application program are tested to obtain an application detection result.
8. An electronic device, comprising: one or more processors; a memory for storing one or more computer programs, characterized in that the one or more processors execute the one or more computer programs to implement the steps of the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Virtual machine recovery method and device, storage medium and electronic device
CN110196749A
Method and apparatus for repairing virtual machine in cloud environment, and electronic device
WO2025152682A1