Data recovery method and device, equipment, storage medium and program product
By dynamically applying for escape space in the all-flash array data storage system, combining hard disk status and data storage information, the problem of low garbage data recycling efficiency in the prior art is solved, and efficient data recycling and flexible adaptability of the storage system is achieved.
Patent Information
- Application Number
- CN202510362775.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-22
Smart Images

Figure CN120353384A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of storage technology, and in particular, to a data recovery method, apparatus, device, storage medium, and program product. Background Art
[0002] With the development of information technology, the amount of data has increased explosively. In the big data era, the demand for efficient and stable data storage is becoming more and more urgent. With the advantages of high-speed reading and writing, low latency, etc., all-flash arrays have become a key technology for coping with massive data storage in fields such as data centers.
[0003] All-flash array data storage generally adopts the Redundant Array of Independent Disks (RAID) technology. Among them, RAID 2.0 technology cuts the physical space of hard disks into multiple data blocks. These data blocks form logical blocks through specific mapping, and then multiple logical blocks are combined into stripes for storing user data. Since there may be both valid and garbage data blocks on a stripe, it is very necessary to selectively recycle the data blocks on the stripe and reorganize new stripes.
[0004] However, the garbage data recovery method in the related technology has the problem of low recovery efficiency. Summary of the Invention
[0005] Based on this, in view of the above technical problems, it is necessary to provide a data recovery method, apparatus, computer device, computer-readable storage medium, and computer program product that can improve data recovery efficiency.
[0006] In a first aspect, this application provides a data recovery method, which includes:
[0007] In a scenario where data recovery using the escape space is triggered, obtain the current storage hard disk information and the storage information of the data storage area in the storage system, and apply for a target escape space according to the hard disk information and the storage information of the data storage area;
[0008] Perform data recovery on the data in the data storage area according to the target escape space.
[0009] The data recovery method provided by the embodiments of the present application, in the scenario of triggering the use of the escape space for data recovery, obtains the current storage hard disk information and the storage information of the data storage area in the storage system, and applies for the target escape space according to the hard disk information and the storage information of the data storage area, and then performs data recovery on the data in the data storage area according to the target escape space. The above method not only considers the hard disk status of each hard disk in the current storage system, but also considers the data storage situation on the data storage area to be recovered in the storage system. By combining the two levels to apply for the escape space, the applied escape space always matches the hard disk status and data storage situation of each hard disk, and is not affected by the failure of the hard disk. That is, if there is a faulty hard disk, the applied escape space corresponds to the hard disks in the normal state. Compared with the existing method of using a fixed escape space for data recovery, the above method can flexibly apply for the escape space without waiting for the faulty hard disk to be repaired before data recovery. Therefore, the above method can greatly improve the data recovery efficiency.
[0010] In some of the embodiments, applying for the target escape space according to the current hard disk information and the storage information of the data storage area includes:
[0011] Determine whether there is a faulty hard disk in the storage system according to the current hard disk information;
[0012] If so, apply for the target escape space according to the hard disk information of the non-faulty hard disks in the storage system and the storage information of the data storage area;
[0013] If not, apply for the target escape space according to the storage information of the data storage area.
[0014] The method described in the embodiments of the present application adopts different application strategies for the faulty situation of the storage hard disks in the storage system, which can enable the storage system to better adapt to various changes. Whether it is a sudden hard disk failure or a change in daily data storage requirements, it can quickly respond to ensure the normal operation of the storage system and improve the reliability of the storage system.
[0015] In some of the embodiments, applying for the target escape space according to the hard disk information of the non-faulty hard disks in the storage system and the storage information of the data storage area includes:
[0016] Determine whether there is a redundant hard disk among the non-faulty hard disks according to the hard disk information of the non-faulty hard disks in the storage system;
[0017] If so, apply for the target escape space according to the storage information of the data storage area;
[0018] If not, apply for the target escape space according to the storage information of the data storage area and the hard disk information of the non-faulty hard disks.
[0019] When applying for a target escape space in the case of a faulty hard disk in the storage system, the method according to the embodiments of the present application can adopt different application strategies according to whether there are redundant hard disks in the non-faulty hard disks, enabling the storage system to better adapt to different resource conditions. Whether the storage system resources are sufficient or relatively tense, the corresponding strategies can be used to meet the needs of data storage and management. When there are redundant hard disks in the storage system, it indicates that the current resources are relatively sufficient, and at this time, the required target escape space can be directly applied for. When there are no redundant hard disks in the storage system, it indicates that the current resources are relatively tense, and at this time, the required target escape space needs to be reduced to achieve efficient data recovery.
[0020] In some embodiments, when applying for a target escape space according to the storage information of the data storage area and the hard disk information of the non-faulty hard disks, it includes:
[0021] Determine the third quantity of data columns and the fourth quantity of parity columns included in the target escape space according to the first quantity of data columns and the second quantity of parity columns included in the data storage area, and the hard disk information of the non-faulty hard disks;
[0022] Apply for the target escape space according to the third quantity and the fourth quantity.
[0023] The method according to the embodiments of the present application determines the third quantity of data columns and the fourth quantity of parity columns of the target escape space according to the first quantity of data columns, the second quantity of parity columns in the data storage area, and the hard disk information of the non-faulty hard disks, which can ensure that the layout of the target escape space is adapted to the layout of the data storage area, avoid applying for an overly large escape space, and thus reduce unnecessary resource occupation.
[0024] In some embodiments, determining the third quantity of data columns and the fourth quantity of parity columns included in the target escape space according to the first quantity of data columns and the second quantity of parity columns included in the data storage area, and the hard disk information of the non-faulty hard disks, includes:
[0025] Determine the reduced column quantity according to the first quantity of data columns and the second quantity of parity columns included in the data storage area, and the hard disk information of the non-faulty hard disks;
[0026] Determine the reduced column type according to the second quantity of parity columns;
[0027] Determine the third quantity of data columns and the fourth quantity of parity columns included in the target escape space according to the reduced column type and the reduced column quantity.
[0028] In some embodiments, determining the reduced column type according to the second quantity of parity columns includes:
[0029] If the second quantity of the parity columns is greater than a preset value, determine that the reduced column type is the parity column type;
[0030] If the second quantity of the parity columns is not greater than the preset value, determine that the reduced column type is the data column type.
[0031] In the embodiments of the present application, for the method described in the embodiments of the present application, when there is no redundant hard disk in the storage system, first perform a parity column reduction operation on the required target escape space, and then perform a data column reduction operation, which can ensure that the data column ratio of the target escape space reaches the maximum, thereby maximizing the utilization rate of the target escape space and further improving the data recovery efficiency.
[0032] In some embodiments, the method further includes:
[0033] Detect whether the storage hard disks in the storage system meet the preset update conditions, and when the storage hard disks in the storage system meet the preset update conditions, perform a regional extension operation on the target escape space after data recovery according to the storage information of the data storage area.
[0034] For the method described in the embodiments of the present application, when the failed hard disk in the storage system is restored or a new hard disk is added, performing a regional extension operation on the target escape space can make the distribution of data on each hard disk more balanced, give full play to the storage capacity of each hard disk, and further improve the resource utilization efficiency of the entire storage system. At the same time, it can avoid the situation that some hard disks are overused while other hard disks are idle, which helps to extend the service life of the hard disks and reduce the storage cost.
[0035] In a second aspect, the present application further provides a data recovery device, which includes:
[0036] An application module, configured to obtain the storage hard disk information and the storage information of the data storage area in the storage system in a scenario where data recovery using the escape space is triggered, and apply for a target escape space according to the hard disk information and the storage information of the data storage area;
[0037] A recovery module, configured to perform data recovery on the data in the data storage area according to the target escape space.
[0038] In a third aspect, the present application further provides a computer device, which includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:
[0039] In a scenario where data recovery using the escape space is triggered, obtain the storage hard disk information and the storage information of the data storage area in the storage system, and apply for a target escape space according to the hard disk information and the storage information of the data storage area;
[0040] Data recycling is performed on the data in the data storage area according to the target escape space.
[0041] In a fourth aspect, the present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:
[0042] In a scenario where triggering the use of the escape space for data recycling, obtain the storage hard disk information in the storage system and the storage information of the data storage area, and apply for the target escape space according to the hard disk information and the storage information of the data storage area;
[0043] Data recycling is performed on the data in the data storage area according to the target escape space.
[0044] In a fifth aspect, the present application also provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the following steps are implemented:
[0045] In a scenario where triggering the use of the escape space for data recycling, obtain the storage hard disk information in the storage system and the storage information of the data storage area, and apply for the target escape space according to the hard disk information and the storage information of the data storage area;
[0046] Data recycling is performed on the data in the data storage area according to the target escape space.
[0047] In the above data recycling method, device, equipment, storage medium and program product, the method obtains the current storage hard disk information in the storage system and the storage information of the data storage area in a scenario where triggering the use of the escape space for data recycling, and applies for the target escape space according to the hard disk information and the storage information of the data storage area, and then performs data recycling on the data in the data storage area according to the target escape space. The above method not only considers the hard disk status of each hard disk in the current storage system, but also considers the data storage situation on the data storage area to be recycled in the storage system. By combining two levels to apply for the escape space, the applied escape space is always matched with the hard disk status and data storage situation of each hard disk, and is not affected by the failure of the hard disk. That is, if there is a faulty hard disk, the applied escape space is correspondingly matched with the hard disks in the normal state. Compared with the existing method of using a fixed escape space for data recycling, the above method can flexibly apply for the escape space and does not need to wait for the faulty hard disk to be repaired before data recycling. Therefore, the above method can greatly improve the data recycling efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 It is the internal structure diagram of a computer device in some embodiments;
[0049] Figure 2 One of the flow diagrams of the data recovery method in some embodiments;
[0050] Figure 3 The structural diagram of the storage space in some embodiments;
[0051] Figure 4 One of the flow diagrams of the data recovery method in some embodiments;
[0052] Figure 5 One of the flow diagrams of the data recovery method in some embodiments;
[0053] Figure 6 One of the flow diagrams of the data recovery method in some embodiments;
[0054] Figure 7 One of the flow diagrams of the data recovery method in some embodiments;
[0055] Figure 8 One of the flow diagrams of the data recovery method in some embodiments;
[0056] Figure 9 One of the flow diagrams of the data recovery method in some embodiments;
[0057] Figure 10 One of the flow diagrams of the data recovery method in some embodiments;
[0058] Figure 11 One of the flow diagrams of the data recovery method in some embodiments;
[0059] Figure 12 One of the flow diagrams of the data recovery method in some embodiments;
[0060] Figure 13 The structural block diagram of the data recovery device in some embodiments. Detailed implementation manners
[0061] In the embodiments of the present application, the term "and / or" describes the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after.
[0062] In the embodiments of the present application, the term "plurality" refers to two or more, and other quantifiers are similar.
[0063] In the embodiments of the present application, the term "at least one" means one or more. For example, at least one of A, B, and C may mean: A exists alone, B exists alone, C exists alone, A and B exist simultaneously, A and C exist simultaneously, B and C exist simultaneously, and A, B, and C exist simultaneously, which are six cases in total.
[0064] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0065] With the development of information technology, the amount of data has grown explosively. In the big data era, the demand for efficient and stable data storage has become increasingly urgent. All-flash arrays, with advantages such as high-speed read and write and low latency, have become a key technology for dealing with massive data storage in fields such as data centers. All-flash array data storage generally adopts the Redundant Array of Independent Disks (RAID) technology. Among them, RAID 2.0 technology cuts the physical space of hard disks into multiple data blocks. These data blocks are mapped into logical blocks through specific mapping, and then multiple logical blocks are combined into stripes for storing user data. Since there may be both valid and garbage data blocks on a stripe, it is very necessary to selectively recycle the data blocks on the stripe and reorganize them into new stripes. However, the garbage data recovery methods in the related technologies have the problem of low recovery efficiency.
[0066] In view of this, the embodiments of the present application propose a data recovery method, device, equipment, storage medium, and program product, which can apply for a target escape space in combination with the current storage hard disk information and the storage information of the data storage area in the storage system, and then use the target escape space for data recovery to improve the data recovery efficiency. The present application provides a data recovery method, aiming to solve the above technical problems. The following embodiments will specifically illustrate the data recovery method described in the present application.
[0067] It should be noted that the beneficial effects brought by the embodiments of the present application or the technical problems solved are not limited to this one, and there may also be other implicit or related problems. For specific details, please refer to the descriptions of the following embodiments.
[0068] Next, the technical solutions of the present application and how the technical solutions of the present application solve the above technical problems will be described in detail with specific embodiments. These specific embodiments below can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. Next, the embodiments of the present application will be described in conjunction with the accompanying drawings.
[0069] In some embodiments, the data recovery method provided by the embodiments of the present application can be applied to a computer device as shown in Figure 1 The computer device can be a terminal or a server. Its internal structure diagram can be as shown in Figure 1 The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit, and an input device. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface, the display unit, and the input device are connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals in a wired or wireless manner. The wireless manner can be implemented through WIFI, a mobile cellular network, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it implements a data recovery method. The display unit of the computer device is used to form a visually visible picture, which can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covered on the display screen, or a button, a trackball, or a touchpad set on the computer device housing, or an external keyboard, a touchpad, or a mouse, etc.
[0070] Those skilled in the art can understand that Figure 1 The structure shown in
[0071] In some embodiments, as shown in Figure 2 A data recovery method is provided. Taking the computer device in Figure 1 as an example, the method includes the following steps:
[0072] S201, in a scenario where triggering the use of the escape space for data recovery, obtain the current storage hard disk information and the storage information of the data storage area in the storage system, and apply for a target escape space according to the hard disk information and the storage information of the data storage area.
[0073] Among them, the escape space refers to the space for recycling data in the storage system when both the garbage collection space (i.e., the GC space) and the reserved space (i.e., the OP space) in the storage space are full, that is, the space for garbage to escape. The target escape space is the escape space for garbage to escape. The storage hard disk can be a solid state drive (SSD for short), or other types of hard disks or disks. The storage hard disk information is used to detect the status of the hard disk group in the storage system and determine whether there are faulty hard disks in the hard disk group. The storage hard disk information can include at least one of the number of storage hard disks, the number of redundant hard disks, the remaining capacity of each storage hard disk, the health status of each storage hard disk, and the working performance of each storage hard disk. Among them, the hard disk group is a combination of hard disks used for garbage escape in the storage system disk array, and the storage system in the embodiment of the present application can be this hard disk group. The data storage area represents the space for storing the data to be recycled, and its form can be a stripe form, that is, the data storage area is the stripe to be recycled. The storage information of the data storage area includes data column information and parity column information. Among them, the data column information stores valid data and / or invalid data, as well as the number of data columns, and the parity column information stores parity information and the number of parity columns. The parity information is determined by the data information in the data column. Optionally, the number of data columns and the number of parity columns in the storage information of the data storage area can be represented by numbers or ratios. The data storage area in the embodiment of the present application can adopt the RAID 2.0 technology of storage pool space management. In this technology, the physical space of the hard disk is cut into data blocks of a fixed size and composed into stripes in units of logical blocks, so that multiple hard disks in the storage system participate in data reading and writing, and at the same time, the stored data is evenly distributed on each hard disk. Optionally, the data storage area in the embodiment of the present application can also adopt other storage pool space management technologies. Since there will be a large amount of duplicate data and garbage data in the stored data, both will be written into the flash memory blocks of each hard disk, and there may be valid data blocks and garbage data blocks on a stripe. Then, it is necessary to selectively recycle and reorganize the data blocks on this stripe into a new stripe to achieve the purpose of recycling valid data and releasing garbage data. This process is the data recycling process.
[0074] In the embodiments of the present application, the proportions of the garbage collection space (i.e., GC space), the reserved space (i.e., OP space), and the escape space in the entire storage space can be determined in advance according to the characteristics of the storage system, the data usage pattern, and the hardware conditions. For example, the escape space can be set to 5% - 10% of the total storage space. And capacity thresholds can be set for the GC space, the OP space, and the escape space respectively. When the usage amounts of the GC space, the OP space, and the escape space reach their respective thresholds, the corresponding warning mechanism is triggered. It should be noted that if the storage system consists of multiple storage hard disks, a partial area of one or more of the storage hard disks can be designated as the escape space. For example, as Figure 3 shown in the schematic structural diagram of the storage space, this storage space is composed of 6 hard disks (disk1 - disk6), and this storage space is divided into a GC space, an OP space, a storage space for parity data, an escape space, a storage space for metadata (i.e., the kSeg space in the figure), and a storage space for hard disk information (i.e., the critical disk space in the figure).
[0075] During the process of data recovery in the system, the computer device can use a system monitoring tool or write a script program to check the remaining space conditions of the GC space and the OP space in the storage system at regular time intervals (such as every minute, every hour, or every day). When it is detected that the remaining space of the GC space and the OP space reaches or is lower than the set trigger threshold, the escape space can be automatically triggered to start data recovery. Optionally, when it is detected that the remaining space of a certain storage hard disk or partition reaches or is lower than the set trigger threshold, an alarm prompt can be given. For example, an email or a text message is sent to notify the administrator, or in the local system, the built - in pop - up prompt function of the system or a third - party pop - up tool is used to issue an alarm to the local user to indicate that the user actively triggers the start of the escape space. It should be noted that the escape space trigger threshold in this embodiment can be determined according to the actual business requirements or storage requirements. This escape space trigger threshold can be a percentage of the remaining available space. For example, it is triggered when the remaining space is lower than 10%. Optionally, this escape space trigger threshold can also be a specific number of bytes. For example, it is triggered when the remaining space is less than 10GB. It should be noted that during the garbage collection process, usually a new strip of the same specification as the old strip is applied for garbage collection. For example, as Figure 4As shown, the data columns of the strip to be recycled (strip 1, strip 2, and strip 3 in the figure, also old strips) are 4 columns, which are distributed on disk1 - disk4 respectively, and the check columns are 2 columns, which are distributed on disk5 - disk6 respectively. When the hard disk group has sufficient capacity, usually an escape space or recycling space with 4 data columns and 2 check columns will be applied for maximum - efficiency data recycling. However, when the hard disk group has insufficient capacity, when applying for an escape space or recycling space, it may be necessary to perform a column - shrinking operation on the desired escape space or recycling space. The desired escape space is an escape space with the same number of data columns and parity columns as those of the old strip.
[0076] In the scenario where data recycling is triggered using the escape space, the computer device can use a system monitoring tool or write a script program to check the storage information of each storage hard disk in the hard disk group and the storage information of the data storage area at regular time intervals (such as every minute, every hour, or every day). Specifically, the method for obtaining the current storage hard disk information in the storage system can be: the computer device can obtain the total number of hard disks in the current hard disk group, and through health checks and status monitoring of each storage hard disk, determine whether there is a faulty hard disk. The computer device can also check whether redundant hard disks are configured in the current hard disk group. If there are redundant hard disks, further determine the redundant hard disks. The method for obtaining the storage information of the data storage area can be: the computer device can first determine the strip to be recycled, and then further determine the data column information and parity column information in the strip to be recycled. After the computer device obtains the current storage hard disk information and the storage information of the data storage area, it can apply for a physical space segment from each storage hard disk in the hard disk group according to the storage hard disk information and the storage information of the data storage area, and logically form the target escape space with the physical spaces applied on each storage hard disk. Among them, the number of data columns and parity columns of the target escape space can be the same as or different from the number of data columns and parity columns of the data storage area.
[0077] S202, perform data recycling on the data in the data storage area according to the target escape space.
[0078] In the embodiments of the present application, after the computer device applies for a target escape space based on the above steps, it can identify valid data and invalid data according to the metadata of the data in the data storage area, and then migrate the valid data to the target escape space in chunks. Among them, the metadata stores at least one of the valid / invalid status of the data, the storage location of the data on the physical disk, and the associated location of the data on the logical disk. The migration methods include the inheritance method and the copy method. After all the valid data is successfully migrated to the target escape space, the computer device can perform a clear or erase operation on all the data in the data storage area. Specifically, different methods can be used for the clear operation. For a simple data storage system, the metadata of the data storage area can be reset to mark the area as available space; for a storage system with high requirements for data security, a thorough erase operation can be performed to overwrite the physical medium of the data storage area multiple times. After the computer device clears the data, it can update the storage space management information of the storage system, mark the data storage area as released, and make it available for subsequent data storage, thus realizing the effective release of the storage space. Optionally, after all the valid data is successfully migrated to the target escape space, the computer device can verify the migration result. For example, by comparing whether the data information in the data storage area and the target escape space is consistent, the integrity and accuracy of the data can be checked. When the verification result indicates that the verification is passed, a clear or erase operation is performed on all the data in the data storage area.
[0079] The data recovery method provided by the embodiments of the present application, in the scenario of triggering the use of the escape space for data recovery, obtains the current storage hard disk information and the storage information of the data storage area in the storage system, and applies for a target escape space according to the hard disk information and the storage information of the data storage area, and then performs data recovery on the data in the data storage area according to the target escape space. The above method not only considers the hard disk status of each hard disk in the current storage system, but also considers the data storage situation on the data storage area to be recovered in the storage system. By combining the two levels to apply for the escape space, the applied escape space is always matched with the hard disk status and the data storage situation of each hard disk, and is not affected by the failure of the hard disk. That is, if there is a faulty hard disk, the applied escape space is correspondingly matched with the hard disks in the normal state. Compared with the existing method of using a fixed escape space for data recovery, the above method can flexibly apply for the escape space without waiting for the faulty hard disk to be repaired before data recovery. Therefore, the above method can greatly improve the data recovery efficiency.
[0080] In some embodiments, a method for applying for a target escape space is also provided, as Figure 5 shown, the "applying for a target escape space according to the current hard disk information and the storage information of the data storage area" in S201 above includes:
[0081] S301. Determine whether there is a faulty hard disk in the storage system according to the current hard disk information.
[0082] Among them, a faulty hard disk is a hard disk that cannot work properly or has low working efficiency. The faulty hard disk includes any one of an abnormal hard disk, a slow hard disk, and a full hard disk.
[0083] In the embodiment of the present application, after the computer device obtains the current hard disk information in the storage system, it can analyze the status of each hard disk in the hard disk group according to the current hard disk information to determine whether there is a faulty hard disk in the hard disk group of the storage system. Optionally, it can be determined whether there is a faulty hard disk in the storage system according to the status of the indicator light of the hard disk. For example, when the indicator light of the hard disk is off or always on and does not flash, it means that the hard disk is faulty. Optionally, it can be determined whether there is a faulty hard disk in the storage system by viewing the log information related to the hard disk. For example, a hard disk without remaining available space or with the remaining available space reaching the capacity threshold can be determined as a faulty hard disk. Among them, the capacity threshold can be determined according to actual business requirements or storage requirements. The capacity threshold can be a percentage of the remaining available space. For example, the capacity threshold is that the remaining available space is less than 10%. Optionally, the capacity threshold can also be a specific number of bytes. For example, the capacity threshold is that the remaining available space is less than 10 GB. Optionally, a hard disk whose read / write speed and / or the number of input / output operations per second (i.e., IOPS) does not meet the preset conditions can be determined as a faulty hard disk. Among them, the preset conditions can include that the read / write speed is less than the read / write speed threshold and / or the number of input / output operations per second is less than the IOPS threshold. The read / write speed threshold and the IOPS threshold can be determined according to actual needs.
[0084] S302. If there is, apply for a target escape space according to the hard disk information of the non-faulty hard disks in the storage system and the storage information of the data storage area.
[0085] Among them, the non-faulty hard disks are the hard disks in the storage system that are in a normal state.
[0086] In the embodiments of the present application, when the computer device determines that there is a faulty hard disk in the storage system based on the above steps, this situation indicates that the storage space capacity composed of the current hard disk group is not very abundant, and it may be necessary to perform a reduction operation on the desired escape space. Then, the computer device can apply for a target escape space according to the number information of non-faulty hard disks and the number of data blocks in the data storage area. For example, when the number of non-faulty hard disks is greater than the number of data blocks in the data storage area, a target escape space with a capacity greater than or equal to the capacity of the data storage area can be applied for. On the contrary, when the number of non-faulty hard disks is less than the number of data blocks in the data storage area, a target escape space with a capacity less than the capacity of the data storage area can be applied for. Optionally, when the number of non-faulty hard disks is greater than the number of data blocks in the data storage area, the computer device can apply for a target escape space with a capacity equal to the capacity of the data storage area based on the non-faulty hard disks with the top performance.
[0087] S303, if not, apply for a target escape space according to the storage information of the data storage area.
[0088] In the embodiments of the present application, when the computer device determines that there is a faulty hard disk in the storage system based on the above steps, this situation indicates that the number of hard disks in the current hard disk group is relatively large, and there is enough space to apply for a target escape space, and it is not necessary to perform a reduction operation on the desired escape space. Then, the computer device can apply for a target escape space with a capacity greater than or equal to the capacity of the data storage area. Optionally, the computer device can also apply for a target escape space with a capacity equal to the capacity of the data storage area based on the hard disks with the top performance.
[0089] The method described in the embodiments of the present application adopts different application strategies for the fault situation of the storage hard disks in the storage system, which can enable the storage system to better adapt to various changes. Whether it is a sudden hard disk failure or a change in daily data storage requirements, it can quickly respond to ensure the normal operation of the storage system and improve the reliability of the storage system.
[0090] In some embodiments, a specific implementation manner for applying for a target escape space is also provided, such as Figure 6 shown, the "applying for a target escape space according to the hard disk information of non-faulty hard disks in the storage system and the storage information of the data storage area" in S302 above includes:
[0091] S401, determine whether there are redundant hard disks among the non-faulty hard disks according to the hard disk information of non-faulty hard disks in the storage system.
[0092] Among them, the redundant hard disk is an additional hard disk set in the storage system to improve data reliability and availability. When a working hard disk in the storage system fails, the redundant hard disk can quickly take over its work, ensuring data integrity and normal system operation, and avoiding data loss and service interruption caused by hard disk failures.
[0093] In the embodiments of the present application, when the computer device determines that there is a faulty hard disk in the storage system, it can further determine whether there is a redundant hard disk among the non-faulty hard disks according to the hard disk information of the non-faulty hard disks in the storage system. For example, in terms of disk-level redundancy, the stripe column width of the hard disk group is = min(number of member disks in the hard disk group – redundant column 1, 25). When the number of hard disks in the hard disk group is greater than 25, it is considered that there is a redundant hard disk among the non-faulty hard disks. Since there are more hard disks in the hard disk group, after a hard disk failure, it does not affect the width of the stripe when selecting a disk, and column reduction does not need to be considered. When the number of hard disks in the hard disk group is not greater than 25, it is considered that there is no redundant hard disk among the non-faulty hard disks. To ensure the continuity of the garbage escape service, column reduction of the stripe is required.
[0094] S402, if not, apply for a target escape space according to the storage information of the data storage area and the hard disk information of the non-faulty hard disks.
[0095] In the embodiments of the present application, if the computer device determines based on the above steps that there is no redundant hard disk in the storage system, this situation indicates that column reduction operation needs to be performed on the desired escape space. Specifically, the desired escape space can be column-reduced according to the number of non-faulty hard disks, the number of data columns, and the number of parity columns in the data storage area to obtain the target escape space.
[0096] The method described in the embodiments of the present application can adopt different application strategies according to whether there is a redundant hard disk among the non-faulty hard disks when applying for a target escape space in the case of a faulty hard disk in the storage system, enabling the storage system to better adapt to different resource conditions. Whether the storage system resources are sufficient or relatively scarce, corresponding strategies can be used to meet the requirements of data storage and management. When there is a redundant hard disk in the storage system, it indicates that the current resources are relatively sufficient, and at this time, the required target escape space can be directly applied for. When there is no redundant hard disk in the storage system, it indicates that the current resources are relatively scarce, and at this time, column reduction processing needs to be performed on the required target escape space to achieve efficient data recovery.
[0097] In some embodiments, a specific implementation manner for applying for a target escape space is also provided, such as Figure 7 shown, "apply for a target escape space according to the storage information of the data storage area and the hard disk information of the non-faulty hard disks" in the above S402 includes:
[0098] S501. Determine a third quantity of data columns and a fourth quantity of parity columns included in the target escape space according to the first quantity of data columns and the second quantity of parity columns included in the data storage area, and the hard disk information of the non-failed hard disks.
[0099] Among them, the first quantity is the quantity of data columns in the data storage area to be recycled (i.e., the strip to be recycled). The second quantity is the quantity of parity columns in the data storage area to be recycled (i.e., the strip to be recycled). The third quantity is the quantity of data columns in the target escape space (i.e., the new strip), and the fourth quantity is the quantity of parity columns in the target escape space (i.e., the new strip).
[0100] In the embodiments of the present application, when the computer device determines that there is no redundant hard disk in the storage system, it performs a column reduction operation on the desired escape space. Specifically, it can first determine the quantity of data columns included in the data storage area and use it as the first quantity, and determine the quantity of parity columns included in the data storage area and use it as the second quantity. Then, according to the first quantity, the second quantity, and the quantity of non-failed hard disks of the non-failed hard disks, determine the third quantity of data columns and the fourth quantity of parity columns included in the target escape space.
[0101] S502. Apply for the target escape space according to the third quantity and the fourth quantity.
[0102] In the embodiments of the present application, after the computer device determines the third quantity of data columns and the fourth quantity of parity columns included in the target escape space based on the above steps, it can determine the sum of the third quantity and the fourth quantity, and then apply for the target escape space corresponding to the sum of the quantities, so as to obtain the target escape space with the sum of the quantities of data columns and parity columns being the sum of the quantities.
[0103] The method described in the embodiments of the present application determines the third quantity of data columns and the fourth quantity of parity columns of the target escape space according to the first quantity of data columns, the second quantity of parity columns in the data storage area, and the hard disk information of the non-failed hard disks, which can ensure that the layout of the target escape space is adapted to the layout of the data storage area, avoid applying for an overly large escape space, and thus reduce unnecessary resource occupation.
[0104] In some embodiments, a specific implementation manner for determining the third quantity of data columns and the fourth quantity of parity columns included in the target escape space is also provided, as Figure 8 shown, the "determine the third quantity of data columns and the fourth quantity of parity columns included in the target escape space according to the first quantity of data columns and the second quantity of parity columns included in the data storage area, and the hard disk information of the non-failed hard disks" in S501 above includes:
[0105] S601. Determine the number of columns to be reduced based on the first number of data columns and the second number of parity columns included in the data storage area, as well as the hard disk information of the non-failed hard disks.
[0106] Among them, the number of columns to be reduced is the number of columns that need to be reduced for the desired escape space.
[0107] In the embodiments of the present application, after the computer device obtains the first number of data columns and the second number of parity columns included in the data storage area, it can determine the sum of the first number and the second number, and then determine the number of non-failed hard disks based on the hard disk information of the non-failed hard disks. Then, calculate the difference between the sum of the first number and the second number and the number of non-failed hard disks to obtain the number of columns to be reduced. For example, if the sum of the first number of data columns and the second number of parity columns included in the data storage area is 6 and the number of non-failed hard disks is 4, then the number of columns to be reduced is 2.
[0108] S602. Determine the type of columns to be reduced based on the second number of parity columns.
[0109] Among them, the type of columns to be reduced includes parity column type and data column type.
[0110] In the embodiments of the present application, after the computer device obtains the second number of parity columns based on the above steps, it can determine the type of columns to be reduced according to the second number of parity columns and the number of columns to be reduced. For example, if the second number is greater than the number of columns to be reduced, then determine the type of columns to be reduced as the parity column type; if the second number is not greater than the number of columns to be reduced, then determine the type of columns to be reduced as the data column type.
[0111] Optionally, as Figure 9 shown, the "determine the type of columns to be reduced according to the second number of parity columns" in S602 above includes:
[0112] S6021. If the second number of parity columns is greater than a preset value, then determine the type of columns to be reduced as the parity column type.
[0113] Among them, the preset value can be set according to actual needs. For example, the preset value is 1, that is, the number of parity columns is at least 1.
[0114] In the embodiments of the present application, after the computer device obtains the second number of parity columns based on the above steps, it can determine the type of columns to be reduced according to the second number of parity columns and the preset value. If the computer device determines that the second number of parity columns is greater than the preset value, then determine the type of columns to be reduced as the parity column type.
[0115] S6022. If the second number of parity columns is not greater than the preset value, then determine the type of columns to be reduced as the data column type.
[0116] In the embodiments of the present application, if the computer device determines that the second quantity of the parity columns is not greater than a preset value, it determines that the reduction column type is the parity column type. It can be understood that, in order to deal with the problem of stripe reduction that may occur during the garbage escape process, if the data columns are reduced first, the utilization rate of the selected stripe space will decrease (that is, the data columns used to store user data in the stripe are relatively reduced compared to the stripe width), which may offset or even exceed the space released by garbage collection, resulting in the used space not decreasing but increasing instead. For example, the stripe width is 10, the data column size is 8, and the parity column size is 2; the proportion of the data column in the stripe is 1 / 5. At this time, if column reduction is performed and the data column is reduced first, the stripe width is 9, the data column size is 7, and the P parity column is 2. The proportion of the data column in the stripe is 7 / 9 < 8 / 10, which means that the utilization rate of the stripe space decreases. Therefore, in order to ensure the utilization rate of the stripe space, when stripe reduction is required, the parity columns are reduced first. When the parity columns are reduced to the preset value, in order to ensure data security, the data columns have to be reduced at this time. Then the general principle is that during the garbage escape process and when a disk failure occurs, the parity columns are reduced first. After the parity columns are reduced to the preset value, the data columns are reduced to ensure that after garbage escape and a hard disk failure, new stripes for carrying valid data can still be allocated for garbage escape, ensuring that garbage can still operate normally.
[0117] S603. Determine a third quantity of data columns and a fourth quantity of parity columns included in the target escape space according to the reduction column type and the reduction column quantity.
[0118] In the embodiments of the present application, after the computer device determines the reduction column type and the reduction column quantity based on the above steps, it can determine a third quantity of data columns and a fourth quantity of parity columns included in the target escape space according to the reduction column type and the reduction column quantity corresponding to the reduction column type. Specifically, the third quantity and the fourth quantity of parity columns included in the target escape space can be determined according to the quantity of data columns in the expected escape space, the quantity of parity columns in the expected escape space, the reduction column type, and the reduction column quantity corresponding to the reduction column type. Specifically, the third quantity is equal to the quantity of data columns in the expected escape space minus the reduction column quantity corresponding to the data column type of the reduction column type, and the fourth quantity is equal to the quantity of parity columns in the expected escape space minus the reduction column quantity corresponding to the parity column type of the reduction column type. For example, the quantity of data columns in the expected escape space is 4, the quantity of parity columns is 2, the reduction column types are the parity column type and the data column type, and the reduction column quantity corresponding to the parity column type is 1, and the reduction column quantity corresponding to the data column type is 2. Then it can be determined that the third quantity of data columns in the target escape space is 4 - 2 = 2, and the fourth quantity of parity columns is 2 - 1 = 1.
[0119] Such as Figure 10As shown in the figure, when recycling data in the storage space composed of a hard disk group including disk1 - disk6, stripe 1, stripe 2, and stripe 3 in the figure are the stripes to be recycled. The number of data columns in the stripes to be recycled is 4, which are respectively distributed on hard disks disk1 - disk4, and the number of verification columns is 2, which are respectively distributed on disk5 - disk6. Then the number of data columns in the desired escape space is 4 and the number of verification columns is 2. However, during the process of garbage escape, hard disk disk4 fails. At this time, a column reduction operation needs to be performed on the desired escape space. At this time, the number of faulty hard disks is 1. Since the number of verification columns in the desired escape space is 2, which is greater than the preset value of 1, the verification columns of the desired escape space can be reduced first. The number of verification columns in the desired escape space is reduced to 2 - 1 = 1, and we get Figure 10 stripe 4 in (the number of verification columns is 1, and the data column data is 4). When using stripe 4 for data recycling (this data recycling process involves data inheritance and data copy), hard disk disk5 fails again. Then at this time, a column reduction operation needs to be performed on stripe 4. At this time, the number of faulty hard disks is 1. Since the number of verification columns in stripe 4 is 1, which is not greater than the preset value of 1, the data columns have to be reduced. The number of data columns in stripe 4 can be reduced to 4 - 1 = 3, and we get Figure 10 stripe 5 in (the number of verification columns is 1, and the data column data is 3).
[0120] In the embodiments of the present application, for the method described in the embodiments of the present application, when there is no redundant hard disk in the storage system, first perform a verification column reduction operation on the required target escape space, and then perform a data column reduction operation. This can ensure that the data column ratio in the target escape space reaches the maximum, thereby maximizing the utilization rate of the target escape space and further improving the data recycling efficiency.
[0121] In some embodiments, the above data recycling method further includes: detecting whether the storage hard disks in the storage system meet the preset update conditions, and when the storage hard disks in the storage system meet the preset update conditions, perform a regional extension operation on the target escape space after data recycling according to the storage information in the data storage area.
[0122] Among them, the preset update conditions include at least one of the available capacity of the storage system reaching the preset capacity threshold, the faulty hard disk being restored, and a new storage hard disk being added to the storage system. The preset capacity threshold can be determined according to actual storage requirements. For example, the used capacity of the storage space is less than 90%. The regional extension operation means performing a column extension operation on the verification columns and / or data columns in the target escape space after data recycling.
[0123] In the embodiments of the present application, during the process of data recovery of a computer device, the available capacity of the storage system in the storage system is monitored in real time. When the available capacity of the storage system reaches a preset capacity threshold, this situation indicates the end of garbage escape, that is, the storage hard disks in the storage system meet the preset update conditions. At this time, the target escape space can be closed, the front-end service can be restored, and all background tasks can be restored to make the system return to the previous normal state. Then, the computer device can perform a column expansion operation on the parity column and / or data column in the target escape space after data recovery according to the situation of the hard disks in the storage space. Specifically, it can first determine whether the faulty hard disk in the storage system has been restored and whether a new storage hard disk has been added to the storage system, which specifically includes two scenarios. The first scenario is that the faulty hard disk has not been restored and no new storage hard disk has been added to the storage system. Then, a column expansion operation can be performed on the parity column of the target escape space after data recovery, expanding the number of parity columns to the number of parity columns in the data storage area, and then performing a balancing operation on other data columns. The second scenario is that the faulty hard disk has been restored and / or a new storage hard disk has been added to the storage system. Then, it can further determine the sum of the number of hard disks with faulty recovery and the number of newly added storage hard disks, and compare this sum with the number of columns reduced in the target escape space after data recovery. If this sum is less than the number of columns reduced, first expand the parity column in the target escape space after data recovery. When the number of parity columns reaches the number of parity columns in the data storage area, then expand the data column in the target escape space after data recovery. If this sum is not less than the number of columns reduced, expand the number of parity columns and data columns in the target escape space after data recovery to the number of parity columns and data columns in the data storage area respectively. As Figure 11 shown, Figure 11 It includes three figures a, b, and c. Strip 1 and Strip 2 in Figure a are the target escape spaces (i.e., new strips) applied during data recovery. When using Strip 1 and Strip 2 for data recovery, hard disks Disk7 and Disk8 failed. At this time, it was necessary to reduce the columns of Strip 1 and Strip 2. Using the above column reduction method, the data columns and parity columns in Strip 1 and Strip 2 were each reduced by 1 column, forming Strip 3, Strip 4, and Strip 5, that is, the target escape space after data recovery was obtained. Figure b shows the first scenario described above, and Figure c shows the second scenario described above.
[0124] The method described in the embodiments of the present application can perform a regional extension operation on the target escape space when the faulty hard disk in the storage system is restored or a new hard disk is added, which can make the distribution of data on each hard disk more balanced, give full play to the storage capacity of each hard disk, and thus improve the resource utilization efficiency of the entire storage system. At the same time, it can avoid the situation where some hard disks are overused while other hard disks are idle, which helps to extend the service life of the hard disks and reduce the storage cost.
[0125] In combination with all of the above embodiments, a data recovery method is further provided. As Figure 12 shown, the method includes:
[0126] S701. In a scenario where the use of the escape space for data recovery is triggered, obtain the current storage hard disk information and the storage information of the data storage area in the storage system, and determine whether there is a faulty hard disk in the storage system according to the current hard disk information.
[0127] S702. If there is a faulty hard disk, determine whether there is a redundant hard disk among the non-faulty hard disks in the storage system according to the hard disk information of the non-faulty hard disks in the storage system.
[0128] S703. If there is a redundant hard disk among the non-faulty hard disks, apply for a target escape space according to the storage information of the data storage area.
[0129] S704. If there is no redundant hard disk among the non-faulty hard disks, determine the reduced column quantity according to the first quantity of the data columns and the second quantity of the parity columns included in the data storage area, as well as the hard disk information of the non-faulty hard disks.
[0130] S705. If the second quantity of the parity columns is greater than a preset value, determine that the reduced column type is the parity column type.
[0131] S706. If the second quantity of the parity columns is not greater than the preset value, determine that the reduced column type is the data column type.
[0132] S707. According to the reduced column type and the reduced column quantity, determine the third quantity of the data columns and the fourth quantity of the parity columns included in the target escape space, and apply for the target escape space according to the third quantity and the fourth quantity.
[0133] S708. If there is no faulty hard disk, apply for a target escape space according to the storage information of the data storage area.
[0134] S709. Recover the data in the data storage area according to the target escape space.
[0135] S710. Detect whether the storage hard disks in the storage system meet the preset update conditions, and perform a regional extension operation on the target escape space after data recovery according to the storage information of the data storage area when the storage hard disks in the storage system meet the preset update conditions.
[0136] In the embodiments of the present application, Problem 1 existing in the prior art: For the reservation of the GC escape space, it is reserved according to the storage pool. When the capacities are uneven, after the space of a single disk is full, when using the GC escape space, it may be impossible to select a dSeg (strip) for GC because the capacity of some disks is full; or more GC escape space is reserved for effective disks, and less is reserved for some disks. After a disk fails, the escape space may become unavailable. Problem 2 existing in the prior art: When reusing the garbage collection space for GC transfer, there are many scenarios where strip contraction occurs, such as disk failure, an increase in the number of hot spare disks, a change in the RAID level, and strip migration of the storage pool across hard disk groups. Most of the prior art chooses not to perform contraction during GC escape, and notifies the user of the above operations, such as immediately inserting a new disk after a disk failure to continue GC escape, the number of hot spare disks cannot increase, and the RAID level is not allowed to change at this time, etc. Various restrictions are imposed to enable GC to continue. If contraction has to be performed, GC escape is terminated. But the disadvantage is that it is not conducive to business continuity. A good storage system should be able to handle various scenarios and various failures. Problem 3 existing in the prior art: When performing contraction during GC escape, the data column D of the strip is preferentially contracted; this will result in a decrease in the space utilization rate of the data columns in the strip, which may offset or even exceed the space released by garbage collection, resulting in the used space not decreasing but increasing instead.
[0137] To solve Problem 1 existing in the above prior art, a certain amount of escape space can be reserved for each hard disk under the storage pool. Even when the capacities are uneven and the available capacity of some hard disks is full, it is still ensured that a single disk can definitely select data blocks and apply for strips using the escape space in the escape mode. To solve Problem 2 and Problem 3 existing in the above prior art, when reusing the garbage collection space for GC transfer, there are many scenarios where strip contraction occurs. For example: At this time, the SSD hard disk capacity is almost full, and the probability of a hard disk failure in this extreme scenario is relatively high. In the prior art, when using the garbage collection escape space for garbage collection, in the case of a hard disk failure, the strips of the storage pool are not contracted. What is adopted is that the user must replace the faulty disk, replace the disk in the same slot, restore the data of the faulty disk, and then continue the garbage collection escape; if the user does not replace the faulty disk or does not insert a new disk into the system, then GC escape cannot continue, and business continuity cannot be guaranteed; Another example is that the number of hot spare disks increases (one or several hard disks pre-configured in the storage system, which do not store user data under normal circumstances but are in a standby state, meaning that the hard disks available for gc to allocate strips in the storage pool decrease), a change in the RAID level, and strip migration of the storage pool across hard disk groups during garbage collection (relocating data to a hard disk group with a smaller strip width, resulting in contraction).
[0138] In order to address the problem of strip column reduction that occurs during the gc escape process, disk selection optimization for strip column reduction during escape can be carried out. In terms of disk-level redundancy, the strip column width member disk count of the disk group is = min(number of member disks in the disk group – redundant column 1, 25). When the number of disks in the disk group is greater than 25, since there are more disks in the disk group, after a disk failure at this time, it does not affect the width of the strip during disk selection, and column reduction does not need to be considered. However, when the number of disks in the disk group is less than 25, in order to ensure the continuity of the gc escape service, strip column reduction is required. However, during the use of the escape space, if the data column is reduced first, it will lead to a decrease in the space utilization rate of the selected strip (that is, the data column used to store user data in the strip is relatively reduced compared to the strip width), which may offset or even exceed the space released by garbage collection, resulting in the used space not decreasing but increasing instead. For example, the strip width is 10, the size of the D data column is 8, and the size of the P parity column is 2; the proportion of D in the strip is 1 / 5. At this time, if column reduction is performed and D is reduced first, the strip width is 9, the size of D is 7, and the P parity column is 2. The proportion of D in the strip is 7 / 9 < 8 / 10, which means the space utilization rate of the strip has decreased. As Figure 4 shown, reduce P first. When P is reduced to 1, in order to ensure data security, D has to be reduced at this time. So the general principle is that during the gc escape process, when a disk failure occurs, reduce the parity column P first. After P is reduced to 1, then reduce the data column D. Such a strip selection strategy can ensure that after a disk failure during GC escape, a new strip for carrying valid data can still be allocated for GC escape, ensuring that GC can still operate normally. The general principle is that when using the escape space, try not to perform column reduction. If column reduction is required, reduce the parity column P first, and then reduce the data column D to ensure that the space utilization rate will not decrease. As Figure 4As shown in the figure, when it is necessary to reduce the stripe during the GC escape process, first reduce P, from P = 2 to P = 1. When the used capacity of the storage pool during GC escape drops to a certain level, the escape space is closed at this time. The front-end business IO is restored, all background tasks are restored, and the system returns to the previous normal state. If there is a disk failure during the GC escape, the stripe is repaired at this time, and the user data on the failed disk is restored to other stripes. In addition, since the parity column P in the stripe is 1, and since the RAID level set by the user at the beginning is 2, in order to restore the data reliability redundancy ability of the stripe, after the GC escape ends, it is necessary to actively perform stripe repair (for example, repair the stripe from P = 1 to P = 2 before). Without changing the stripe width, expand the width parity column from 1 to 2. In order to reduce the number of data migrations and improve the balance efficiency, a combination of inheritance + copy is used. If the data blocks of two stripes come from the same hard disk, this data block needs to be inherited. For example, in the figure, by modifying the mapping of the d13 data block to stripe 3 to stripe 4, and the mapping of the previous stripe 4 to stripe 3 at this position to achieve the inheritance effect. If the two data blocks do not come from the same hard disk, then by copying, copy the data of d12 to d17. The d16 data block on stripe 3 later participates in the migration of the next stripe because stripe 4 is full. If there is a new disk expansion later, then actively perform stripe expansion and balance, and the method of stripe expansion and balance also uses the above combination of inheritance + copy, and the logic is similar.
[0139] In the method described in the embodiments of the present application, when the user-visible space in the all-flash array storage system is exhausted and the OP space begins to be used, and the garbage collection is not timely, resulting in the exhaustion of the OP space, the GC escape space reserved in advance on each disk in the storage pool is used to ensure that new stripes can be continuously provided for GC transfer, and ensure that the garbage collection can still continue to work. There are many scenarios in the GC escape that will cause stripe reduction. The embodiments of the present application allow the stripe to be reduced. In order to ensure that the space utilization rate of the stripe selected by the disk selection algorithm decreases as much as possible, first reduce the parity column P, and then reduce the data column D to ensure that the used space will not increase instead of decrease. After the GC escape is completed, if there is no disk expansion or new disk, then the P of the stripe that was previously reduced can be expanded without changing the stripe width to restore the previous redundancy level of the stripe. At the same time, perform stripe balance for each hard disk group in the storage; if there is disk expansion or new disk, then perform expansion balance.
[0140] The methods described in the above steps are all described in the foregoing embodiments. For detailed content, please refer to the foregoing description and will not be repeated here.
[0141] It should be understood that although each step in the flowcharts involved in the above-described embodiments is shown in sequence according to the indication of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear description in this article, there is no strict order limit for the execution of these steps, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same moment, but can be executed at different moments. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.
[0142] Based on the same inventive concept, an embodiment of the present application also provides a data recovery device for implementing the above-mentioned data recovery method. The solution provided by this device to solve the problem is similar to the solution described in the above method. Therefore, the specific limitations in one or more embodiments of the following data recovery devices can refer to the limitations on the data recovery method in the above text, and will not be repeated here.
[0143] In some embodiments, as Figure 13 shown, a data recovery device is provided, including:
[0144] An application module 11, configured to obtain the storage hard disk information and the storage information of the data storage area in the storage system in a scenario where the use of the escape space for data recovery is triggered, and apply for a target escape space according to the hard disk information and the storage information of the data storage area.
[0145] A recovery module 12, configured to perform data recovery on the data in the data storage area according to the target escape space.
[0146] In some embodiments, the above application module includes:
[0147] A determination unit, configured to determine whether there is a faulty hard disk in the storage system according to the current hard disk information.
[0148] A first application unit, configured to, if there is, apply for a target escape space according to the hard disk information of the non-faulty hard disks in the storage system and the storage information of the data storage area.
[0149] A second application unit, configured to, if not, apply for a target escape space according to the storage information of the data storage area.
[0150] In some embodiments, the above first application unit includes:
[0151] A determination subunit, configured to determine whether there is a redundant hard disk among the non-faulty hard disks according to the hard disk information of the non-faulty hard disks in the storage system.
[0152] A first application subunit, configured to, if any exists, apply for a target escape space according to the storage information of the data storage area.
[0153] A second application subunit, configured to, if none exists, apply for a target escape space according to the storage information of the data storage area and the hard disk information of the non-faulty hard disks.
[0154] In some embodiments, the above-mentioned second application subunit is specifically configured to determine a third quantity of data columns and a fourth quantity of parity columns included in the target escape space according to a first quantity of data columns and a second quantity of parity columns included in the data storage area, and the hard disk information of the non-faulty hard disks; and apply for the target escape space according to the third quantity and the fourth quantity.
[0155] In some embodiments, the above-mentioned second application subunit is specifically configured to determine a reduced column quantity according to a first quantity of data columns and a second quantity of parity columns included in the data storage area, and the hard disk information of the non-faulty hard disks; determine a reduced column type according to the second quantity of parity columns; and determine a third quantity of data columns and a fourth quantity of parity columns included in the target escape space according to the reduced column type and the reduced column quantity.
[0156] In some embodiments, the above-mentioned second application subunit is specifically configured to, if the second quantity of parity columns is greater than a preset value, determine that the reduced column type is a parity column type; and if the second quantity of parity columns is not greater than the preset value, determine that the reduced column type is a data column type.
[0157] In some embodiments, the above-mentioned data recovery device is further configured to detect whether the storage hard disks in the storage system meet a preset update condition, and perform a regional extension operation on the target escape space after data recovery according to the storage information of the data storage area when the storage hard disks in the storage system meet the preset update condition.
[0158] Each module in the above-mentioned data recovery device can be implemented in whole or in part by software, hardware, and their combination. The above-mentioned modules can be embedded in the processor in the computer device in hardware form or be independent of it, or can be stored in the memory in the computer device in software form, so that the processor can call and execute the operations corresponding to the above-mentioned modules.
[0159] In some embodiments, a computer device is provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps of the data recovery method described in any of the above embodiments are implemented.
[0160] In some embodiments, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of the data recovery method described in any of the above embodiments are implemented.
[0161] In some embodiments, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the steps of the data recovery method described in any of the above embodiments are implemented.
[0162] Those of ordinary skill in the art can understand that all or part of the processes of the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memories can include read-only memory (ROM), magnetic tapes, floppy disks, flash memories, optical memories, high-density embedded non-volatile memories, resistive random access memories (ReRAM), magnetoresistive random access memories (MRAM), ferroelectric random access memories (FRAM), phase change memories (PCM), graphene memories, etc. Volatile memories can include random access memory (RAM) or external cache memories, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processing units, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.
[0163] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0164] The above-described embodiments merely represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.
Claims
1. A data recovery method, characterized in that, The method includes: In a scenario where data recovery using an escape space is triggered, obtain the current storage hard disk information and the storage information of the data storage area in the storage system, and apply for a target escape space according to the hard disk information and the storage information of the data storage area; Perform data recovery on the data in the data storage area according to the target escape space.
2. The method according to claim 1, characterized in that, The applying for a target escape space according to the current hard disk information and the storage information of the data storage area includes: Determine whether there are faulty hard disks in the storage system according to the current hard disk information; If there are, apply for the target escape space according to the hard disk information of the non-faulty hard disks in the storage system and the storage information of the data storage area; If not, apply for the target escape space according to the storage information of the data storage area.
3. The method according to claim 2, wherein The applying for the target escape space according to the hard disk information of the non-faulty hard disks in the storage system and the storage information of the data storage area includes: Determine whether there are redundant hard disks among the non-faulty hard disks according to the hard disk information of the non-faulty hard disks in the storage system; If there are, apply for the target escape space according to the storage information of the data storage area; If not, apply for the target escape space according to the storage information of the data storage area and the hard disk information of the non-faulty hard disks.
4. The method according to claim 3, wherein The applying for the target escape space according to the storage information of the data storage area and the hard disk information of the non-faulty hard disks includes: Determine a third quantity of data columns and a fourth quantity of parity columns included in the target escape space according to a first quantity of data columns and a second quantity of parity columns included in the data storage area, and the hard disk information of the non-faulty hard disks; Apply for the target escape space according to the third quantity and the fourth quantity.
5. The method according to claim 4, wherein The determining a third quantity of data columns and a fourth quantity of parity columns included in the target escape space according to a first quantity of data columns and a second quantity of parity columns included in the data storage area, and the hard disk information of the non-faulty hard disks includes: Determine a reduced column quantity according to a first quantity of data columns and a second quantity of parity columns included in the data storage area, and the hard disk information of the non-faulty hard disks; Determine a reduced column type according to the second quantity of parity columns; Determine a third quantity of data columns and a fourth quantity of parity columns included in the target escape space according to the reduced column type and the reduced column quantity.
6. The method according to claim 5, wherein The determining a reduced column type according to the second quantity of parity columns includes: If the second quantity of parity columns is greater than a preset value, determine that the reduced column type is a parity column type; If the second quantity of parity columns is not greater than the preset value, determine that the reduced column type is a data column type.
7. The method according to any one of claims 1-6, characterized in that, The method further includes: Detect whether the storage hard disks in the storage system meet a preset update condition, and perform an area extension operation on the target escape space after data recovery according to the storage information of the data storage area when the storage hard disks in the storage system meet the preset update condition.
8. A data recovery device, characterized in that, The device includes: An application module, configured to obtain storage hard disk information and storage information of a data storage area in a storage system in a scenario where triggering the use of an escape space for data recovery, and apply for a target escape space according to the hard disk information and the storage information of the data storage area; A recovery module, configured to perform data recovery on the data in the data storage area according to the target escape space.
9. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.