A method, system, device and medium for data reconstruction of a distributed file system
The method allows for precise control of data reconstruction in distributed file systems by prioritizing storage pools, ensuring high-priority pools are reconstructed first and adjusting based on priority changes, optimizing data recovery efficiency.
Patent Information
- Application Number
- CN202111336172.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-11
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2041-11-11
AI Technical Summary
In the prior art, in the process of data reconstruction between storage pools, distributed file systems cannot perform refined control according to the order of priority, resulting in unreasonable allocation of storage resources.
Set a data reconstruction level for the storage pool, and record the storage pool where the data object is located when it needs to be reconstructed. By obtaining the highest priority storage pool for resource reservation and data reconstruction, ensure that the high priority storage pool is reconstructed first, and the low priority storage pool is reconstructed after the high priority storage pool reconstruction is completed.
The data reconstruction of the storage pool in the distributed file system is implemented in priority order, and the refined control of data reconstruction is realized, which improves the utilization efficiency of storage resources and the efficiency of data recovery.
Smart Images

Figure CN114168556B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of distributed file storage, and particularly relates to a method, system, device and medium for reconstructing data in a distributed file system. Background Art
[0002] A distributed file system describes storage devices hierarchically according to data center - rack - host - disk. File data is split into storage objects and stored on different disks through the CRUSH algorithm. Among them, CRUSH: (Controlled, Scalable, Decentralized Placement of Replicated Data, a controlled, scalable, and distributed replicated data placement algorithm). Each disk is managed by an OSD service to implement functions such as data reading and writing and data recovery. Among them, OSD: (Object - based Storage Device, object storage device). The file system pools and manages all resources, and each storage pool has different data management policies such as replication policies, failure domains, and parity policies. To achieve policy isolation between storage pools, the upper - layer data is not directly mapped to disks, but PG is introduced to implement two - level mapping. Among them, PG (Placement Group, a carrier for placing objects). Data objects are first mapped to PG through a pseudo - random hash function to achieve the first - level mapping; PG is mapped to OSD through a pseudo - random hash function to obtain a list of OSD members, including a primary OSD and the remaining secondary OSDs.
[0003] When a storage device in the cluster fails, such as a disk being damaged, a host powering off, a network card failure, etc., causing one or more OSDs to be unavailable, the OSD members mapped by PG will change, and data needs to be written to the newly added OSD through data recovery. Or when a failed OSD rejoins the cluster, the version of the data on the disk lags behind the replicas on other disks, and data recovery is also required to update the data to the latest. Data recovery is carried out in units of PG. There is a PG instance on each member OSD of each PG. The PG on the primary OSD is called the primary PG, and the PG on the secondary OSD is called the secondary PG. The primary PG controls the data recovery process. PG manages different states through a state machine. In the prior art, a data recovery priority is set for the storage pool, but the priority only affects the speed of data reconstruction. The higher the priority, the larger the amount of data that can be reconstructed per unit time. If there are several storage pools sharing storage resources, the storage pools cannot determine the order of data reconstruction according to the priority level. Summary of the Invention
[0004] To solve the above technical problems, the present invention proposes a method, system, device and medium for data reconstruction in a distributed file system, which can enable the storage pool of the distributed file system to perform data reconstruction according to the order of priority levels, realizing fine-grained control of data reconstruction.
[0005] To achieve the above object, the present invention adopts the following technical solutions:
[0006] A method for data reconstruction in a distributed file system, comprising the following steps:
[0007] Set a data reconstruction level for the storage pool, and record the storage pool where the data object is located when the data object needs to be reconstructed;
[0008] On the premise that the data recovery priority in the placement group is not to stop data recovery, obtain the storage pool with the highest priority. When the priority of the currently running storage pool is not less than the level of the highest priority storage pool, perform resource reservation;
[0009] After reserving resources, perform data reconstruction on the premise that the data reconstruction priority in the placement group is not less than the data priority in the storage pool.
[0010] Further, the method further includes: before the placement group performs resource reservation, if the data recovery priority in the placement group is to stop data recovery, or the data recovery priority in the placement group is not to stop data recovery, but the priority of the currently running storage pool is less than the level of the highest priority storage pool, then no resource reservation is performed and it enters a temporary state.
[0011] Further, the method further includes: after reserving resources, if it is checked that the data reconstruction priority in the placement group is less than the data priority in the storage pool or the data reconstruction priority in the placement group is to stop data reconstruction, then stop data reconstruction and enter a temporary state.
[0012] Further, the method for setting the data reconstruction level for the storage pool is:
[0013] Set the data reconstruction level for the storage pool through the command line or the interface; the data reconstruction levels include recovery priority, adaptive, read / write priority, stop data recovery.
[0014] Further, the step of recording the storage pool where the data object is located when the data object needs to be reconstructed includes:
[0015] When the placement group state machine enters the working state, if data reconstruction is required according to the difference in the placement group log or data reconstruction is required according to the full scan result of the object, then record the storage pool where the data object is located.
[0016] Further, after data reconstruction is completed for all placement groups in the storage pool, it is deleted from the storage pool where the recorded data object is located.
[0017] The present invention also provides a system for distributed file system data reconstruction, including a setting module, a resource reservation module, and a data reconstruction module;
[0018] The setting module is used to set the data reconstruction level for the storage pool and record the storage pool where the data object is located when the data object needs to be reconstructed;
[0019] The resource reservation module is used to obtain the highest-priority storage pool on the premise that the data recovery priority in the placement group is not to stop data recovery. When the priority of the currently running storage pool is not less than the level of the highest-priority storage pool, resource reservation is performed;
[0020] The data reconstruction module is used to perform data reconstruction after resources are reserved on the premise that the data reconstruction priority in the placement group is not less than the data priority in the storage pool.
[0021] Further, the system further includes a deletion module;
[0022] The deletion module is used to delete from the storage pool where the recorded data object is located after data reconstruction is completed for all placement groups in the storage pool.
[0023] The present invention also provides a device, including:
[0024] A memory for storing a computer program;
[0025] A processor for implementing the method steps when executing the computer program.
[0026] The present invention also provides a readable storage medium, on which a computer program is stored, and the computer program, when executed by a processor, implements the method steps as described above.
[0027] The effects provided in the summary of the invention are only the effects of the embodiments, rather than all the effects of the invention. One of the above technical solutions has the following advantages or beneficial effects:
[0028] The present invention provides a method, system, device, and medium for data reconstruction in a distributed file system. The method includes setting a data reconstruction level for a storage pool and recording the storage pool where a data object is located when the data object needs to be reconstructed; obtaining the highest-priority storage pool on the premise that the data recovery priority in a placement group is not to stop data recovery, and making a resource reservation when the priority of the currently running storage pool is not less than the level of the highest-priority storage pool; after obtaining the reserved resources, performing data reconstruction on the premise that the data reconstruction priority in the placement group is not less than the data priority in the storage pool. Based on a method for data reconstruction in a distributed file system, a system, device, and storage medium for data reconstruction in a distributed file system are also provided. In the present invention, if a low-priority storage pool starts data reconstruction first and a high-priority storage pool starts reconstruction later, the data reconstruction of the low-priority storage pool can be stopped and the data reconstruction of the high-priority storage pool can be started immediately. After the data reconstruction of the high-priority storage pool is completed, the low-priority storage pool can start reconstruction. If the priority of the storage pool is modified during the data reconstruction process, the order of data reconstruction can be determined immediately according to the modified priority. If the data reconstruction priority is set to stop data recovery, the storage pool does not perform data reconstruction, and the storage pool that is currently performing data reconstruction stops data reconstruction immediately after being set to stop data recovery. The storage pool of the present invention can perform data reconstruction in the order of the priority levels. Users can set the reconstruction priority for the storage pool according to their needs, and the storage pool with a higher priority performs data reconstruction first, realizing fine-grained control of data reconstruction. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Such as Figure 1 is a flowchart of a method for data reconstruction in a distributed file system according to Embodiment 1 of the present invention;
[0030] Such as Figure 2 is a schematic diagram of a system for data reconstruction in a distributed file system according to Embodiment 2 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0031] To clearly illustrate the technical features of the present solution, the present invention will be described in detail below through specific embodiments and in conjunction with their accompanying drawings. The following disclosure provides many different embodiments or examples for implementing different structures of the present invention. To simplify the disclosure of the present invention, the components and settings of specific examples are described below. In addition, the present invention may repeat reference numerals and / or letters in different examples. This repetition is for the purpose of simplification and clarity, and does not itself indicate the relationship between the various embodiments and / or settings discussed. It should be noted that the components illustrated in the accompanying drawings are not necessarily drawn to scale. The present invention omits the description of well-known components and processing techniques and processes to avoid unnecessarily limiting the present invention.
[0032] Embodiment 1
[0033] Embodiment 1 of the present invention proposes a method for data reconstruction in a distributed file system. The distributed file system performs data reconstruction according to the storage pool priority, where data reconstruction includes data recovery priority, adaptability, read / write priority, and stop data recovery. The storage pool with a higher priority performs data reconstruction first, achieving fine-grained control of data reconstruction. If a storage pool with a lower priority starts data reconstruction first and a storage pool with a higher priority starts reconstruction later, the data reconstruction of the storage pool with a lower priority can be stopped and the data reconstruction of the storage pool with a higher priority can be started immediately. After the data reconstruction of the storage pool with a higher priority is completed, the storage pool with a lower priority can start reconstruction. If the priority of the storage pool is modified during the data reconstruction process, the order of data reconstruction can be determined immediately according to the modified priority. If the data reconstruction priority is set to stop data recovery, the storage pool does not perform data reconstruction, and the storage pool that is performing data reconstruction stops data reconstruction immediately after being set to stop data recovery.
[0034] Such as Figure 1 is a flowchart of a method for data reconstruction in a distributed file system according to Embodiment 1 of the present invention.
[0035] In step S101, set the data reconstruction level for the storage pool; set the data reconstruction level for the storage pool through the command line or interface, including data recovery priority, adaptability, read / write priority, and stop data recovery.
[0036] In step S102, record the storage pool where the data object is located when the data object needs to be reconstructed.
[0037] When the placement group state machine enters the working state, if recovery or backfill is required, record the storage pool where the data object is located.
[0038] Among them, recovery is to perform data reconstruction according to the differences in the placement group logs;
[0039] Backfill is to perform data reconstruction according to the full-scan results of the object.
[0040] In step S103, before the placement group makes a resource reservation, judge the data recovery priority in the placement group. If it is stop data recovery, execute S110. If it is not stop data recovery, execute step S104.
[0041] In step S104, obtain the storage pool with the highest priority recorded in step S102;
[0042] In step S105, compare the priority of the currently running storage pool with the level of the highest-priority storage pool. When the priority of the currently running storage pool is not less than the level of the highest-priority storage pool, execute step S106. When the priority of the currently running storage pool is less than the level of the highest-priority storage pool, execute step S110.
[0043] In step S106, perform resource reservation.
[0044] In step S107, after reserving resources, check the data reconstruction priority in the placement group; after reserving resources, the group starts data reconstruction. The data reconstruction is carried out in a loop, and several objects can be recovered each time.
[0045] In step S108, determine whether the data reconstruction priority is not less than the priority of the storage pool. If it is not less than the priority of the storage pool, execute step S109. If it is less than the priority of the storage pool, execute step S111.
[0046] In step S109, perform data reconstruction.
[0047] In step S110, do not perform resource reservation.
[0048] In step S111, do not perform data reconstruction.
[0049] In step S112, jump to the temporary state, wait for a certain period of time, and then re-enter step S103.
[0050] After all the placement groups in the storage pool in Embodiment 1 of the present invention have completed data reconstruction, delete them from the storage pool where the data objects are recorded.
[0051] Embodiment 1 of the present invention proposes a method for data reconstruction in a distributed file system. According to the storage pool priority, it is judged whether the placement group before the start of reconstruction can perform resource reservation, and after resource reservation, according to the storage pool priority, it is judged whether the placement group during reconstruction can continue data reconstruction. The present invention can enable the distributed file system storage pool to perform data reconstruction according to the priority order, realizing fine-grained control of data reconstruction.
[0052] Embodiment 2
[0053] Based on the method for data reconstruction in a distributed file system proposed in Embodiment 1 of the present invention, Embodiment 2 of the present invention also proposes a system for data reconstruction in a distributed file system. The system includes: a setting module, a resource reservation module, and a data reconstruction module;
[0054] The setting module is used to set the data reconstruction level for the storage pool and record the storage pool where the data object is located when the data object needs to be reconstructed;
[0055] The resource reservation module is used to obtain the storage pool with the highest priority on the premise that the data recovery priority in the placement group is not to stop data recovery. When the priority of the currently running storage pool is not less than the level of the highest priority storage pool, resource reservation is performed.
[0056] The data reconstruction module is used to perform data reconstruction after resource reservation on the premise that the data reconstruction priority in the placement group is not less than the data priority in the storage pool.
[0057] The system also includes a deletion module;
[0058] The deletion module is used to delete from the storage pool where the recorded data object is located after all placement groups in the storage pool have completed data reconstruction.
[0059] Among them, the process implemented by the setting module is: setting the data reconstruction level for the storage pool through the command line or the interface; the data reconstruction level includes recovery priority, adaptive, read / write priority, and stop data recovery.
[0060] When the placement group state machine enters the working state, if recovery or backfill is required, record the storage pool where the data object is located.
[0061] Among them, recovery is to perform data reconstruction based on the differences in the placement group logs;
[0062] Backfill is to perform data reconstruction based on the full scan results of the objects.
[0063] In the resource reservation module, before the placement group performs resource reservation, if the data recovery priority in the placement group is to stop data recovery, or the data recovery priority in the placement group is not to stop data recovery, but the priority of the currently running storage pool is less than the level of the highest priority storage pool, then no resource reservation is performed and it enters the temporary state.
[0064] In the data reconstruction module, after resource reservation is obtained, if it is checked that the data reconstruction priority in the placement group is less than the data priority in the storage pool or the data reconstruction priority in the placement group is to stop data reconstruction, then data reconstruction is stopped and it enters the temporary state.
[0065] The distributed file system performs data reconstruction according to the storage pool priority, where data reconstruction includes data recovery priority, adaptability, read / write priority, and stopping data recovery; the storage pool with a higher priority performs data reconstruction first, realizing fine-grained control of data reconstruction. If the storage pool with a lower priority starts data reconstruction first and the storage pool with a higher priority starts reconstruction later, the data reconstruction of the storage pool with a lower priority can be stopped and the data reconstruction of the storage pool with a higher priority can be started immediately. After the data reconstruction of the storage pool with a higher priority is completed, the storage pool with a lower priority can start reconstruction. If the priority of the storage pool is modified during the data reconstruction process, the order of data reconstruction can be determined immediately according to the modified priority. If the data reconstruction priority is set to stop data recovery, the storage pool does not perform data reconstruction, and the storage pool that is performing data reconstruction stops data reconstruction immediately after being set to stop data recovery
[0066] In Embodiment 2 of the present invention, a system for data reconstruction of a distributed file system is proposed. Before the reconstruction starts, it is judged whether the placement group can reserve resources according to the storage pool priority, and after the resource reservation, it is judged whether the placement group during the reconstruction can continue data reconstruction according to the storage pool priority. The present invention can enable the storage pools of the distributed file system to perform data reconstruction in the order of priority from high to low, realizing fine-grained control of data reconstruction.
[0067] Embodiment 3
[0068] The present invention also proposes a device, including:
[0069] A memory for storing a computer program;
[0070] A processor for implementing the following method steps when executing the computer program:
[0071] Such as Figure 1 is a flowchart of a method for data reconstruction of a distributed file system according to Embodiment 1 of the present invention.
[0072] In step S101, a data reconstruction level is set for the storage pool; the data reconstruction level of the storage pool is set through a command line or an interface, including data recovery priority, adaptability, read / write priority, and stopping data recovery.
[0073] In step S102, the storage pool where the data object is located is recorded when the data object needs to be reconstructed.
[0074] When the placement group state machine enters the working state, if recovery or backfill is required, the storage pool where the data object is located is recorded.
[0075] Among them, to perform recovery is to perform data reconstruction according to the differences in the placement group logs;
[0076] Backfill requires data reconstruction based on the full scan results of the object.
[0077] In step S103, before resource reservation for the placement group, judge the data recovery priority in the placement group. If data recovery is stopped, execute S110. If data recovery is not stopped, execute step S104.
[0078] In step S104, obtain the storage pool with the highest priority recorded in step S102;
[0079] In step S105, compare the priority of the current running storage pool with the level of the highest priority storage pool. When the priority of the current running storage pool is not less than the level of the highest priority storage pool, execute step S106. When the priority of the current running storage pool is less than the level of the highest priority storage pool, execute step S110.
[0080] In step S106, perform resource reservation.
[0081] In step S107, after resource reservation, check the data reconstruction priority in the placement group; after resource reservation, the placement group starts data reconstruction. Data reconstruction is carried out in a loop, and several objects can be recovered each time.
[0082] In step S108, judge whether the data reconstruction priority is not less than the priority of the storage pool. If it is not less than the priority of the storage pool, execute step S109. If it is less than the priority of the storage pool, execute step S111.
[0083] In step S109, perform data reconstruction.
[0084] In step S110, do not perform resource reservation.
[0085] In step S111, do not perform data reconstruction.
[0086] In step S112, jump to the temporary state, wait for a certain time, and then re-enter step S103.
[0087] After all the placement groups in the storage pool in Embodiment 1 of the present invention have completed data reconstruction, they are deleted from the storage pool where the data objects are recorded.
[0088] Embodiment 3 of the present invention proposes a device for data reconstruction of a distributed file system, which judges whether the placement group before reconstruction can perform resource reservation according to the storage pool priority, and after resource reservation, judges whether the placement group during reconstruction can continue data reconstruction according to the storage pool priority. The present invention can enable the distributed file system storage pool to perform data reconstruction in the order of priority from high to low, realizing fine-grained control of data reconstruction.
[0089] It should be noted that the technical solution of the present invention also provides an electronic device, including: a communication interface capable of interacting with other devices such as network devices; a processor connected to the communication interface to achieve information interaction with other devices, and when running a computer program, executing a method for reconstructing distributed file system data provided by one or more of the above technical solutions, and the computer program is stored on a memory. Of course, in actual application, each component in the electronic device is coupled together through a bus system. It can be understood that the bus system is used to realize the connection and communication between these components. The bus system includes not only a data bus, but also a power bus, a control bus, and a status signal bus. The memory in the embodiments of the present application is used to store various types of data to support the operation of the electronic device. Examples of these data include: any computer program for operating on the electronic device. It can be understood that the memory can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM, Read Only Memory), a programmable read-only memory (PROM, Programmable Read-Only Memory), an erasable programmable read-only memory (EPROM, Erasable Programmable Read-Only Memory), an electrically erasable programmable read-only memory (EEPROM, Electrically Erasable Programmable Read-Only Memory), a ferromagnetic random access memory (FRAM, ferromagnetic random access memory), a flash memory (FlashMemory), a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD-ROM, Compact Disc Read-Only Memory); the magnetic surface memory can be a disk memory or a tape memory. The volatile memory can be a random access memory (RAM, Random AccessMemory), which is used as an external cache.By way of example and not limitation, many forms of RAM are available, such as static random access memory (SRAM), synchronous static random access memory (SSRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), sync link dynamic random access memory (SLDRAM), direct rambus random access memory (DRRAM). The memories described in the embodiments of the present application are intended to include but not limited to these and any other suitable types of memories. The methods disclosed in the embodiments of the present application above can be applied to a processor or implemented by a processor. The processor may be an integrated circuit chip with the ability to process signals. In the implementation process, the steps of the above methods can be completed by the integrated logic circuit in the hardware of the processor or instructions in the form of software. The above processor may be a general-purpose processor, a DSP (Digital Signal Processing, that is, a chip capable of implementing digital signal processing technology), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or any conventional processor, etc. Combining the steps of the methods disclosed in the embodiments of the present application, it can be directly embodied as being executed and completed by a hardware decoding processor, or executed and completed by a combination of hardware and software modules in the decoding processor. The software module may be located in a storage medium, and the storage medium is located in the memory. The processor reads the program in the memory and combines its hardware to complete the steps of the foregoing methods. When the processor executes the program, it implements the corresponding processes in the various methods of the embodiments of the present application. For the sake of brevity, it will not be elaborated here.
[0090] Embodiment 4
[0091] The present invention also provides a readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method steps are as follows:
[0092] Such as Figure 1 It is a flowchart of a method for data reconstruction of a distributed file system according to Embodiment 1 of the present invention.
[0093] In step S101, a data reconstruction level is set for the storage pool; the data reconstruction level of the storage pool is set through a command line or an interface, including data recovery priority, adaptability, read / write priority, and stop data recovery.
[0094] In step S102, when a data object needs to be reconstructed, the storage pool where the data object is located is recorded.
[0095] When the placement group state machine enters the working state, if recovery or backfill is required, the storage pool where the data object is located is recorded.
[0096] Among them, recovery is to reconstruct data according to the differences in the placement group logs;
[0097] Backfill is to reconstruct data according to the full scan results of the object.
[0098] In step S103, before resource reservation for the placement group, the data recovery priority in the placement group is judged. If it is to stop data recovery, S110 is executed. If it is not to stop data recovery, step S104 is executed.
[0099] In step S104, the storage pool with the highest priority recorded in step S102 is obtained;
[0100] In step S105, the priority of the currently running storage pool is compared with the level of the highest priority storage pool. When the priority of the currently running storage pool is not less than the level of the highest priority storage pool, step S106 is executed. When the priority of the currently running storage pool is less than the level of the highest priority storage pool, step S110 is executed.
[0101] In step S106, resource reservation is performed.
[0102] In step S107, after the resource is reserved, the data reconstruction priority in the placement group is checked;
[0103] In this step: after the resource is reserved, the group starts data reconstruction. The data reconstruction is performed in a loop, and several objects can be recovered each time.
[0104] In step S108, it is determined whether the data reconstruction priority is not less than the priority of the storage pool. If it is not less than the priority of the storage pool, step S109 is executed; if it is less than the priority of the storage pool, step S111 is executed.
[0105] In step S109, data reconstruction is performed.
[0106] In step S110, resource reservation is not performed.
[0107] In step S111, data reconstruction is not performed.
[0108] In step S112, it jumps to the temporary state and re-enters step S103 after waiting for a certain period of time.
[0109] After all placement groups in the storage pool in Embodiment 1 of the present invention complete data reconstruction, they are deleted from the storage pool where the recorded data objects are located.
[0110] Embodiment 4 of the present invention proposes a storage medium for data reconstruction of a distributed file system. According to the storage pool priority, it is determined whether the placement group before the start of reconstruction can perform resource reservation, and after resource reservation, according to the storage pool priority, it is determined whether the placement group during reconstruction can continue data reconstruction. The present invention can enable the storage pool of the distributed file system to perform data reconstruction in the order of priority from high to low, realizing fine-grained control of data reconstruction.
[0111] The embodiment of the present application also provides a storage medium, namely a computer storage medium, specifically a computer-readable storage medium, such as a memory storing a computer program. The above computer program can be executed by a processor to complete the steps of the foregoing method. The computer-readable storage medium can be a FRAM, ROM, PROM, EPROM, EEPROM, Flash Memory, magnetic surface memory, optical disc, or CD-ROM, etc.
[0112] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps including those of the above method embodiments; and the foregoing storage medium includes various media that can store program codes, such as removable storage devices, ROM, RAM, magnetic disks, or optical discs. Alternatively, if the above integrated units of the present application are implemented in the form of software function modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present application, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing an electronic device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in the embodiments of the present application. And the foregoing storage medium includes various media that can store program codes, such as removable storage devices, ROM, RAM, magnetic disks, or optical discs.
[0113] For the description of the relevant parts in the processing device and storage medium for data reconstruction of the distributed file system provided in the embodiments of the present application, reference can be made to the detailed description of the corresponding parts in the method for data reconstruction of the distributed file system provided in Embodiment 1 of the present application, and details will not be repeated here.
[0114] It should be noted that, in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variation thereof is intended to cover a non-exclusive inclusion, so that a process, method, article or device including a series of elements includes the inherent elements thereof. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the said element. In addition, the parts of the above technical solutions provided in the embodiments of the present application that are consistent with the corresponding technical solutions in the prior art in terms of implementation principles are not described in detail to avoid excessive elaboration.
[0115] Although the specific implementation manners of the present invention have been described above in conjunction with the accompanying drawings, it is not a limitation to the protection scope of the present invention. For those skilled in the art, other different forms of modifications or deformations can be made based on the above description. It is not necessary and impossible to list all the implementation manners here. Based on the technical solution of the present invention, various modifications or deformations that can be made by those skilled in the art without creative efforts are still within the protection scope of the present invention.
Claims
1. A method for reconstructing data in a distributed file system, characterized in that, Including the following steps: Set the data reconstruction priority for the storage pool, and record the storage pool where the data object is located when the data object needs to be reconstructed; On the premise that the data recovery priority in the placement group is not to stop data recovery, obtain the storage pool with the highest priority, where the storage pool with the highest priority is the storage pool with the highest priority where the data object is located when the data object needs to be reconstructed; when the priority of the currently running storage pool is not less than the level of the storage pool with the highest priority, perform resource reservation; After reserving resources, perform data reconstruction on the premise that the data reconstruction priority in the placement group is not less than the priority of the storage pool; The method further includes: before the placement group performs resource reservation, if the data recovery priority in the placement group is to stop data recovery, or the data recovery priority in the placement group is not to stop data recovery, but the priority of the currently running storage pool is less than the level of the storage pool with the highest priority, then no resource reservation is performed and it enters a temporary state; The method further includes: after reserving resources, if it is detected that the data reconstruction priority in the placement group is less than the data priority in the storage pool or the data reconstruction priority in the placement group is to stop data reconstruction, then data reconstruction is stopped and it enters a temporary state.
2. The method for reconstructing data of a distributed file system according to claim 1, characterized in that The method for setting the data reconstruction priority for the storage pool is: Set the data reconstruction priority for the storage pool through the command line or the interface; the data reconstruction priority includes data recovery priority, adaptability, read / write priority, stop data recovery.
3. The method for reconstructing data of a distributed file system according to claim 1, characterized in that The step of recording the storage pool where the data object is located when the data object needs to be reconstructed includes: When the placement group state machine enters the working state, if data reconstruction needs to be performed according to the difference in the placement group log or needs to be performed according to the full scan result of the object, then record the storage pool where the data object is located.
4. A method for reconstructing data of a distributed file system according to any one of claims 1 to 3, characterized in that, After all placement groups in the storage pool have completed data reconstruction, delete from the storage pool that records the location of the data object.
5. A system for reconstructing data in a distributed file system, which is used to execute the method for reconstructing data in a distributed file system according to any one of claims 1 to 4, characterized in that, Including a setting module, a resource reservation module, and a data reconstruction module; The setting module is used to set the data reconstruction priority for the storage pool and record the storage pool where the data object is located when the data object needs to be reconstructed; The resource reservation module is used to obtain the storage pool with the highest priority on the premise that the data recovery priority in the placement group is not to stop data recovery, where the storage pool with the highest priority is the storage pool with the highest priority where the data object is located when the data object needs to be reconstructed; when the priority of the currently running storage pool is not less than the level of the storage pool with the highest priority, perform resource reservation; The data reconstruction module is used to perform data reconstruction after reserving resources on the premise that the data reconstruction priority in the placement group is not less than the priority of the storage pool.
6. The system for reconstructing data of a distributed file system according to claim 5, characterized in that, The system further includes a deletion module; The deletion module is used to delete from the storage pool that records the location of the data object after all placement groups in the storage pool have completed data reconstruction.
7. A device, characterized in that, Including: A memory for storing a computer program; A processor for implementing the method steps described in any one of claims 1 to 4 when executing the computer program.
8. A readable storage medium, characterized in that, A computer program is stored on the readable storage medium, and when the computer program is executed by a processor, the method steps described in any one of claims 1 to 4 are implemented.
Citation Information
Patent Citations
Method and device used for memory equipment management
CN105892934A
Data backup state-based data recovery method and system
CN106339276A