Data recovery method, device and equipment

By creating multiple coroutines under the same thread and collaborating on data recovery tasks, the problem of low data recovery efficiency in distributed storage systems is solved, and fast and efficient data recovery is achieved.

CN120034580APending Publication Date: 2025-05-23XINHUASAN INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510106315.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

After the node failure of the existing distributed storage system, data recovery efficiency is low and it takes a long time to complete data recovery.

Method used

By creating multiple coroutines under the same thread, they are used to obtain the main information list, the backup information list, the missing information list, the pulling of objects to be restored and the recovery data, so as to realize adaptive data recovery of coroutine scheduling.

Benefits of technology

Improves the efficiency of data recovery, shortens recovery time, ensures that data can be quickly restored after a failure occurs, and reduces the impact on user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120034580A_ABST
    Figure CN120034580A_ABST
Patent Text Reader

Abstract

The invention provides a data recovery method, device and equipment, and the method comprises the steps: obtaining a main information list through a first-class coroutine, and sending a comparison request to a backup OSD through the first-class coroutine, the comparison request being used for enabling the backup OSD to obtain a backup information list; a missing information list is obtained through a second type of coroutines, the main information list comprises object information of the main OSD, the backup information list comprises object information of the backup OSD, and the missing information list comprises first object information which exists in the backup information list and does not exist in the main information list; sending a pulling request to the backup OSD through a third type coroutine, wherein the pulling request is used for enabling the backup OSD to obtain a to-be-recovered object corresponding to the first object information; and receiving a to-be-recovered object through the fourth-class coroutine, and recovering missing data of the main OSD based on the to-be-recovered object through the fourth-class coroutine. Through the scheme of the invention, the data reconstruction efficiency is improved, and the switching time among a plurality of threads is saved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of communication technology, and in particular to a data recovery method, device and equipment. Background Art

[0002] In the digital age, the speed at which data is generated is immeasurable. A reliable, efficient, and scalable storage solution is needed to handle massive amounts of data. Distributed storage systems have become an important means to meet this demand. A distributed storage system is a system that stores data in multiple nodes. A distributed storage system has the following characteristics: High availability and fault tolerance: Continue to provide services when a node fails, improving availability and fault tolerance. Scalability: Expand the capacity and performance of the distributed storage system by adding more nodes to adapt to growing data needs. Load balancing: Data is evenly distributed to different nodes to ensure that the load on each node is relatively balanced and to avoid single point failures. Data redundancy and backup: Protect data from hardware failures or data corruption by using data redundancy and backup strategies.

[0003] After a node fails, the data stored in the node needs to be restored to ensure that the data is not lost. However, the efficiency of data recovery is low and it takes a long time to complete the data recovery. Summary of the invention

[0004] The present application provides a data recovery method, which is applied to a master OSD. The method includes:

[0005] Obtaining the main information list through a first type of coroutine created under the thread, and sending a comparison request to the backup OSD through the first type of coroutine, wherein the comparison request is used to enable the backup OSD to obtain the backup information list;

[0006] Obtaining a missing information list through a second type of coroutine created under the thread; wherein the main information list includes object information of the main OSD, the backup information list includes object information of the backup OSD, and the missing information list includes first object information that exists in the backup information list but does not exist in the main information list;

[0007] Sending a pull request to the backup OSD through a third-type coroutine created under the thread, wherein the pull request is used to enable the backup OSD to obtain the object to be restored corresponding to the first object information;

[0008] The object to be restored is received by a fourth type of coroutine created under the thread, and the missing data of the primary OSD is restored based on the object to be restored by the fourth type of coroutine.

[0009] The present application provides a data recovery device, applied to a master OSD, the device comprising:

[0010] A creation module, used to create a first type of coroutine under a thread, create a second type of coroutine under the thread, create a third type of coroutine under the thread, and create a fourth type of coroutine under the thread;

[0011] A processing module, used to obtain a main information list through the first type of coroutine, send a comparison request to a backup OSD through the first type of coroutine, the comparison request is used to enable the backup OSD to obtain a backup information list; obtain a missing information list through the second type of coroutine, wherein the main information list includes object information of the main OSD, the backup information list includes object information of the backup OSD, and the missing information list includes first object information that exists in the backup information list but does not exist in the main information list; send a pull request to the backup OSD through the third type of coroutine, the pull request is used to enable the backup OSD to obtain an object to be restored corresponding to the first object information; receive the object to be restored through the fourth type of coroutine, and restore the missing data of the main OSD based on the object to be restored through the fourth type of coroutine.

[0012] The present application provides an electronic device, comprising: a processor and a machine-readable storage medium, wherein the machine-readable storage medium stores machine-executable instructions that can be executed by the processor; the processor is used to execute the machine-executable instructions to implement the data recovery method of the above example of the present application.

[0013] The present application provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the data recovery method of the above example of the present application is implemented.

[0014] The present application provides a machine-readable storage medium, which stores machine-executable instructions that can be executed by a processor; wherein the processor is used to execute the machine-executable instructions, and when executing the machine-executable instructions, implements the data recovery method of the above example of the present application.

[0015] It can be seen from the above technical solution that in the embodiment of the present application, by creating multiple coroutines (such as the first type of coroutine, the second type of coroutine, the third type of coroutine and the fourth type of coroutine) under the same thread, the reconstruction task is completed through multiple coroutines, and the missing data is recovered, so that adaptive data recovery based on coroutine scheduling can be achieved, and the data recovery process is divided into different coroutine tasks to improve the efficiency of data reconstruction. Since the reconstruction task is completed by multiple coroutines instead of by multiple threads, the switching time between multiple threads is saved, and data recovery is completed in a shorter time, and the efficiency of data recovery is higher. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 It is a flowchart of a data recovery method in one embodiment of the present application;

[0017] Figure 2A is a schematic diagram of switching between different threads in one implementation of the present application;

[0018] Figure 2B is a schematic diagram of different coroutines under the same thread in one implementation manner of the present application;

[0019] Figure 3 is a schematic diagram of the structure of a distributed storage system in one embodiment of the present application;

[0020] Figure 4A It is a flowchart of a data recovery method in one embodiment of the present application;

[0021] Figure 4B It is a flowchart of a data recovery method in one embodiment of the present application;

[0022] Figure 4C It is a flowchart of a data recovery method in one embodiment of the present application;

[0023] Figure 5A is a schematic diagram of multiple priority scheduling pools in one implementation of the present application;

[0024] Figure 5B is a schematic diagram of scheduling coroutines of scheduling pools with different priorities in one implementation of the present application;

[0025] Figure 6 is a structural schematic diagram of a data recovery device in one embodiment of the present application;

[0026] Figure 7 It is a hardware structure diagram of an electronic device in one embodiment of the present application. DETAILED DESCRIPTION

[0027] In the embodiment of the present application, a data recovery method is proposed, which can be applied to the primary OSD (Object Storage Device). The primary OSD is a device that supports the primary OSD. In this embodiment, the device that supports the primary OSD is called the primary OSD. Figure 1 FIG. 4 is a flow chart of the method, which includes:

[0028] Step 101: obtain a main information list through a first-class coroutine created under a thread, and send a comparison request to a backup OSD through the first-class coroutine, where the comparison request is used to enable the backup OSD to obtain a backup information list.

[0029] In an example, the first type of coroutine can be one coroutine, that is, the operations of "obtaining the main information list" and "sending a comparison request to the backup OSD" are executed through the same coroutine. Alternatively, the first type of coroutine can be multiple coroutines, taking two coroutines as an example, that is, the operation of "obtaining the main information list" is executed through one coroutine, and the operation of "sending a comparison request to the backup OSD" is executed through another coroutine.

[0030] For example, considering that "obtaining the master information list" and "sending a comparison request to the backup OSD" are continuous tasks, that is, there are no other tasks between these two operations, the operations of "obtaining the master information list" and "sending a comparison request to the backup OSD" can be executed by the same coroutine.

[0031] Step 102, obtain the missing information list through the second type of coroutine created under the thread; wherein, the main information list includes the object information of the main OSD, the backup information list includes the object information of the backup OSD, and the missing information list includes the first object information that exists in the backup information list but does not exist in the main information list.

[0032] In one example, obtaining the missing information list through the second type of coroutine created under the thread may include but is not limited to: receiving the missing information list sent by the backup OSD through the second type of coroutine; wherein the comparison request includes the main information list, and the missing information list is generated by the backup OSD based on the main information list and the backup information list after the backup OSD obtains the backup information list. In this case, the second type of coroutine can be a coroutine, that is, the operation of "receiving the missing information list sent by the backup OSD" is performed through a coroutine.

[0033] Alternatively, a backup information list sent by the backup OSD is received through the second type of coroutine, and a missing information list is generated based on the main information list and the backup information list through the second type of coroutine.

[0034] In this case, the second type of coroutine can be one coroutine, that is, the operations of "receiving the backup information list sent by the backup OSD" and "generating the missing information list based on the main information list and the backup information list" are executed through the same coroutine. Alternatively, the second type of coroutine can be multiple coroutines. Taking two coroutines as an example, the operation of "receiving the backup information list sent by the backup OSD" is executed through one coroutine, and the operation of "generating the missing information list based on the main information list and the backup information list" is executed through another coroutine. For example, considering that "receiving the backup information list sent by the backup OSD" and "generating the missing information list based on the main information list and the backup information list" are continuous tasks, that is, there are no other tasks between these two operations, the operations of "receiving the backup information list sent by the backup OSD" and "generating the missing information list based on the main information list and the backup information list" can be executed through the same coroutine.

[0035] Step 103: Send a pull request to the backup OSD through the third type of coroutine created under the thread, where the pull request is used to enable the backup OSD to obtain the object to be restored corresponding to the first object information.

[0036] Step 104: Receive the object to be restored through the fourth type of coroutine created under the thread, and restore the missing data of the primary OSD based on the object to be restored through the fourth type of coroutine.

[0037] In one example, the fourth type of coroutine can be one coroutine, that is, the operations of "receiving the object to be recovered" and "recovering the missing data of the primary OSD based on the object to be recovered" can be executed through the same coroutine. Alternatively, the fourth type of coroutine can be multiple coroutines. Taking two coroutines as an example, the operation of "receiving the object to be recovered" can be executed through one coroutine, and the operation of "recovering the missing data of the primary OSD based on the object to be recovered" can be executed through another coroutine. For example, considering that "receiving the object to be recovered" and "recovering the missing data of the primary OSD based on the object to be recovered" are continuous tasks, that is, there are no other tasks between these two operations, therefore, the operations of "receiving the object to be recovered" and "recovering the missing data of the primary OSD based on the object to be recovered" can be executed through the same coroutine.

[0038] In an example, after recovering the missing data of the primary OSD based on the object to be recovered, the recovery information list can be obtained through the fifth type of coroutine created under the thread, the recovery information list includes second object information that exists in the primary information list and does not exist in the backup information list, and the missing object of the backup OSD corresponding to the second object information is obtained through the fifth type of coroutine; the missing object is sent to the backup OSD through the fifth type of coroutine, so that the backup OSD can recover the missing data of the backup OSD based on the missing object.

[0039] In one example, the fifth type of coroutine is one coroutine, that is, the operations of "obtaining the recovery information list", "obtaining the missing object of the backup OSD corresponding to the second object information" and "sending the missing object to the backup OSD" are executed through the same coroutine. Alternatively, the fifth type of coroutine is two coroutines, that is, the operation of "obtaining the recovery information list" is executed through one coroutine, and the operations of "obtaining the missing object of the backup OSD corresponding to the second object information" and "sending the missing object to the backup OSD" are executed through another coroutine. Alternatively, the fifth type of coroutine is two coroutines, that is, the operations of "obtaining the recovery information list" and "obtaining the missing object of the backup OSD corresponding to the second object information" are executed through one coroutine, and the operation of "sending the missing object to the backup OSD" is executed through another coroutine. Alternatively, the fifth type of coroutine is multiple coroutines, taking three coroutines as an example, that is, the operation of "obtaining the recovery information list" is executed through one coroutine, the operation of "obtaining the missing object of the backup OSD corresponding to the second object information" is executed through another coroutine, and the operation of "sending the missing object to the backup OSD" is executed through another coroutine.

[0040] For example, considering that "obtaining the recovery information list", "obtaining the missing object of the backup OSD corresponding to the second object information" and "sending the missing object to the backup OSD" are continuous tasks, the operations of "obtaining the recovery information list", "obtaining the missing object of the backup OSD corresponding to the second object information" and "sending the missing object to the backup OSD" can be executed by the same coroutine.

[0041] In an example, after sending the missing object to the backup OSD through the fifth type of coroutine, the sixth type of coroutine created under the thread can also receive the recovery completion message sent by the backup OSD, and the sixth type of coroutine can determine that the backup OSD has successfully recovered the missing data based on the recovery completion message.

[0042] In an example, the sixth type of coroutine can be one coroutine, that is, the operations of "receiving the recovery completion message sent by the backup OSD" and "determining that the backup OSD has successfully recovered the missing data based on the recovery completion message" are executed through the same coroutine. Alternatively, the sixth type of coroutine can be multiple coroutines, taking two coroutines as an example, that is, the operation of "receiving the recovery completion message sent by the backup OSD" is executed through one coroutine, and the operation of "determining that the backup OSD has successfully recovered the missing data based on the recovery completion message" is executed through another coroutine. For example, considering that "receiving the recovery completion message sent by the backup OSD" and "determining that the backup OSD has successfully recovered the missing data based on the recovery completion message" are continuous tasks, the operations of "receiving the recovery completion message sent by the backup OSD" and "determining that the backup OSD has successfully recovered the missing data based on the recovery completion message" can be executed through the same coroutine.

[0043] In an example, the first type of coroutine, the second type of coroutine, the third type of coroutine, the fourth type of coroutine, the fifth type of coroutine and the sixth type of coroutine can be created under the same thread. For example, when the data recovery process is triggered, the first type of coroutine, the second type of coroutine, the third type of coroutine, the fourth type of coroutine, the fifth type of coroutine and the sixth type of coroutine can be created. For example, when the data recovery process is triggered, the first type of coroutine can be created. After the processing of the first type of coroutine is completed, the second type of coroutine can be created. After the processing of the second type of coroutine is completed, the third type of coroutine can be created. After the processing of the third type of coroutine is completed, the fourth type of coroutine can be created. After the processing of the fourth type of coroutine is completed, the fifth type of coroutine can be created. After the processing of the fifth type of coroutine is completed, the sixth type of coroutine can be created. The above is just an example, and there is no restriction on the creation timing of these coroutines, as long as these coroutines are created under the same thread.

[0044] In one example, the main OSD includes at least two priority scheduling pools. After creating a reconstruction task coroutine under a thread, the reconstruction task coroutine is scheduled to the priority scheduling pool corresponding to the reconstruction task. The reconstruction task is used to recover the missing data of the main OSD and / or the backup OSD; wherein the reconstruction task coroutine includes a first type of coroutine, a second type of coroutine, a third type of coroutine, a fourth type of coroutine, a fifth type of coroutine, and a sixth type of coroutine. Based on the weight value corresponding to each priority scheduling pool, coroutines are traversed in sequence from at least two priority scheduling pools as the current coroutine, and the task corresponding to the current coroutine is executed through the current coroutine.

[0045] Among them, when the current coroutine is a reconstruction task coroutine, the reconstruction task can be executed through the reconstruction task coroutine to restore the missing data of the primary OSD and / or the backup OSD.

[0046] In one example, at least two priority scheduling pools include a first-level priority scheduling pool, a second-level priority scheduling pool, and a third-level priority scheduling pool. The first weight value corresponding to the first-level priority scheduling pool is greater than the second weight value corresponding to the second-level priority scheduling pool, and the second weight value is greater than the third weight value corresponding to the third-level priority scheduling pool; the read-write task corresponds to the first-level priority scheduling pool, the reconstruction task corresponds to the second-level priority scheduling pool, and the remaining tasks other than the read-write task and the reconstruction task correspond to the third-level priority scheduling pool.

[0047] Based on the weight value corresponding to each priority scheduling pool, traversing the coroutines from at least two priority scheduling pools in sequence as the current coroutine may include but is not limited to: traversing the coroutines with the first weight value from the first-level priority scheduling pool as the current coroutine, traversing the coroutines with the second weight value from the second-level priority scheduling pool as the current coroutine, traversing the coroutines with the third weight value from the third-level priority scheduling pool as the current coroutine, and returning to execute traversing the coroutines with the first weight value from the first-level priority scheduling pool as the current coroutine;

[0048] Among them, when traversing the coroutines from the first-level priority scheduling pool, if the number of coroutines in the first-level priority scheduling pool is less than the first weight value, the remaining coroutines are supplemented by the second-level priority scheduling pool and / or the third-level priority scheduling pool; when traversing the coroutines from the second-level priority scheduling pool, if the number of coroutines in the second-level priority scheduling pool is less than the second weight value, the remaining coroutines are supplemented by the third-level priority scheduling pool.

[0049] It can be seen from the above technical solution that in the embodiment of the present application, by creating multiple coroutines (such as the first type of coroutine, the second type of coroutine, the third type of coroutine and the fourth type of coroutine) under the same thread, the reconstruction task is completed through multiple coroutines, and the missing data is recovered, so that adaptive data recovery based on coroutine scheduling can be achieved, and the data recovery process is divided into different coroutine tasks to improve the efficiency of data reconstruction. Since the reconstruction task is completed by multiple coroutines instead of by multiple threads, the switching time between multiple threads is saved, and data recovery is completed in a shorter time, and the efficiency of data recovery is higher.

[0050] The data recovery method of the embodiment of the present application is described below in conjunction with specific application scenarios.

[0051] In distributed storage systems, fault tolerance and self-recovery are usually the primary goals. On the one hand, distributed storage systems require continuous and stable performance and do not want to have a significant impact on performance after a failure. On the other hand, some business scenarios require data to be restored as soon as possible to ensure that data is not lost. Therefore, data self-recovery after a failure is crucial to distributed storage systems. However, the efficiency of data recovery is low, and it takes a long time to complete data recovery. In addition, during the data recovery process, read and write operations cannot be performed, resulting in users being unable to write or read data, affecting the user experience.

[0052] In view of the above findings, an adaptive data recovery method based on coroutine scheduling is proposed in the embodiment of the present application. In a distributed storage system, the data recovery process is divided into different coroutine tasks through the scheduling optimization of coroutines, which can improve the efficiency of data reconstruction. The efficiency of data recovery is high, and data recovery can be completed in a short time. In addition, the reconstruction tasks and read-write tasks are divided into scheduling pools of different priorities, and weighted scheduling is used to ensure that the impact of the reconstruction tasks on the read-write tasks does not exceed a certain proportion, reduce the impact of the reconstruction tasks on the read-write tasks, give priority to the read-write tasks, and perform read-write operations during the data recovery process. Users can write and read data, thereby improving the user experience.

[0053] For example, see Figure 2A The figure shows the schematic diagram of different thread switching. Thread 1 and thread 2 can be created. Thread 1 is used to execute task 1, and thread 2 is used to execute task 2. First, task 1 is executed through thread 1. When thread 1 is waiting, the operating system will block thread 1 and switch to thread 2 to continue processing. Then, task 2 is executed through thread 2. When thread 2 is waiting, the operating system will block thread 2 and switch to thread 1 to continue processing. Then, task 1 is executed through thread 1, and so on.

[0054] When the operating system blocks thread 1 and switches thread 1 to thread 2, the processing of thread 1 is in user state, the processing of thread 2 is in user state, and the processing of switching thread 1 to thread 2 is in kernel state. Therefore, it is necessary to switch from user state to kernel state to perform thread switching, and then switch from kernel state to user state. Since the switching time between user state and kernel state is very long, see Figure 2A As shown in the figure, the switching overhead between threads will be very large. Obviously, when the number of threads is large, the threads will occupy a lot of memory space, and too many thread switches will take up a lot of time.

[0055] For example, see Figure 2B The figure shows a schematic diagram of different coroutines under the same thread. Coroutine 1, coroutine 2 and coroutine 3 can be created. Coroutine 1 is used to execute task 1, coroutine 2 is used to execute task 2, and coroutine 3 is used to execute task 3. First, execute task 1 through coroutine 1. After task 1 is completed, switch to coroutine 2 for further processing. Then, execute task 2 through coroutine 2. After task 2 is completed, switch to coroutine 3 for further processing. Then, execute task 3 through coroutine 3, and so on.

[0056] Coroutines run on threads and are lightweight threads. Coroutine scheduling is performed in user space. The coroutines in this embodiment can also be replaced by ULT (User Level Thread). After a coroutine is executed, you can choose to give up and let another coroutine run on the current thread. Coroutines do not increase the number of threads. They just run multiple coroutines in a time-sharing multiplexing manner based on threads. The switching of coroutines is completed in user state, and the cost of switching is much smaller than the cost of switching threads from user state to kernel state.

[0057] In summary, when switching from coroutine 1 to coroutine 2, since the processing of coroutine 1 and coroutine 2 are both in user state, and the processing of coroutine 1 switching to coroutine 2 is also in kernel state, there is no switching between user state and kernel state. Figure 2B As shown, the switching overhead between coroutines is small.

[0058] A distributed storage system can include multiple nodes (storage nodes, also called storage servers) through which data is stored. Object-based storage is a network storage architecture, and the device based on object storage is an object storage device, referred to as OSD. When a distributed storage system uses object storage for data storage, the storage node is an OSD, that is, data is stored through multiple OSDs.

[0059] Object storage is also called object-oriented storage or cloud storage. It has the advantages of high-speed direct access and distributed data sharing. It provides a storage architecture with high performance, high reliability, cross-platform and secure data sharing. Object storage allows the computing infrastructure to be separated from storage requirements.

[0060] When a distributed storage system uses object storage, the main operation is the object (Object). The object is the basic unit of data storage in the distributed storage system. An object is actually a combination of file data and a set of attribute information (MetaData). Each object has a globally unique identifier, and the OSD accesses the object through the globally unique identifier. For example, an object consists of three parts: Key, Value, and Metadata. The Key is the globally unique identifier of the object, which can be understood as the file name and is used to retrieve the object, similar to a URL address. The Value is the data itself. Metadata is metadata, similar to the label (attribute) of an object (data), which can be various descriptive information of the object. For example, for a picture of a person (that is, the Value of the object is the picture of the person), Metadata can be name, gender, age, shooting location, shooting time, etc.

[0061] From the above, it can be seen that a distributed storage system can include multiple OSDs, and a distributed storage system can store the same data (object) in multiple OSDs. Figure 3 As shown, it is a schematic diagram of the structure of a distributed storage system. Taking the distributed storage system including 3 OSDs as an example, one OSD can be used as the main OSD, and the remaining OSDs can be used as backup OSDs. For example, OSD1 is used as the main OSD, and OSD2 and OSD3 are used as backup OSDs. When the same object is stored in multiple OSDs, OSD1 can store object a1, object a2, object a3, and object a4, OSD2 can store object a1, object a2, object a3, and object a4, and OSD3 can store object a1, object a2, object a3, and object a4.

[0062] In a distributed storage system, objects stored in OSD may be lost when an upgrade, maintenance, restart, power failure, or disk failure occurs, which will trigger data reconstruction, also known as data self-recovery. Data reconstruction is used to recover data (objects) missing from OSD. For example, when object a1 of OSD1 is lost, OSD1 can obtain object a1 from OSD2 or OSD3, and recover the lost object of OSD1 based on object a1. Obviously, based on the multi-copy mechanism or the erasure mechanism, the failed copy data can be reconstructed to ensure data integrity.

[0063] In the embodiment of the present application, a data recovery method is proposed. By creating multiple coroutines under the same thread, the reconstruction task can be completed through multiple coroutines to recover the missing data. Figure 4A FIG. 1 is a flow chart of the data recovery method, which may include:

[0064] Step 401: The main OSD creates a first-class coroutine under a thread, and obtains a main information list through the first-class coroutine. The main information list may include object information of the main OSD, such as a globally unique identifier (Key) of the object.

[0065] In an example, in order to achieve data reconstruction, the master OSD can create a thread for the data reconstruction task, and the thread is subsequently recorded as thread P1. The master OSD can create a first-class coroutine under thread P1, and after creating the first-class coroutine, the first-class coroutine can be run to obtain the main information list through the first-class coroutine.

[0066] For example, the master OSD scans the global unique identifiers of all local objects of the master OSD through the first type of coroutine, and adds the global unique identifiers of these objects to the main information list, which may include the global unique identifiers of all objects of the master OSD. For example, when the master OSD stores objects a1, a2, a3, and a4, if object a4 is lost, the master information list may include the global unique identifier of object a1, the global unique identifier of object a2, and the global unique identifier of object a3.

[0067] Step 402: The primary OSD sends a comparison request to the backup OSD through the first type of causal routine.

[0068] In an example, when the primary OSD runs the first type of coroutine, it can also send a comparison request to each backup OSD through the first type of coroutine, and the comparison request is used to enable the backup OSD to obtain the backup information list. For example, when OSD1 is the primary OSD and OSD2 and OSD3 are backup OSDs, OSD1 sends a comparison request to OSD2 through the first type of coroutine, and OSD1 sends a comparison request to OSD3 through the first type of coroutine.

[0069] Since the first type of coroutine is used to scan the globally unique identifiers of all objects and the first type of coroutine is used to send a comparison request to the backup OSD, the first type of coroutine is also called a coroutine for scanning and sending statistical information.

[0070] When running the first type of coroutine, there is no order relationship between the operation of sending a comparison request to the backup OSD through the first type of coroutine and the operation of obtaining the main information list through the first type of coroutine. For example, send a comparison request to the backup OSD through the first type of coroutine first, or obtain the main information list through the first type of coroutine first.

[0071] Step 403: After receiving the comparison request, the backup OSD obtains a backup information list, which may include object information of the backup OSD, such as a globally unique identifier (Key) of the object.

[0072] For example, since the comparison request is used to enable the backup OSD to obtain the backup information list, after receiving the comparison request, the backup OSD can scan the global unique identifiers of all objects local to the backup OSD, and add the global unique identifiers of these objects to the backup information list, and the backup information list can include the global unique identifiers of all objects of the backup OSD. For example, when the backup OSD stores objects a1, a2, a3, and a4, if object a3 is lost, the backup information list can include the global unique identifier of object a1, the global unique identifier of object a2, and the global unique identifier of object a4.

[0073] In an example, in order to achieve data reconstruction, the backup OSD can create a thread for the data reconstruction task, which is subsequently recorded as thread P2. The backup OSD can create coroutine X1 under thread P2, and after creating coroutine X1, it can run coroutine X1. For example, the backup OSD receives a comparison request through coroutine X1, and scans the global unique identifiers of all objects local to the backup OSD to obtain a backup information list.

[0074] Step 404: The backup OSD sends a backup information list to the primary OSD. For example, the backup OSD may send a comparison response to the primary OSD, and the comparison response may include a backup information list. For example, when the backup OSD is running coroutine X1, the backup OSD may send a backup information list to the primary OSD through coroutine X1.

[0075] Step 405: The master OSD creates a second type of coroutine under the thread, and the second type of coroutine runs in the same thread as the first type of coroutine. The master OSD receives the backup information list through the second type of coroutine, and generates a missing information list based on the main information list and the backup information list through the second type of coroutine. The missing information list includes object information (such as the global unique identifier of the object) that exists in the backup information list but does not exist in the main information list.

[0076] The master OSD can also create a second type of coroutine under thread P1, and the second type of coroutine and the first type of coroutine run under the same thread P1. After the processing of the first type of coroutine is completed, the first type of coroutine can be exited and the second type of coroutine can be run, that is, the first type of coroutine actively gives up and allows the second type of coroutine to run on thread P1.

[0077] When the second type of coroutine is running, the primary OSD receives the backup information list sent by the backup OSD through the second type of coroutine, such as receiving a comparison response sent by the backup OSD, and obtaining the backup information list from the comparison response.

[0078] Then, the master OSD can compare the master information list and the backup information list through the second type of coroutine, find the object information that exists in the backup information list but does not exist in the master information list, call the object information the first object information, and add the first object information (globally unique identifier of the object) to the missing information list, that is, the missing information list may include the first object information that exists in the backup information list but does not exist in the master information list. Obviously, the missing information list can represent the objects that are missing from the master OSD.

[0079] For example, the main information list includes the globally unique identifier of object a1, the globally unique identifier of object a2, and the globally unique identifier of object a3, and the backup information list includes the globally unique identifier of object a1, the globally unique identifier of object a2, and the globally unique identifier of object a4. Then, the first object information that exists in the backup information list and does not exist in the main information list is the globally unique identifier of object a4. Based on this, the missing information list can include the globally unique identifier of object a4, indicating that the object missing from the primary OSD is object a4.

[0080] When the primary OSD compares the primary information list with the backup information list through the second type of coroutine, if there are multiple backup OSDs, the multiple backup OSDs will return multiple backup information lists. Therefore, the primary OSD compares the primary information list with the multiple backup information lists through the second type of coroutine to obtain a missing information list.

[0081] For example, since the second type of coroutine is used to receive the backup information list and the second type of coroutine is used to generate the missing information list, the second type of coroutine can also be called a coroutine for obtaining the information list.

[0082] In summary, based on steps 401 to 405, the objects missing from the primary OSD can be found.

[0083] In another possible implementation, after the primary OSD obtains the primary information list through the first type of coroutine, when sending a comparison request to the backup OSD through the first type of coroutine, the comparison request may include the primary information list.

[0084] After the backup OSD obtains the backup information list through the coroutine X1, the coroutine X1 can generate a missing information list based on the main information list and the backup information list, and the missing information list can include the first object information that exists in the backup information list but does not exist in the main information list. Then, the backup OSD can send the missing information list to the main OSD through the coroutine X1. On this basis, the main OSD receives the missing information list through the second type of coroutine, and then learns the objects missing from the main OSD based on the missing information list.

[0085] Based on the missing information list, a data recovery method is proposed in the embodiment of the present application. Figure 4B FIG. 1 is a flow chart of the data recovery method, which may include:

[0086] Step 406: The primary OSD creates a third-type coroutine under the thread, and the third-type coroutine and the second-type coroutine run under the same thread. The primary OSD sends a pull request to the backup OSD through the third-type coroutine.

[0087] The master OSD can also create a third type of coroutine under thread P1, and the third type of coroutine and the second type of coroutine run under the same thread P1. After the processing of the second type of coroutine is completed, the second type of coroutine can be exited and the third type of coroutine can be run, that is, the second type of coroutine actively gives up and allows the third type of coroutine to run on thread P1.

[0088] When running the third type of coroutine, the primary OSD can send a pull request to the backup OSD through the third type of coroutine. The pull request may include the first object information (i.e., the first object information in the missing information list), and the pull request is used to enable the backup OSD to obtain the object to be restored corresponding to the first object information.

[0089] For example, assuming that the primary OSD needs to obtain the object to be restored corresponding to the first object information from OSD2 (backup OSD), the primary OSD can send a pull request to OSD2 through the third type of coroutine.

[0090] Since the third type of coroutine is used to send a pull request to the backup OSD, and the pull request is used to pull the object missing from the primary OSD, the third type of coroutine can also be called a coroutine for pulling the object missing from itself.

[0091] Step 407: After receiving the pull request, the backup OSD obtains the object to be restored corresponding to the first object information (the object missing from the primary OSD is called the object to be restored, that is, the object corresponding to the first object information).

[0092] For example, since the pull request is used to enable the backup OSD to obtain the object to be restored corresponding to the first object information, after receiving the pull request, the backup OSD can also obtain the first object information from the pull request, and find the object corresponding to the first object information from all local objects of the backup OSD. The object corresponding to the first object information is the object to be restored that is missing from the primary OSD.

[0093] In an example, the backup OSD may create a coroutine X2 under thread P2 and run the coroutine X2. The backup OSD receives a pull request through the coroutine X2, obtains the first object information from the pull request, and obtains the object to be restored corresponding to the first object information from all local objects of the backup OSD.

[0094] Step 408: The backup OSD sends the object to be restored to the primary OSD (ie, sends data). For example, when the backup OSD is running the coroutine X2, the backup OSD can send the object to be restored to the primary OSD through the coroutine X2.

[0095] Step 409: The master OSD creates a fourth type of coroutine under the thread, and the fourth type of coroutine and the third type of coroutine run under the same thread. The master OSD receives the object to be restored through the fourth type of coroutine, and restores the missing data of the master OSD based on the object to be restored through the fourth type of coroutine, that is, stores the object to be restored.

[0096] The master OSD can also create a fourth type of coroutine under thread P1, and the fourth type of coroutine and the third type of coroutine run under the same thread P1. After the processing of the third type of coroutine is completed, the third type of coroutine can be exited and the fourth type of coroutine can be run, that is, the third type of coroutine actively gives up and allows the fourth type of coroutine to run on thread P1.

[0097] When the fourth type of coroutine is running, the primary OSD can receive the object to be restored through the fourth type of coroutine, and the object to be restored is the object missing from the primary OSD. Then, the primary OSD restores the data missing from the primary OSD based on the object to be restored through the fourth type of coroutine, that is, the object to be restored is stored through the fourth type of coroutine. After the object to be restored is stored in the primary OSD, the data missing from the primary OSD is restored.

[0098] Since the fourth type of coroutine is used to receive the object to be restored and restore the missing data of the primary OSD based on the object to be restored, the fourth type of coroutine can also be called a coroutine that restores its own missing objects.

[0099] In one example, if there are multiple first object information in the missing information list, then in steps 406-409, the master OSD can simultaneously obtain the objects to be restored corresponding to all the first object information, and restore the missing data of the master OSD based on these objects to be restored, thereby completing the data recovery of the master OSD. Alternatively, in steps 406-409, the master OSD can obtain an object to be restored corresponding to the first object information, and restore the missing data of the master OSD based on the object to be restored, and repeat steps 406-409 to obtain the object to be restored corresponding to the next first object information, and restore the missing data of the master OSD based on the object to be restored, and so on, until all the objects to be restored are obtained and the data recovery of the master OSD is completed.

[0100] In summary, based on steps 406 to 409, data recovery of the primary OSD can be completed.

[0101] On the basis of the data recovery of the main OSD, if the backup OSD also has missing data, that is, the backup OSD needs to be restored, then a data recovery method is proposed in the embodiment of the present application, see Figure 4C FIG. 1 is a flow chart of the data recovery method, which may include:

[0102] Step 410: The primary OSD creates a fifth-type coroutine under a thread, and the fifth-type coroutine runs in the same thread as the fourth-type coroutine. The primary OSD obtains the recovery information list of the backup OSD through the fifth-type coroutine, and the recovery information list includes the second object information that exists in the primary information list but does not exist in the backup information list.

[0103] The master OSD can also create a fifth type of coroutine under thread P1, and the fifth type of coroutine and the fourth type of coroutine run under the same thread P1. After the fourth type of coroutine is processed, the fourth type of coroutine can be exited and the fifth type of coroutine can be run, that is, the fourth type of coroutine actively gives up and allows the fifth type of coroutine to run on thread P1.

[0104] When running the fifth type of coroutine, the primary OSD can obtain the recovery information list of the backup OSD through the fifth type of coroutine. For the sake of distinction, the missing information list of the backup OSD is called the recovery information list.

[0105] For example, the primary OSD can compare the main information list and the backup information list through the fifth type of coroutine to find the object information that exists in the main information list but does not exist in the backup information list, call the object information the second object information, and add the second object information (the object's globally unique identifier) ​​to the recovery information list of the backup OSD. Obviously, the recovery information list can represent the objects that are missing from the backup OSD.

[0106] For another example, after the primary OSD obtains the primary information list through the first type of coroutine, when sending a comparison request to the backup OSD through the first type of coroutine, the comparison request includes the primary information list. After the backup OSD obtains the backup information list through coroutine X1, it generates a recovery information list based on the primary information list and the backup information list through coroutine X1, and the recovery information list includes second object information that exists in the primary information list but does not exist in the backup information list. The backup OSD sends the recovery information list to the primary OSD through coroutine X1, and the primary OSD stores the recovery information list. On this basis, in step 410, the primary OSD obtains the stored recovery information list through the fifth type of coroutine, and then learns the objects that are missing from the backup OSD based on the recovery information list.

[0107] Step 411: The primary OSD obtains the missing object of the backup OSD corresponding to the second object information through the fifth type of coroutine, and the primary OSD sends the missing object to the backup OSD through the fifth type of coroutine.

[0108] For each second object information in the recovery information list, the primary OSD finds the object corresponding to the second object information from all local objects of the primary OSD through the fifth type of coroutine. The object is the object that the backup OSD is missing. For the sake of convenience, the object that the backup OSD is missing is called a missing object.

[0109] After obtaining the missing object of the backup OSD, the primary OSD sends the missing object to the backup OSD through the fifth type of goroutine, thereby sending the missing object of the backup OSD to the backup OSD.

[0110] For example, since the fifth type of coroutine is used to obtain the missing object of the backup OSD, and the fifth type of coroutine is used to send the missing object to the backup OSD, the fifth type of coroutine is also called a coroutine for sending an object.

[0111] Step 412: The backup OSD receives the missing object (ie, the object that the backup OSD has lost), and the backup OSD recovers the data that the backup OSD has lost based on the missing object, that is, stores the missing object.

[0112] In an example, the backup OSD can create a coroutine X3 under thread P2 and run the coroutine X3. On this basis, the backup OSD receives the missing object through the coroutine X3, and the missing object is the object that the backup OSD has lost. Then, the backup OSD recovers the data that the backup OSD has lost based on the missing object through the coroutine X3, that is, stores the missing object through the coroutine X3. After the missing object is stored in the backup OSD, the data that the backup OSD has lost is recovered.

[0113] Step 413: After restoring the missing data of the backup OSD, the backup OSD sends a restoration completion message to the primary OSD, where the restoration completion message indicates that the missing data has been successfully restored.

[0114] For example, after the backup OSD completes the recovery of the missing data of the backup OSD through the coroutine X3, the backup OSD can also send a recovery completion message to the primary OSD through the coroutine X3.

[0115] Step 414: The primary OSD creates a sixth-type coroutine under the thread, and the sixth-type coroutine runs in the same thread as the fifth-type coroutine. The primary OSD receives the recovery completion message sent by the backup OSD through the sixth-type coroutine, and determines based on the recovery completion message that the backup OSD has successfully recovered the missing data.

[0116] The master OSD can also create a sixth-type coroutine under thread P1, and the sixth-type coroutine and the fifth-type coroutine run under the same thread P1. After the fifth-type coroutine is processed, the fifth-type coroutine can be exited and the sixth-type coroutine can be run, that is, the fifth-type coroutine actively gives up and allows the sixth-type coroutine to run on thread P1.

[0117] After receiving the recovery completion message through the sixth type of coroutine, the primary OSD learns that the backup OSD has successfully recovered the missing data, that is, the data recovery of the backup OSD has been successfully completed.

[0118] When the data of multiple backup OSDs need to be restored, the primary OSD receives a recovery completion message after learning that the data recovery of the backup OSD is complete, and can restore the data of the next backup OSD, repeating steps 410 to 414, and so on, until the data recovery of all backup OSDs is complete.

[0119] Since the sixth type of coroutine is used to receive a recovery completion message and determine based on the recovery completion message that the backup OSD has successfully recovered the data, the sixth type of coroutine is also called a coroutine that completes the object recovery process.

[0120] In summary, based on steps 410 to 414, data recovery of the backup OSD can be completed.

[0121] After completing the data recovery of the primary OSD and the backup OSD, wait for the next triggering of the data reconstruction task and re-execute steps 401-414. That is, the data reconstruction task can be executed periodically, and each time the data reconstruction task is executed, steps 401-414 are executed.

[0122] In summary, it can be seen that by executing each stage of data recovery in the form of coroutine tasks, there is no need to switch from user state to kernel state, thereby optimizing the time of task switching. Especially in the application scenario of massive small files (objects), the time for recovering a single object is very short. By dividing the coroutine tasks, each task does not need to switch between user state and kernel state, which greatly improves the efficiency of data recovery.

[0123] In an example, in order to give priority to read and write tasks so that the data recovery process can also perform read and write operations, the main OSD (backup OSD) can include at least two priority scheduling pools, and schedule the coroutines of the read and write tasks to the priority scheduling pool corresponding to the read and write tasks, and schedule the coroutines of the reconstruction tasks to the priority scheduling pool corresponding to the reconstruction tasks, thereby dividing the reconstruction tasks and read and write tasks into different priority scheduling pools.

[0124] For example, the master OSD (backup OSD) includes a first-level priority scheduling pool and a second-level priority scheduling pool, the read-write task corresponds to the first-level priority scheduling pool, the reconstruction task corresponds to the second-level priority scheduling pool, and the remaining tasks other than the read-write task and the reconstruction task can correspond to the first-level priority scheduling pool or the second-level priority scheduling pool. For another example, the master OSD (backup OSD) includes a first-level priority scheduling pool, a second-level priority scheduling pool, and a third-level priority scheduling pool, the read-write task corresponds to the first-level priority scheduling pool, the reconstruction task corresponds to the second-level priority scheduling pool, and the remaining tasks other than the read-write task and the reconstruction task correspond to the third-level priority scheduling pool. For another example, the master OSD (backup OSD) includes a first-level priority scheduling pool, a second-level priority scheduling pool, a third-level priority scheduling pool, and a fourth-level priority scheduling pool, the read-write task corresponds to the first-level priority scheduling pool, the reconstruction task corresponds to the second-level priority scheduling pool, and the remaining tasks other than the read-write task and the reconstruction task correspond to the third-level priority scheduling pool and the fourth-level priority scheduling pool. In this embodiment, there is no restriction on the number of priority scheduling pools, and a primary OSD (backup OSD) including a first-level priority scheduling pool, a second-level priority scheduling pool, and a third-level priority scheduling pool is taken as an example.

[0125] See also Figure 5A As shown in the figure, it is a schematic diagram of multiple priority scheduling pools. The first-level priority scheduling pool can also be called a high-priority scheduling pool, and the coroutines of read and write tasks can be scheduled to the high-priority scheduling pool. For example, each time a coroutine for a read and write task is created, the coroutine scheduler schedules the coroutine to the high-priority scheduling pool, such as coroutine 0, coroutine 1, coroutine 2, coroutine 3, and coroutine 4 are all coroutines for read and write tasks.

[0126] The second-level priority scheduling pool can also be called the medium priority scheduling pool, and the coroutine of the reconstruction task (which can be called the reconstruction task coroutine) can be scheduled to the medium priority scheduling pool. For example, each time a coroutine for a reconstruction task (which can also be called a data recovery task, and the reconstruction task is used to recover the missing data of the primary OSD and / or backup OSD) is created, the coroutine scheduler can schedule the coroutine to the medium priority scheduling pool, such as coroutine 5, coroutine 6, coroutine 7, coroutine 8 and coroutine 9 are all coroutines for the reconstruction task.

[0127] For the main OSD, the reconstruction task coroutine may include the first type of coroutine, the second type of coroutine, the third type of coroutine, the fourth type of coroutine, the fifth type of coroutine and the sixth type of coroutine. Based on this, after the main OSD creates the reconstruction task coroutine under the thread, it needs to schedule the reconstruction task coroutine to the priority scheduling pool (i.e., the medium priority scheduling pool) corresponding to the reconstruction task. For example, when the main OSD creates the first type of coroutine under the thread, it does not directly run the first type of coroutine, but first schedules the first type of coroutine to the medium priority scheduling pool, until the coroutine scheduler schedules to the first type of coroutine, and then runs the first type of coroutine to perform the above processing. In addition, when the main OSD creates the second type of coroutine under the thread, it does not directly run the second type of coroutine, but first schedules the second type of coroutine to the medium priority scheduling pool, until the coroutine scheduler schedules to the second type of coroutine, and then runs the second type of coroutine to perform the above processing, and so on.

[0128] For the backup OSD, the reconstruction task coroutine may include coroutine X1, coroutine X2 and coroutine X3. Based on this, after the backup OSD creates the reconstruction task coroutine under the thread, it needs to schedule the reconstruction task coroutine to the priority scheduling pool (i.e., the medium priority scheduling pool) corresponding to the reconstruction task. For example, when the backup OSD creates coroutine X1 under the thread, it first schedules coroutine X1 to the medium priority scheduling pool, and does not run coroutine X1 to perform the above processing until the coroutine scheduler schedules it to the coroutine X1, and so on.

[0129] The third-level priority scheduling pool can also be called the low-priority scheduling pool, which can schedule the coroutines of tasks other than the read-write task and the reconstruction task to the low-priority scheduling pool. For example, each time a coroutine for other tasks is created, the coroutine scheduler schedules the coroutine to the low-priority scheduling pool, such as coroutine 10, coroutine 11, coroutine 12, coroutine 13, and coroutine 14 are all coroutines for other tasks.

[0130] In one example, on the basis of dividing multiple priority scheduling pools, a coroutine scheduling strategy based on weight priority can be used to divide different priority scheduling pools into different weight values. For example, the weight value corresponding to the first-level priority scheduling pool is called the first weight value, the weight value corresponding to the second-level priority scheduling pool is called the second weight value, and the weight value corresponding to the third-level priority scheduling pool is called the third weight value.

[0131] The first weight value may be greater than the second weight value (the first weight value may also be equal to the second weight value), and the second weight value may be greater than the third weight value (the second weight value may also be equal to the third weight value). For example, the first weight value may be 3, the second weight value may be 2, and the third weight value may be 1. For another example, the first weight value may be 5, the second weight value may be 3, and the third weight value may be 1, and so on.

[0132] From the above, it can be seen that the priority scheduling pool can be divided into a high priority scheduling pool, a medium priority scheduling pool and a low priority scheduling pool, and the priority scheduling pools corresponding to the tasks can be divided based on the importance of different tasks. For example, the coroutines of read and write tasks are divided into the high priority scheduling pool, the coroutines of reconstruction tasks are divided into the medium priority scheduling pool, and the coroutines of other tasks (i.e. non-urgent tasks, such as scanning, detection, verification, etc.) are divided into the low priority scheduling pool. The first weight value of the high priority scheduling pool is set to 3, the second weight value of the medium priority scheduling pool is set to 2, and the third weight value of the low priority scheduling pool is set to 1.

[0133] In one example, based on the weight value corresponding to each priority scheduling pool, the coroutines can be traversed in sequence from at least two priority scheduling pools as the current coroutine, and the task corresponding to the current coroutine can be executed through the current coroutine. Among them, when the current coroutine is a coroutine of a read-write task, the read-write operation can be performed based on the current coroutine. When the current coroutine is a reconstruction task coroutine, the reconstruction task can be executed through the current coroutine (i.e., the reconstruction task coroutine) to restore the missing data of the primary OSD and / or backup OSD, see steps 401-414. When the current coroutine is a coroutine of other tasks, related operations can be performed through the current coroutine.

[0134] For example, when the coroutine scheduler is scheduling coroutines, it schedules coroutines in different priority scheduling pools in rounds. Each round of scheduling is scheduled according to the preset weight value. In each round, coroutines in different priority scheduling pools are scheduled based on the weight values ​​of different priority scheduling pools. If the number of coroutines in a priority scheduling pool is insufficient, coroutines are supplemented from the next priority scheduling pool.

[0135] In an example, taking the first-level priority scheduling pool, the second-level priority scheduling pool and the third-level priority scheduling pool as examples, first traverse the first weight value coroutines from the first-level priority scheduling pool as the current coroutine, then traverse the second weight value coroutines from the second-level priority scheduling pool as the current coroutine, then traverse the third weight value coroutines from the third-level priority scheduling pool as the current coroutine, then return to execute the first weight value coroutines from the first-level priority scheduling pool as the current coroutine, and so on, and repeat the above steps continuously to traverse the coroutines from each priority scheduling pool as the current coroutine.

[0136] When traversing the coroutines from the first-level priority scheduling pool, if the number of coroutines in the first-level priority scheduling pool is less than the first weight value, the remaining coroutines are supplemented through the second-level priority scheduling pool and / or the third-level priority scheduling pool. For example, if the sum of the number of coroutines in the first-level priority scheduling pool and the number of coroutines in the second-level priority scheduling pool is not less than the first weight value, the remaining coroutines are supplemented through the second-level priority scheduling pool. If the sum of the number of coroutines in the first-level priority scheduling pool and the number of coroutines in the second-level priority scheduling pool is less than the first weight value, the remaining coroutines are supplemented through the second-level priority scheduling pool and the third-level priority scheduling pool.

[0137] When traversing the coroutines from the second-level priority scheduling pool, if the number of coroutines in the second-level priority scheduling pool is less than the second weight value, the remaining coroutines can be supplemented by the third-level priority scheduling pool.

[0138] See also Figure 5B As shown, this is a schematic diagram of scheduling coroutines in scheduling pools of different priorities.

[0139] In the first round, 3 coroutines (with the first weight value of 3) are traversed from the first-level priority scheduling pool as the current coroutine. Figure 5A As shown, the first-level priority scheduling pool includes coroutine 0, coroutine 1, coroutine 2, coroutine 3 and coroutine 4. According to the first-in-first-out scheduling order, the coroutine scheduler first traverses coroutine 0 from the first-level priority scheduling pool as the current coroutine, and executes the task corresponding to coroutine 0 through coroutine 0. After the processing of coroutine 0 is completed (coroutine 0 needs to be deleted from the first-level priority scheduling pool), the coroutine scheduler traverses coroutine 1 from the first-level priority scheduling pool as the current coroutine, and executes the task corresponding to coroutine 1 through coroutine 1. After the processing of coroutine 1 is completed, the coroutine scheduler traverses coroutine 2 from the first-level priority scheduling pool as the current coroutine, and executes the task corresponding to coroutine 2 through coroutine 2.

[0140] At this point, the scheduling process of the first-level priority scheduling pool is completed, and 2 coroutines (the second weight value is 2) are traversed from the second-level priority scheduling pool as the current coroutine. Figure 5A As shown, the second-level priority scheduling pool includes coroutine 5, coroutine 6, coroutine 7, coroutine 8 and coroutine 9. According to the first-in-first-out scheduling order, the coroutine scheduler traverses coroutine 5 from the second-level priority scheduling pool as the current coroutine, and executes the task corresponding to coroutine 5 through coroutine 5. After the processing of coroutine 5 is completed, the coroutine scheduler traverses coroutine 6 from the second-level priority scheduling pool as the current coroutine, and executes the task corresponding to coroutine 6 through coroutine 6.

[0141] At this point, the scheduling process of the second-level priority scheduling pool is completed, and one coroutine (the third weight value is 1) is traversed from the third-level priority scheduling pool as the current coroutine. Figure 5AAs shown, the third-level priority scheduling pool includes coroutine 10, coroutine 11, coroutine 12, coroutine 13 and coroutine 14. According to the first-in-first-out scheduling order, the coroutine scheduler traverses coroutine 10 from the third-level priority scheduling pool as the current coroutine, and executes the task corresponding to coroutine 10 through coroutine 10. So far, the first round of scheduling process is completed, and 6 coroutines can be traversed from all priority scheduling pools as the current coroutines to execute the tasks corresponding to these coroutines.

[0142] In the second round, first traverse 3 coroutines from the first-level priority scheduling pool as the current coroutine. Considering that there are only 2 coroutines in the first-level priority scheduling pool, it is necessary to supplement the remaining coroutines through the second-level priority scheduling pool, that is, to supplement 1 coroutine 7. Based on this, the coroutine scheduler traverses coroutine 3 from the first-level priority scheduling pool as the current coroutine, and executes the task corresponding to coroutine 3 through coroutine 3. After the processing of coroutine 3 is completed, the coroutine scheduler traverses coroutine 4 from the first-level priority scheduling pool as the current coroutine, and executes the task corresponding to coroutine 4 through coroutine 4. After the processing of coroutine 4 is completed, the coroutine scheduler traverses coroutine 7 from the second-level priority scheduling pool as the current coroutine, that is, coroutine 7 is the coroutine supplemented to the first-level priority scheduling pool, and the task corresponding to coroutine 7 is executed through coroutine 7.

[0143] At this point, the scheduling process of the first-level priority scheduling pool is completed, and two coroutines are traversed from the second-level priority scheduling pool as the current coroutine. The coroutine scheduler traverses coroutine 8 from the second-level priority scheduling pool as the current coroutine, and executes the task corresponding to coroutine 8 through coroutine 8. The coroutine scheduler traverses coroutine 9 from the second-level priority scheduling pool as the current coroutine, and executes the task corresponding to coroutine 9 through coroutine 9.

[0144] At this point, the scheduling process of the second-level priority scheduling pool is completed, and 1 coroutine is traversed from the third-level priority scheduling pool as the current coroutine. The coroutine scheduler traverses coroutine 11 from the third-level priority scheduling pool as the current coroutine, and executes the task corresponding to coroutine 11 through coroutine 11. At this point, the second round of scheduling process is completed, and 6 coroutines can be traversed from all priority scheduling pools as the current coroutine.

[0145] In the third round, three coroutines are first traversed from the first-level priority scheduling pool as the current coroutine. Considering that the first-level priority scheduling pool is empty and the second-level priority scheduling pool is empty, the remaining coroutines are supplemented through the third-level priority scheduling pool, that is, coroutine 12, coroutine 13 and coroutine 14. Based on this, the coroutine scheduler traverses coroutine 12 from the third-level priority scheduling pool as the current coroutine, and executes the task through coroutine 12. The coroutine scheduler traverses coroutine 13 from the third-level priority scheduling pool as the current coroutine, and executes the task through coroutine 13. The coroutine scheduler traverses coroutine 14 from the third-level priority scheduling pool as the current coroutine, and executes the task through coroutine 14. At this point, the scheduling process of all coroutines is completed.

[0146] After all coroutines are scheduled, if there is no scheduled task, the coroutine scheduler can sleep and wait until a new scheduled task is inserted, and then return to the above process to re-execute the scheduling.

[0147] As can be seen from the above technical solutions, in the embodiment of the present application, by creating multiple coroutines under the same thread, the reconstruction task can be completed by multiple coroutines, and the missing data can be restored, so that adaptive data recovery based on coroutine scheduling can be realized, and the data recovery process is divided into different coroutine tasks to improve the efficiency of data reconstruction. Since the reconstruction task is completed by multiple coroutines instead of by multiple threads, the switching time between multiple threads is saved, and data recovery is completed in a short time, and the efficiency of data recovery is high. In a distributed storage system, the data recovery process is divided into different coroutine tasks through the scheduling optimization of coroutines, which can improve the efficiency of data reconstruction. Through weighted coroutine scheduling, the reading business and the reconstruction business are divided into different priority scheduling pools, and the scheduling method is used to ensure that the impact of the reconstruction business on the read and write business does not exceed a certain proportion, thereby reducing the fluctuation of the business during the reconstruction process. In the scheduling of coroutine tasks, a dynamic weighted method is adopted, and the reconstruction business is fully scheduled in the scene without user business, so as to support an adaptive method and improve the reconstruction efficiency in the scene with no business or a small amount of user business.

[0148] Based on the same application concept as the above method, a data recovery device is proposed in the embodiment of the present application and applied to the main OSD, see Figure 6 FIG. 1 is a schematic diagram of the structure of the device, wherein the device comprises:

[0149] A creation module 61, used to create a first type of coroutine under a thread, create a second type of coroutine under the thread, create a third type of coroutine under the thread, and create a fourth type of coroutine under the thread;

[0150] Processing module 62 is used to obtain a main information list through the first type of coroutine, send a comparison request to the backup OSD through the first type of coroutine, and the comparison request is used to enable the backup OSD to obtain the backup information list; obtain a missing information list through the second type of coroutine, wherein the main information list includes object information of the main OSD, the backup information list includes object information of the backup OSD, and the missing information list includes first object information that exists in the backup information list but does not exist in the main information list; send a pull request to the backup OSD through the third type of coroutine, and the pull request is used to enable the backup OSD to obtain the object to be restored corresponding to the first object information; receive the object to be restored through the fourth type of coroutine, and restore the missing data of the main OSD based on the object to be restored through the fourth type of coroutine.

[0151] In one example, when the processing module 62 obtains the missing information list through the second-class coroutine, it is specifically used to: receive the missing information list sent by the backup OSD through the second-class coroutine; wherein the comparison request includes the main information list, and the missing information list is generated by the backup OSD based on the main information list and the backup information list after the backup OSD obtains the backup information list; or, receive the backup information list sent by the backup OSD through the second-class coroutine, and generate the missing information list based on the main information list and the backup information list through the second-class coroutine.

[0152] The creation module 61 is further used to create a fifth type of coroutine under the thread and a sixth type of coroutine under the thread; the processing module 62 is further used to obtain a recovery information list through the fifth type of coroutine, the recovery information list including second object information that exists in the main information list and does not exist in the backup information list, and obtain the missing object of the backup OSD corresponding to the second object information through the fifth type of coroutine; send the missing object to the backup OSD through the fifth type of coroutine, so that the backup OSD recovers the missing data of the backup OSD based on the missing object; and receive the recovery completion message sent by the backup OSD through the sixth type of coroutine, and determine through the sixth type of coroutine that the backup OSD has successfully recovered the missing data based on the recovery completion message.

[0153] In one example, the device further includes (in Figure 6 Not shown):

[0154] A coroutine scheduling module, used for scheduling the reconstruction task coroutine to a priority scheduling pool corresponding to the reconstruction task after creating the reconstruction task coroutine under the thread; wherein the reconstruction task is used to recover the missing data of the primary OSD and / or the backup OSD; wherein the primary OSD includes at least two priority scheduling pools, and the reconstruction task coroutine includes the first type of coroutine, the second type of coroutine, the third type of coroutine, the fourth type of coroutine, the fifth type of coroutine and the sixth type of coroutine;

[0155] The coroutine scheduling module is used to traverse the coroutines from the at least two priority scheduling pools in sequence as the current coroutine based on the weight value corresponding to each priority scheduling pool, and execute the task corresponding to the current coroutine through the current coroutine; wherein, when the current coroutine is the reconstruction task coroutine, the reconstruction task is executed through the reconstruction task coroutine to restore the missing data of the primary OSD and / or backup OSD.

[0156] In one example, the at least two priority scheduling pools may include a first-level priority scheduling pool, a second-level priority scheduling pool, and a third-level priority scheduling pool, wherein a first weight value corresponding to the first-level priority scheduling pool may be greater than a second weight value corresponding to the second-level priority scheduling pool, and the second weight value may be greater than a third weight value corresponding to the third-level priority scheduling pool; wherein the read-write task may correspond to the first-level priority scheduling pool, the reconstruction task may correspond to the second-level priority scheduling pool, and the remaining tasks other than the read-write task and the reconstruction task may correspond to the third-level priority scheduling pool;

[0157] The coroutine scheduling module is based on the weight value corresponding to each priority scheduling pool, and is specifically used to traverse the coroutines from at least two priority scheduling pools in sequence as the current coroutine: traverse the coroutines with the first weight value from the first-level priority scheduling pool as the current coroutine, traverse the coroutines with the second weight value from the second-level priority scheduling pool as the current coroutine, traverse the coroutines with the third weight value from the third-level priority scheduling pool as the current coroutine, and return to execute the coroutines with the first weight value traversed from the first-level priority scheduling pool as the current coroutine; wherein, when traversing the coroutines from the first-level priority scheduling pool, if the number of coroutines in the first-level priority scheduling pool is less than the first weight value, the remaining coroutines are supplemented by the second-level priority scheduling pool and / or the third-level priority scheduling pool; when traversing the coroutines from the second-level priority scheduling pool, if the number of coroutines in the second-level priority scheduling pool is less than the second weight value, the remaining coroutines are supplemented by the third-level priority scheduling pool.

[0158] Based on the same application concept as the above method, an electronic device (such as a main OSD) is proposed in the embodiment of the present application, see Figure 7As shown, the electronic device includes: a processor 71 and a machine-readable storage medium 72, the machine-readable storage medium 72 stores machine-executable instructions that can be executed by the processor 71; the processor 71 is used to execute the machine-executable instructions to implement the data recovery method disclosed in the above example of this application.

[0159] Based on the same application concept as the above method, an embodiment of the present application also provides a machine-readable storage medium, on which a number of computer instructions are stored. When the computer instructions are executed by a processor, the data recovery method disclosed in the above example of the present application can be implemented.

[0160] The above-mentioned machine-readable storage medium may be any electronic, magnetic, optical or other physical storage device, which may contain or store information, such as executable instructions, data, etc. For example, the machine-readable storage medium may be: RAM (Radom Access Memory), volatile memory, non-volatile memory, flash memory, storage drive (such as hard disk drive), solid state drive, any type of storage disk (such as CD, DVD, etc.), or similar storage medium, or a combination thereof.

[0161] Based on the same application concept as the above method, an embodiment of the present application further provides a computer program product, which may include a computer program. When the computer program is executed by a processor, it implements the data recovery method disclosed in the above example of the present application.

[0162] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the embodiments of the present application may adopt the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0163] The above is only an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included in the scope of the claims of the present application.

Claims

1. A data recovery method, characterized in that: Applied to the master OSD, the method comprises: Obtaining the main information list through a first type of coroutine created under the thread, and sending a comparison request to the backup OSD through the first type of coroutine, wherein the comparison request is used to enable the backup OSD to obtain the backup information list; Obtaining a missing information list through a second type of coroutine created under the thread; wherein the main information list includes object information of the main OSD, the backup information list includes object information of the backup OSD, and the missing information list includes first object information that exists in the backup information list but does not exist in the main information list; Sending a pull request to the backup OSD through a third-type coroutine created under the thread, wherein the pull request is used to enable the backup OSD to obtain the object to be restored corresponding to the first object information; The object to be restored is received by a fourth type of coroutine created under the thread, and the missing data of the primary OSD is restored based on the object to be restored by the fourth type of coroutine.

2. The method according to claim 1, characterized in that The obtaining of the missing information list by the second type of coroutine created under the thread includes: Receiving the missing information list sent by the backup OSD through the second-type coroutine; wherein the comparison request includes the main information list, and the missing information list is generated by the backup OSD based on the main information list and the backup information list after obtaining the backup information list; Alternatively, the backup information list sent by the backup OSD is received through the second-type coroutine, and the missing information list is generated based on the main information list and the backup information list through the second-type coroutine.

3. The method according to claim 1, characterized in that After the fourth type of coroutine is used to recover the missing data of the primary OSD based on the object to be recovered, the method further includes: Acquire a recovery information list through a fifth type of coroutine created under the thread, the recovery information list including second object information that exists in the main information list and does not exist in the backup information list, and acquire a missing object of the backup OSD corresponding to the second object information through the fifth type of coroutine; The missing object is sent to the backup OSD through the fifth type of coroutine, so that the backup OSD recovers the missing data of the backup OSD based on the missing object.

4. The method according to claim 3, characterized in that After sending the missing object to the backup OSD through the fifth type of coroutine, the method also includes: receiving a recovery completion message sent by the backup OSD through a sixth type of coroutine created under the thread, and determining through the sixth type of coroutine that the backup OSD has successfully recovered the missing data based on the recovery completion message.

5. The method according to any one of claims 1 to 4, characterized in that: The master OSD includes at least two priority scheduling pools, and the method further includes: After creating a reconstruction task coroutine under the thread, scheduling the reconstruction task coroutine to a priority scheduling pool corresponding to the reconstruction task, the reconstruction task is used to recover the missing data of the primary OSD and / or the backup OSD; wherein the reconstruction task coroutine includes the first type of coroutine, the second type of coroutine, the third type of coroutine, the fourth type of coroutine, the fifth type of coroutine and the sixth type of coroutine; Based on the weight value corresponding to each priority scheduling pool, traverse the coroutines from the at least two priority scheduling pools in sequence as the current coroutine, and execute the task corresponding to the current coroutine through the current coroutine; Wherein, when the current coroutine is the reconstruction task coroutine, the reconstruction task is executed by the reconstruction task coroutine to restore the missing data of the primary OSD and / or the backup OSD.

6. The method according to claim 5, characterized in that The at least two priority scheduling pools include a first-level priority scheduling pool, a second-level priority scheduling pool, and a third-level priority scheduling pool, wherein a first weight value corresponding to the first-level priority scheduling pool is greater than a second weight value corresponding to the second-level priority scheduling pool, and the second weight value is greater than a third weight value corresponding to the third-level priority scheduling pool; The read / write task corresponds to the first-level priority scheduling pool, the reconstruction task corresponds to the second-level priority scheduling pool, and the remaining tasks except the read / write task and the reconstruction task correspond to the third-level priority scheduling pool; The step of traversing the coroutines from the at least two priority scheduling pools in sequence as the current coroutine based on the weight value corresponding to each priority scheduling pool includes: traversing the coroutines with a first weight value from the first-level priority scheduling pool as the current coroutine, traversing the coroutines with a second weight value from the second-level priority scheduling pool as the current coroutine, traversing the coroutines with a third weight value from the third-level priority scheduling pool as the current coroutine, and returning to execute traversing the coroutines with a first weight value from the first-level priority scheduling pool as the current coroutine; Among them, when traversing the coroutines from the first-level priority scheduling pool, if the number of coroutines in the first-level priority scheduling pool is less than the first weight value, the remaining coroutines are supplemented by the second-level priority scheduling pool and / or the third-level priority scheduling pool; when traversing the coroutines from the second-level priority scheduling pool, if the number of coroutines in the second-level priority scheduling pool is less than the second weight value, the remaining coroutines are supplemented by the third-level priority scheduling pool.

7. A data recovery device, characterized in that: Applied to the main OSD, the device comprises: A creation module, used to create a first type of coroutine under a thread, create a second type of coroutine under the thread, create a third type of coroutine under the thread, and create a fourth type of coroutine under the thread; A processing module, used to obtain a main information list through the first type of coroutine, send a comparison request to a backup OSD through the first type of coroutine, the comparison request is used to enable the backup OSD to obtain a backup information list; obtain a missing information list through the second type of coroutine, wherein the main information list includes object information of the main OSD, the backup information list includes object information of the backup OSD, and the missing information list includes first object information that exists in the backup information list but does not exist in the main information list; send a pull request to the backup OSD through the third type of coroutine, the pull request is used to enable the backup OSD to obtain an object to be restored corresponding to the first object information; receive the object to be restored through the fourth type of coroutine, and restore the missing data of the main OSD based on the object to be restored through the fourth type of coroutine.

8. The device according to claim 7, characterized in that When the processing module obtains the missing information list through the second type of coroutine, it is specifically used to: receive the missing information list sent by the backup OSD through the second type of coroutine; wherein the comparison request includes the main information list, and the missing information list is generated by the backup OSD based on the main information list and the backup information list after the backup OSD obtains the backup information list; or, receive the backup information list sent by the backup OSD through the second type of coroutine, and generate the missing information list based on the main information list and the backup information list through the second type of coroutine; Among them, the creation module is also used to create a fifth type of coroutine under the thread, and to create a sixth type of coroutine under the thread; the processing module is also used to obtain a recovery information list through the fifth type of coroutine, the recovery information list includes second object information that exists in the main information list and does not exist in the backup information list, and obtain the missing object of the backup OSD corresponding to the second object information through the fifth type of coroutine; send the missing object to the backup OSD through the fifth type of coroutine, so that the backup OSD recovers the missing data of the backup OSD based on the missing object; and receive the recovery completion message sent by the backup OSD through the sixth type of coroutine, and determine through the sixth type of coroutine that the backup OSD has successfully recovered the missing data based on the recovery completion message.

9. The device according to claim 7 or 8, characterized in that The device also includes: A coroutine scheduling module, used for scheduling the reconstruction task coroutine to a priority scheduling pool corresponding to the reconstruction task after creating the reconstruction task coroutine under the thread; wherein the reconstruction task is used to recover the missing data of the primary OSD and / or the backup OSD; wherein the primary OSD includes at least two priority scheduling pools, and the reconstruction task coroutine includes the first type of coroutine, the second type of coroutine, the third type of coroutine, the fourth type of coroutine, the fifth type of coroutine and the sixth type of coroutine; The coroutine scheduling module is used to traverse the coroutines from the at least two priority scheduling pools in sequence as the current coroutine based on the weight value corresponding to each priority scheduling pool, and execute the task corresponding to the current coroutine through the current coroutine; wherein, when the current coroutine is the reconstruction task coroutine, the reconstruction task is executed through the reconstruction task coroutine to recover the missing data of the primary OSD and / or the backup OSD; Among them, the at least two priority scheduling pools include a first-level priority scheduling pool, a second-level priority scheduling pool and a third-level priority scheduling pool, the first weight value corresponding to the first-level priority scheduling pool is greater than the second weight value corresponding to the second-level priority scheduling pool, and the second weight value is greater than the third weight value corresponding to the third-level priority scheduling pool; the read-write task corresponds to the first-level priority scheduling pool, the reconstruction task corresponds to the second-level priority scheduling pool, and the remaining tasks other than the read-write task and the reconstruction task correspond to the third-level priority scheduling pool; the coroutine scheduling module, based on the weight value corresponding to each priority scheduling pool, sequentially traverses the coroutines from at least two priority scheduling pools as the current coroutine, and is specifically used to: select a coroutine from the first-level priority scheduling pool The first priority pool is traversed as the current coroutine, the second weight value coroutines are traversed as the current coroutine from the second-level priority scheduling pool, the third weight value coroutines are traversed as the current coroutine from the third-level priority scheduling pool, and the coroutines traversed as the current coroutine from the first-level priority scheduling pool are returned for execution. The coroutines traversed as the current coroutines are traversed from the first-level priority scheduling pool; when traversing the coroutines from the first-level priority scheduling pool, if the number of coroutines in the first-level priority scheduling pool is less than the first weight value, the remaining coroutines are supplemented by the second-level priority scheduling pool and / or the third-level priority scheduling pool; when traversing the coroutines from the second-level priority scheduling pool, if the number of coroutines in the second-level priority scheduling pool is less than the second weight value, the remaining coroutines are supplemented by the third-level priority scheduling pool.

10. An electronic device, characterized in that: include: a processor and a machine-readable storage medium storing machine-executable instructions executable by the processor; The processor is used to execute machine executable instructions to implement the method described in any one of claims 1-6.