Task processing method and storage device

Through the design of main coroutines and backup coroutines, the task processing speed reduction caused by frequent switching of coroutines in high concurrency scenarios is solved, and more efficient task processing is achieved, especially in storage devices, which significantly improves task completion efficiency.

CN120276811APending Publication Date: 2025-07-08CHENGDU HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410022477.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-05
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

In high concurrency scenarios, frequent switching of coroutines leads to a reduced task processing speed. The existing technology cannot effectively reduce the number of coroutines and scheduling time, affecting task processing efficiency.

Method used

The design of main coroutines and standby coroutines is adopted, and the coroutines with a single processing core handles multiple tasks, and the coroutines are upgraded as the main coroutines when the main coroutines are blocked, reducing the number of coroutines and the number of switches, and optimizing task queue operations using lock-free queues and lock-up mechanisms.

Benefits of technology

By reducing the number of coroutines and frequent switching, shorten the task completion time, improve task processing efficiency, and reduce resource overhead and task interruption risks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120276811A_ABST
    Figure CN120276811A_ABST
Patent Text Reader

Abstract

The invention provides a task processing method which can reduce the number of coroutines of a processing core and reduce the coroutine scheduling duration so as to improve the task processing efficiency. The method comprises the steps that after a main coroutine and a standby coroutine are created in a processing core, the main coroutine is used for polling a task queue of the processing core. The invention further provides a storage device, a computer readable storage medium and a computer program product capable of implementing the method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technologies, and in particular, to a task processing method and a storage device. Background Art

[0002] A coroutine, also known as a lightweight thread, can be understood as a function that implements cross-execution of tasks.

[0003] A task processing method is roughly as follows: after receiving an input / output (IO) task, multiple IO subtasks are generated according to the IO task, a coroutine is created for each IO subtask, and the IO subtasks are processed using the coroutines in the coroutine order.

[0004] In a high-concurrency scenario, a large number of IO subtasks are generated. At this time, a large number of coroutines are required to process the IO subtasks, and frequent switching of coroutines will cause the task processing speed to decrease. Summary of the Invention

[0005] This application provides a task processing method, which can reduce the number of coroutines of a processing core, reduce the coroutine scheduling duration, and improve the task processing efficiency. This application also provides a storage device, a storage device cluster, a computer-readable storage medium, and a computer program product that can implement the above method.

[0006] In a first aspect, a task processing method is provided, and the method includes: after creating a main coroutine and a backup coroutine in a first processing core, polling a task queue of the first processing core using the main coroutine.

[0007] Implementing in this way can use a single coroutine of a processing core to process multiple tasks. Since it is not necessary to create a coroutine for each task, the number of coroutines in a single processing core can be reduced, and it is not necessary to frequently switch coroutines when processing multiple tasks. Therefore, the task completion duration can be shortened. The main coroutine and the backup coroutine are scheduled by a thread of the first processing core, and the backup coroutine is used to be upgraded to the main coroutine when the main coroutine is blocked. In this way, the backup coroutine can be quickly upgraded to the main coroutine, so that the processing core can keep the main coroutine polling the task queue and reduce task interruption.

[0008] In some possible implementation manners, the task processing method of the present application further includes: before polling the task queue of the first processing core using the main coroutine, dividing the IO tasks obtained by the first processing core into multiple groups of IO subtasks, and adding at least one group of IO subtasks to the first task queue of the first processing core. Among them, the IO task is a read request or a write request, and the first task queue is a lock-free queue. The IO tasks obtained by the first processing core are sent by the host, or sent by other processing cores of the storage device, or sent by other storage devices. Since the main coroutine and the standby coroutine are scheduled by the same thread and the standby coroutine is upgraded only when the main coroutine is blocked, the main coroutine and the standby coroutine do not conflict, so a lock-free queue can be used, thereby saving the operations of locking and unlocking, and improving the task processing efficiency.

[0009] In some possible implementation manners, the task processing method of the present application further includes: obtaining multiple IO subtasks through a second processing core of the storage device; locking the second task queue of the first processing core; adding the first IO task stored by the second processing core to the second task queue of the first processing core; when the main coroutine is in a waiting state, waking up the main coroutine; using the second IO subtask stored by the second processing core as the to-be-processed IO subtask; adding the to-be-processed IO subtask to the second task queue; when the main coroutine is in a waiting state, waking up the main coroutine; when the main coroutine is in a running state and the to-be-processed IO subtask is not the last IO subtask, updating the to-be-processed IO subtask to the next IO subtask of the to-be-processed IO subtask, and triggering the step of adding the to-be-processed IO subtask to the second task queue; when the main coroutine is in a running state and the to-be-processed IO subtask is the last IO subtask, unlocking the second task queue. The second processing core and the first processing core belong to the same storage device and the second processing core is different from the first processing core. The tasks submitted across cores usually have delays. In order to poll this task, this task needs to be added to the locked second task queue. Multiple IO tasks entering the queue require one locking and one unlocking. Compared with performing one locking and one unlocking for each task entering the queue, this implementation manner can reduce the number of lockings and unlockings and improve the task processing efficiency. During the process of the task entering the queue, when the main coroutine is in a running state, the step of waking up the main coroutine is not executed. Compared with the method of waking up the main coroutine for each task entering the queue, this implementation manner can reduce the number of times of waking up the main coroutine and improve the queueing speed.

[0010] In some possible implementation manners, when the main coroutine is blocked, the main coroutine is modified into an ordinary coroutine and the backup coroutine is modified into the main coroutine; the modified main coroutine polls the task queue; a coroutine is created and the created coroutine is set as the backup coroutine. When the coroutine state of the main coroutine is the running state and the main coroutine executes a down operation, a sleep operation, a yield operation, or a combination of the above operations, it is determined that the main coroutine is blocked. At this time, the backup coroutine is upgraded to the main coroutine, so that one main coroutine is maintained by the processing core to poll the task queue, preventing task processing interruption.

[0011] In some possible implementation manners, when the task polled by the main coroutine is a read task and the read task misses the cache, the main coroutine is blocked.

[0012] In some possible implementation manners, the task processing method of the present application further includes: using a third processing core to receive short tasks sent by the host and using the third processing core to process the short tasks. A short task is a task whose processing duration is less than the CPU time slice, such as a network request task and a database query task. The third processing core and the first processing core belong to the same storage device and the third processing core is different from the first processing core. In this way, different processing cores can be used to execute short tasks and IO tasks, accelerating the processing efficiency of short tasks.

[0013] In a second aspect, a storage device is provided. The storage device includes a processor, and the processor is used to create a main coroutine and a backup coroutine in a first processing core. The main coroutine is used to poll the task queue of the first processing core; the backup coroutine is used to be upgraded to the main coroutine when the main coroutine is blocked.

[0014] In some possible implementation manners, the processor is further used to divide the IO tasks obtained by the first processing core into multiple groups of IO subtasks; add at least one group of IO subtasks to a first task queue of the first processing core.

[0015] In some possible implementation manners, the processor is further used to obtain multiple IO subtasks through a second processing core; lock a second task queue of the first processing core; add a first IO subtask stored in the second processing core to the second task queue; when the main coroutine is in a waiting state, wake up the main coroutine; use a second IO subtask stored in the second processing core as a to-be-processed IO subtask; add the to-be-processed IO subtask to the second task queue; when the main coroutine is in a waiting state, wake up the main coroutine; when the main coroutine is in a running state and the to-be-processed IO subtask is not the last IO subtask, update the to-be-processed IO subtask to the next IO subtask of the to-be-processed IO subtask, triggering the step of adding the to-be-processed IO subtask to the second task queue; when the main coroutine is in a running state and the to-be-processed IO subtask is the last IO subtask, unlock the second task queue.

[0016] In some possible implementations, the processor is further configured to, when the main coroutine is blocked, modify the main coroutine into an ordinary coroutine and modify the standby coroutine into the main coroutine; create a coroutine, and set the created coroutine as the standby coroutine; and use the modified main coroutine to poll the task queue.

[0017] In some possible implementations, the processor is further configured to block the main coroutine when the task polled by the main coroutine is a read task and the read task misses the cache.

[0018] In some possible implementations, the processor is further configured to use the third processing core to receive short tasks sent by the host and use the third processing core to process the short tasks.

[0019] For the glossary in the second aspect, the steps and technical effects executed by the processor can refer to the corresponding descriptions in the first aspect.

[0020] A third aspect provides a storage device, including a processor and a memory; the processor is configured to execute instructions stored in the memory, so that the storage device executes the method described in the first aspect or any possible implementation manner of the first aspect.

[0021] A fourth aspect provides a storage device cluster, which includes at least one storage device, and each storage device includes a processor and a memory; the processors of the at least one storage device are configured to execute instructions stored in the memories of the at least one storage device, so that the storage device cluster executes the method described in the first aspect or any possible implementation manner of the first aspect.

[0022] A fifth aspect provides a computer-readable storage medium, which includes computer program instructions. When the computer program instructions are executed by a computing device, the computing device executes the method described in the first aspect or any possible implementation manner of the first aspect.

[0023] A sixth aspect provides a computer program product including instructions. When the instructions are run by a computing device, the computing device is caused to execute the method described in the first aspect or any possible implementation manner of the first aspect.

[0024] A seventh aspect provides a chip system, which includes a processor and a memory connected to each other. The processor can run instructions stored in the memory, so that the chip system executes the method described in the first aspect or any possible implementation manner of the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 It is a schematic diagram of an application scenario in an embodiment of the present application;

[0026] Figure 2 It is another schematic diagram of an application scenario in an embodiment of the present application;

[0027] Figure 3 is a schematic diagram of an existing task processing method;

[0028] Figure 4 is a flowchart of the task processing method in an embodiment of the present application;

[0029] Figure 5 is a schematic diagram of the task processing method in an embodiment of the present application;

[0030] Figure 6 is a schematic diagram of the processing method of cross-core tasks in an embodiment of the present application;

[0031] Figure 7 is another schematic diagram of the task processing method in an embodiment of the present application;

[0032] Figure 8 is a schematic diagram of the task processing method in a distributed storage system in an embodiment of the present application;

[0033] Figure 9 is a schematic diagram of the task processing method in a centralized storage system in an embodiment of the present application;

[0034] Figure 10 is a structural diagram of a storage device in an embodiment of the present application;

[0035] Figure 11 is another structural diagram of the storage device in an embodiment of the present application. Detailed implementation manners

[0036] The task processing method of the present application can be applied to a data processing system including a storage device, such as a distributed storage system or a centralized storage system.

[0037] Refer to Figure 1 , in one embodiment, as Figure 1 shown, the data processing system includes a computing device 100 and a storage cluster connected through a network 120. The number of computing devices 100 in the data processing system can be, but is not limited to, 3. The computing device 100 can be, but is not limited to, a desktop computer, a tablet computer, a mobile phone, a virtual machine. The storage cluster includes one or more servers 110( Figure 1 three servers 110 are shown in, but not limited to three servers 110), and the servers 110 can communicate with each other. The server 110 is a device that has both computing power and storage capacity, such as a server, a desktop computer, etc. Exemplarily, the server 110 can be an ARM server or an X86 server. In terms of hardware, as Figure 1As shown, the server 110 includes at least a processor 112, a memory 113, a network card 114, and a hard disk 115. The processor 112, the memory 113, the network card 114, and the hard disk 115 are connected through a bus. Among them, the processor 112 and the memory 113 are used to provide computing resources. Specifically, the processor 112 is a central processing unit (CPU), which is used to process data access requests from outside the server 110 (computing devices, application servers, or other servers), and is also used to process requests generated inside the server 110. Exemplarily, when the processor 112 receives a write data request, it temporarily stores the data in these write data requests in the memory 113. When the total amount of data in the memory 113 reaches a certain threshold, the processor 112 sends the data stored in the memory 113 to the hard disk 115 for persistent storage. In addition, the processor 112 is also used to calculate or process data, such as metadata management, deduplication, data compression, data verification, virtualized storage space, and address translation, etc. Figure 1 Only one processor 112 is shown in the figure. In actual applications, the number of processors 112 is often multiple. Among them, one processor 112 has one or more processing cores. The number of processors and the number of processing cores are not limited in this embodiment.

[0038] Memory 113 refers to the internal memory that directly exchanges data with the processor. It can read and write data at any time and is very fast. It serves as the temporary data memory for the operating system or other running programs. Memory includes at least two types of memories. For example, memory can be either random access memory or read only memory (ROM). For instance, random access memory can be dynamic random access memory (DRAM) or storage class memory (SCM). DRAM is a semiconductor memory and, like most random access memories (RAM), belongs to a type of volatile memory device. SCM is a composite storage technology that combines the characteristics of traditional storage devices and memories. Storage class memory can provide faster read and write speeds than hard disks, but its access speed is slower than that of DRAM and it is also cheaper than DRAM. However, DRAM and SCM are only exemplary illustrations in this embodiment. Memory can also include other random access memories, such as static random access memory (SRAM), etc. For read only memory, for example, it can be programmable read only memory (PROM), erasable programmable read only memory (EPROM), etc. Additionally, memory 113 can also be a dual in-line memory module or a dual in-line memory module (DIMM), that is, a module composed of DRAM, or it can be a solid state disk (SSD). In practical applications, multiple memories 113 and different types of memories 113 can be configured in the storage server 110. The number and type of memories 113 are not limited in this embodiment. In addition, the memory 113 can be configured to have a power retention function. The power retention function means that when the system experiences a power outage and then powers on again, the data stored in the memory 113 will not be lost. Memory with the power retention function is called non-volatile memory.

[0039] The network card 114 is used to communicate with the computing device 100 or other servers 110. The hard disk 115 is used to provide storage resources, such as storing data. It can be a disk or other types of storage media, such as a solid state disk or a shingled magnetic recording hard disk, etc.

[0040] Figure 2The data processing system shown includes a computing device 200, a switch 210, and a storage system 220. The number of computing devices 200 and switches 210 is not limited to Figure 2 the number shown. The storage system 220 is a centralized storage system with separated disk control. Its engine 221 includes one or more controllers, such as controller 0 and controller 1. Each controller includes a front-end interface 225, a CPU 223, a memory 224, and a back-end interface 226. The engine 221 may not have hard disk slots, and the hard disk 234 needs to be placed in the hard disk enclosure 230. The back-end interface 226 communicates with the hard disk enclosure 230. The back-end interface 226 exists in the form of an adapter card in the engine 221. Two or more back-end interfaces 226 can be used simultaneously on one engine 221 to connect multiple hard disk enclosures. Alternatively, the adapter card can also be integrated on the motherboard. In this case, the adapter card can communicate with the processor 122 through the PCIe bus. It should be noted that Figure 2 one engine 221 is shown. However, in practical applications, the storage system may include two or more engines 221, and redundancy or load balancing is performed among multiple engines 221.

[0041] The hard disk enclosure 230 includes a control unit 231 and several hard disks 234. The control unit 231 can have various forms. In one case, the hard disk enclosure 230 belongs to an intelligent disk enclosure, such as Figure 2As shown, the control unit 231 includes a CPU and a memory. The CPU is used to perform operations such as address conversion and data reading and writing. The memory is used to temporarily store the data to be written to the hard disk 234, or the data read from the hard disk 234 to be sent to the controller. In another case, the control unit 231 is a programmable electronic component, such as a data processing unit (DPU). The DPU has the generality and programmability of the CPU, but is more specialized and can operate efficiently on network data packets, storage requests, or analysis requests. The DPU is distinguished from the CPU by a greater degree of parallelism (requiring the processing of a large number of requests). Optionally, the DPU here can also be replaced by a graphics processing unit (GPU), a neural-network processing units (NPU), or other processing chips. Generally, the number of control units 231 can be one, or two or more. When the hard disk enclosure 230 includes at least two control units 231, there may be an ownership relationship between the hard disk 234 and the control unit 231. If there is an ownership relationship between the hard disk 234 and the control unit 231, then each control unit can only access the hard disk that belongs to it, which often involves forwarding read / write data requests between the control units 231, resulting in a longer data access path. In addition, if the storage space is insufficient, when adding a new hard disk 234 to the hard disk enclosure 230, it is necessary to re-bind the ownership relationship between the hard disk 234 and the control unit 231, and the operation is complex, resulting in poor scalability of the storage space. Therefore, in another implementation, the functions of the control unit 231 can be offloaded to the network card 232. In other words, in this implementation, the hard disk enclosure 230 does not have a control unit 231 inside, but the network card 232 is used to complete data reading and writing, address conversion, and other computing functions. At this time, the network card 232 is a smart network card. It can include a CPU and a memory. In some application scenarios, the network card 232 may also have a persistent memory medium, such as persistent memory (PM), or non-volatile random access memory (NVRAM), or phase change memory (PCM), etc. The CPU is used to perform operations such as address conversion and data reading and writing. The memory is used to temporarily store the data to be written to the hard disk 234, or the data read from the hard disk 234 to be sent to the controller. It can also be a programmable electronic component, such as a DPU. Optionally, the DPU here can also be replaced by a GPU, an NPU, or other processing chips.There is no ownership relationship between the network card 232 and the hard disks 234 in the hard disk enclosure 230. The network card 232 can access any hard disk 234 in the hard disk enclosure 230. Therefore, it is more convenient to expand the hard disk when the storage space is insufficient.

[0042] According to the type of communication protocol between the engine 221 and the hard disk enclosure 230, the hard disk enclosure 230 may be a serial attached small computer system interface (SAS) hard disk enclosure, a non-volatile memory host controller interface specification express (NVMe) hard disk enclosure, an IP hard disk enclosure, or other types of hard disk enclosures. The SAS hard disk enclosure uses the SAS 3.0 protocol, and each enclosure supports 25 SAS hard disks. The engine 221 is connected to the hard disk enclosure 230 through an on-board SAS interface or a SAS interface module. The NVMe hard disk enclosure is like a complete computer system, and the NVMe hard disks are inserted into the NVMe hard disk enclosure. The NVMe hard disk enclosure is then connected to the engine 221 through an RDMA port.

[0043] It should be noted that the storage cluster of the present application is not limited to Figure 1 the storage cluster shown, and it can also be other forms of storage clusters. The centralized storage system of the present application is not limited to Figure 2 the storage system shown, and it can also be a storage system with integrated disk control.

[0044] To ensure that the software can run on different types of devices, multiple functions of the software can be implemented by different software components, that is, software componentization. When the software includes multiple software components, the scheduling overhead will also increase. The following compares the time overhead of the software for performing IO tasks and the time overhead of performing IO tasks by software components when the functions of the software are implemented by software components.

[0045] The software component for processing the IO stream can be called an IO stream component. For software that does not include an IO stream component, the storage device creates 2 coroutines according to the instructions of the software, and the duration of processing the IO tasks through the 2 coroutines is 180 microseconds.

[0046] In an example of the existing task processing method, for software that includes an IO stream component, the storage device creates coroutines 1 to 5 according to the instructions of the software, as Figure 3As shown in the figure. The IO task is split into IO subtasks 1 to IO subtask 5, and then coroutine 1 is called to process IO subtask 1. After IO subtask 1 hits the cache, coroutine 2 is called to process IO subtask 2. After IO subtask 2 hits the cache, coroutine 3 is called to process IO subtask 3. In the case where IO subtask 3 does not hit the cache, a disk read operation is performed on the hard disk, and coroutine 4 is called to process IO subtask 4. After IO subtask 4 hits the cache, coroutine 5 is called to process IO subtask 5. It can be seen that the processing core creates a coroutine for each IO subtask. In a high-concurrency scenario, there will be a large number of IO subtasks. At this time, a large number of coroutines need to be created on a single processing core, which requires a large amount of computing resources. Taking the time overhead of creating a coroutine and the time overhead of calling a coroutine as 10 microseconds, the duration of processing the IO task by 5 coroutines is 210 microseconds. It can be seen from this that adding 3 coroutines to process the task will increase the time overhead by 16%.

[0047] Regarding the problem of large time overhead existing in the above method, the present application provides a task processing method, which can use a single coroutine to process multiple tasks, thereby reducing the number of coroutines in a single processing core, reducing the number of coroutine switches, thereby shortening the task completion duration and improving the task completion efficiency. The task processing method of the present application will be introduced below. Refer to Figure 4 , an embodiment of the task processing method of the present application includes the following steps:

[0048] S401. Create a main coroutine and a backup coroutine in the first processing core.

[0049] In this embodiment, the processor of the storage device includes multiple processing cores, and the first processing core is any one of the multiple processing cores. When the computing device runs a synchronous programming program and the synchronous programming program includes coroutines, the computing device can send a coroutine creation instruction to the storage device, and the storage device creates a main coroutine and a backup coroutine in the first processing core according to the coroutine creation instruction. The main coroutine is used to poll the task queue, and the backup coroutine is used to upgrade to the main coroutine when the main coroutine is blocked. It should be noted that when the main coroutine is blocked, a new coroutine cannot be created using the main coroutine. The backup coroutine needs to be created before the main coroutine executes the task. The backup coroutine is also called a slave coroutine.

[0050] S402. Use the main coroutine to poll the task queue of the first processing core.

[0051] The first processing core may include multiple task queues, and the task queue may be, but is not limited to, a circular task queue. Polling means periodically accessing each task queue, extracting tasks from the task queue, and then processing the tasks. After the main coroutine polls the task queue, when there are no tasks in all the task queues, the main coroutine enters the waiting state. When a new task enters the queue, the main coroutine is awakened.

[0052] In this embodiment, in the initialization stage, each processing core is configured with 2 coroutines. Compared with the existing method of creating one coroutine for each task, the task processing method of the present application can reduce the number of coroutines, thereby reducing resource overhead. Moreover, by using the main coroutine to poll the task queue of the processing core, it is not necessary to frequently switch coroutines when processing multiple tasks, so the time overhead of switching coroutines can be reduced, thereby improving the task processing efficiency.

[0053] Secondly, the main coroutine and the backup coroutine are scheduled by the same thread of the first processing core. When the main coroutine is blocked, the backup coroutine is upgraded to the main coroutine, so that the processing core can keep the main coroutine polling the task queue and reduce the task interruption duration.

[0054] The task queue of the processing core can be divided into a local submission task queue and a cross-core submission task queue according to the task source. The number of the local submission task queue and the number of the cross-core submission task queue can be set according to the actual situation, and the present application does not make any limitation. The tasks in the local submission task queue are the tasks generated by the current processing core, and the tasks in the cross-core submission task queue are the tasks obtained from other processing cores. The following introduces the tasks submitted locally. In an optional embodiment, before S402, the task processing method of the present application further includes: dividing the IO tasks obtained by the first processing core into multiple groups of IO subtasks, and adding at least one group of IO subtasks to the first task queue of the first processing core.

[0055] In this embodiment, the IO task can be a read request or a write request. When the first processing core receives a read request, after dividing the read request into multiple groups of subtasks, one or more groups of subtasks are added to the first task queue, and then the main coroutine of the first processing core polls and processes the tasks in the first task queue. The subtasks in the task queue can be regarded as a complete task. The subtasks are generated by the first processing core, and its corresponding first task queue is the local submission task queue. Since the main coroutine and the backup coroutine are scheduled by the same thread and the backup coroutine is upgraded only when the main coroutine is blocked, the main coroutine and the backup coroutine will not conflict. Therefore, a lock-free first task queue can be used, and the operations of locking and unlocking can be saved when tasks are enqueued and dequeued, thereby improving the task processing efficiency. The processing process of the write request is similar to that of the read request and will not be elaborated here.

[0056] For the convenience of understanding, the following introduces the task processing method in the IO concurrency scenario. Refer to Figure 5, in one example, coroutine 1 and coroutine 2 are created on a processing core. Coroutine 1 is the main coroutine and coroutine 2 is the backup coroutine. The IO task is split into IO subtasks 1 to 5, and then coroutine 1 is called to process IO subtask 1. After IO subtask 1 hits the cache, coroutine 1 is called to process IO subtask 2. After IO subtask 2 hits the cache, coroutine 1 is called to process IO subtask 3. In the case where IO subtask 3 misses the cache, coroutine 1 is modified to an ordinary coroutine, a disk read operation is performed on the hard disk of the storage device, and coroutine 2 is upgraded to the main coroutine. Coroutine 2 is called to process IO subtask 4. After IO subtask 4 hits the cache, coroutine 2 is called to process IO subtask 5. The time overhead for creating a coroutine and the time overhead for calling a coroutine are both taken as 10 microseconds. For software including an IO stream component, the storage device creates 2 coroutines according to the data processing method of the present application. When the main coroutine is unblocked, the task completion duration for using 1 coroutine to process 5 IO subtasks is about 170 microseconds. When the main coroutine is blocked, the task completion duration for using 2 coroutines to process 5 IO subtasks is about 180 microseconds. Compared with Figure 3 the time overhead of the method shown (i.e., 210 microseconds), the task completion duration can be significantly reduced.

[0057] When the number of IO subtasks of a processing core is greater than a preset number, the IO subtasks can be sent to other processing cores of the storage device or other storage devices. The addition of other processing cores to the IO subtasks is introduced below. In another alternative embodiment, before S402, the task processing method of the present application further includes:

[0058] Step A: Obtain a plurality of IO subtasks through the second processing core of the storage device.

[0059] Specifically, the IO subtasks obtained by the second processing core can be generated according to the IO tasks sent by the computing device or the IO subtasks sent by other processing cores.

[0060] Step B: Lock the second task queue of the first processing core.

[0061] Step C: Add the first IO subtasks stored in the second processing core to the second task queue.

[0062] Step D: When the main coroutine is in a waiting state, wake up the main coroutine.

[0063] Step E: Use the second IO subtasks stored in the second processing core as the to-be-processed IO subtasks.

[0064] Step F: Add the to-be-processed IO subtasks to the second task queue.

[0065] Step G: When the main coroutine is in a waiting state, wake up the main coroutine.

[0066] Step H: When the main coroutine is in the running state and the to-be-processed IO subtask is not the last IO subtask, update the to-be-processed IO subtask to the next IO subtask, and trigger Step F.

[0067] Step I: When the main coroutine is in the running state and the to-be-processed IO subtask is the last IO subtask among the multiple IO subtasks, unlock the second task queue.

[0068] In this embodiment, the IO tasks from other processing cores belong to asynchronous tasks, and the period for receiving asynchronous tasks is longer than the time slice of the CPU. Therefore, the tasks submitted across cores need to be added to the locked second task queue. Adding multiple IO tasks to the second task queue for one locking and one unlocking can reduce the number of lockings and unlockings compared with locking and unlocking once for each task enqueueing, thus improving the task processing efficiency.

[0069] During the task enqueueing process, when the main coroutine is in the running state, the step of waking up the main coroutine is not executed. Compared with the method of waking up the main coroutine for each task enqueueing, this implementation can reduce the number of times of waking up the main coroutine and improve the enqueueing speed.

[0070] Refer to Figure 6 , in an example, the task processing method of the present application further includes: Processing core 2 is the processing core that sends the IO subtasks. Processing core 2 sends the IO subtask group to other processing cores. For example, it sends the IO subtask group 1 to processing core 1 and the IO subtask group 2 to processing core 3. Add the IO subtask group 1 to task queue 1. When the coroutine 11 is blocked during the task execution process, change the coroutine 11 from the main coroutine to an ordinary coroutine, change the coroutine 12 from the standby coroutine to the main coroutine. After creating the coroutine 13, set the coroutine 13 as the standby coroutine, and then let the coroutine 12 execute the task extracted from the task queue 1. After adding the IO subtask group 2 to task queue 2, let the coroutine 31 execute the task extracted from the task queue 2, and the coroutine 32 serves as the standby coroutine.

[0071] The main coroutine may be blocked during the task processing, and at this time, a new coroutine needs to be used to execute the task. The following introduces the task processing method including the task scheduling process. Refer to Figure 7 , in an example, the task processing method of the present application includes the following steps:

[0072] S701. Create coroutines.

[0073] After the computing device sends a coroutine creation instruction to the storage device, the storage device creates coroutine 1 and coroutine 2 in the processing cores of the storage device according to the instruction, sets coroutine 1 as the main coroutine, and sets coroutine 2 as the standby coroutine.

[0074] S702. When both the run queue and the task queue are empty queues, set the coroutine status of coroutine 1 to the waiting state.

[0075] S703. The computing device sends a task to the storage device.

[0076] S704. When the run queue is empty and the task queue has tasks, move the tasks in the task queue to the run queue.

[0077] S705. Set the coroutine status of coroutine 1 to the running state.

[0078] S706. Wake up the standby coroutine.

[0079] When the standby coroutine is coroutine 2, wake up coroutine 2, and at this time coroutine 2 is in the waiting state.

[0080] S707. Coroutine 1 executes the task taken out from the run queue.

[0081] S708. When coroutine 1 is blocked, set coroutine 1 as an ordinary coroutine and set coroutine 2 as the main coroutine.

[0082] When the task polled by the main coroutine is a read task and the read task misses the cache, the main coroutine performs a down operation to block the main coroutine. When the main coroutine is in the running state and the main coroutine performs a down operation, a sleep operation, or a yield operation, that is, the main coroutine is blocked. In this application, the main coroutine flag can be set as a global variable, the coroutine flag of coroutine 2 can be set as the main coroutine flag, and at the same time, the main coroutine flag of coroutine 1 is cancelled. When the coroutine number of the target coroutine is less than the coroutine number of the main coroutine, the target coroutine is an ordinary coroutine. When the coroutine number of the target coroutine is greater than the coroutine number of the main coroutine, the target coroutine is a standby coroutine.

[0083] It can be seen from S708 that when the main coroutine is blocked, the standby coroutine can be upgraded to the main coroutine, so that the processing core maintains a main coroutine to poll the run queue to prevent task processing interruption. After the storage device sets coroutine 2 as the main coroutine, S709 is triggered.

[0084] S709. Create a coroutine.

[0085] S710. Set coroutine 3 as a standby coroutine.

[0086] After the storage device creates coroutine 3, set coroutine 3 as a standby coroutine.

[0087] S711. When both the run queue and the task queue are empty queues, set the coroutine status of coroutine 2 to the waiting state.

[0088] S712. When the running queue is empty and there are tasks in the task queue, move the tasks in the task queue to the running queue.

[0089] S713. Set the coroutine status of coroutine 2 to the running status.

[0090] S714. Wake up the standby coroutine.

[0091] As can be seen from S710, at this time the standby coroutine is coroutine 3, that is, wake up coroutine 3, and at this time coroutine 3 is in the waiting state.

[0092] S715. Coroutine 2 executes the tasks taken out from the running queue.

[0093] When coroutine 2 is blocked, upgrade coroutine 3 to the main coroutine, and then create a new coroutine as the standby coroutine. The specific process can refer to S708 to S710, and the subsequent steps after the coroutine is blocked can be deduced by analogy, which will not be elaborated here.

[0094] The following combines Figure 1 the storage system shown to introduce the task processing method of the present application. Refer to Figure 8 , the computing device 100 sends the IO task to the server 110A through the network 120. The processing core 1121 of the server 110A divides the IO task into multiple groups of IO subtasks. After enqueuing the IO subtasks, use the main coroutine to poll the task queue. When the main coroutine is blocked, send the IO subtask to the hard disk 115 of the server 110A, read the data corresponding to the IO subtask from the hard disk 115 of the server 110A, the processing core 1121 of the server 110A modifies the standby coroutine to the main coroutine, and then uses the main coroutine to poll the task queue.

[0095] After dividing the IO task into multiple groups of IO subtasks, a part of the IO subtasks can be assigned to other processing cores, such as the processing core 1121 of the server 110B. After enqueuing the received IO subtasks, the processing core 1121 of the server 110B uses the main coroutine to poll the task queue. When the main coroutine is blocked, the processing core 1121 of the server 110B sends the IO subtask to the hard disk 115 of the server 110B, reads the data corresponding to the IO subtask from the hard disk 115 of the server 110B, modifies the standby coroutine to the main coroutine, and then uses the main coroutine to poll the task queue.

[0096] The following combines Figure 2 the storage system shown to introduce the task processing method of the present application. Refer to Figure 9Upon this, the computing device 200 sends the IO tasks to the storage system 220 via the switch 210. The processing core 01 in the controller 0 of the storage system 220 divides the IO tasks into multiple groups of IO subtasks. After enqueuing the IO subtasks, it uses the main coroutine to poll the task queue. When the main coroutine is blocked, it sends the IO subtasks to the hard disk enclosure 230. The hard disk enclosure 230 sends the data corresponding to the IO subtasks to the processing core 01. The processing core 01 changes the standby coroutine to the main coroutine, and then uses the main coroutine to poll the task queue.

[0097] After dividing the IO tasks into multiple groups of IO subtasks, a part of the IO subtasks can be allocated to the processing cores of other controllers, such as the processing core 11 of the controller 1. After enqueuing the received IO subtasks, the processing core 11 uses the main coroutine to poll the task queue. When the main coroutine is blocked, it sends the IO subtasks to the hard disk enclosure 230. The hard disk enclosure 230 sends the data corresponding to the IO subtasks to the processing core 11. The processing core 11 changes the standby coroutine to the main coroutine, and then uses the main coroutine to poll the task queue.

[0098] In an alternative embodiment, the task processing method of the present application further includes: using a third processing core to receive short tasks sent by the host and using the third processing core to process the short tasks.

[0099] A short task is a task whose processing duration is less than the CPU time slice, such as a network request task or a database query task. The third processing core and the first processing core belong to the same storage device and are different from each other. In this way, different processing cores can be used to execute short tasks and IO tasks, thereby improving the processing efficiency of short tasks.

[0100] The following introduces the storage device of the present application. Refer to Figure 10 In one embodiment, the storage device 1000 of the present application includes one or more processors 1001, and each processor 1001 includes one or more processing cores; the processor 1001 is used to create a main coroutine and a standby coroutine in the first processing core. The main coroutine is used to poll the task queue of the first processing core; the standby coroutine is used to be upgraded to the main coroutine when the main coroutine is blocked.

[0101] In an alternative embodiment, the processor 1001 is further used to divide the IO tasks into multiple groups of IO subtasks and add at least one group of IO subtasks to the first task queue of the first processing core. The first task queue is a lock-free queue.

[0102] In an alternative embodiment, the processor 1001 is further configured to obtain a plurality of IO subtasks through the second processing core; lock the second task queue of the first processing core; add the first IO task stored in the second processing core to the second task queue; when the main coroutine is in a waiting state, wake up the main coroutine; use the second IO subtask stored in the second processing core as the to-be-processed IO subtask; add the to-be-processed IO subtask to the second task queue; when the main coroutine is in a waiting state, wake up the main coroutine; when the main coroutine is in a running state and the to-be-processed IO subtask is not the last IO subtask, update the to-be-processed IO subtask to the next IO subtask of the to-be-processed IO subtask, and trigger the step of adding the to-be-processed IO subtask to the second task queue; when the main coroutine is in a running state and the to-be-processed IO subtask is the last IO subtask, unlock the second task queue.

[0103] In an alternative embodiment, the processor 1001 is further configured to, when the main coroutine is blocked, modify the main coroutine into an ordinary coroutine and modify the backup coroutine into the main coroutine; create a coroutine, set the created coroutine as the backup coroutine, and use the modified main coroutine to poll the task queue.

[0104] In an alternative embodiment, the processor 1001 is further configured to block the main coroutine when the task polled by the main coroutine is a read task and the read task misses the cache.

[0105] See Figure 11 , in an embodiment, the storage device 1100 of the present application includes: a bus 1102, a processor 1104, a memory 1106, and a communication interface 1108. The processor 1104, the memory 1106, and the communication interface 1108 communicate with each other through the bus 1102. The computing device 1100 may be a server or a terminal device. It should be understood that the present application does not limit the number of processors and memories in the computing device 1100.

[0106] The bus 1102 may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The bus may be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, Figure 11 only one line is shown in, but it does not mean that there is only one bus or one type of bus. The bus 1104 may include a path for transmitting information between various components (for example, the memory 1106, the processor 1104, the communication interface 1108) of the computing device 1100.

[0107] The processor 1104 may include any one or more of processors such as a CPU, GPU, MP, or DSP. The memory 1106 may include volatile memory, such as random access memory (RAM). The processor 1104 may also include non-volatile memory (NVM), such as ROM, flash memory, a hard disk drive (HDD), or a solid state drive (SSD).

[0108] The memory 1106 stores executable program code, and the processor 1104 executes the executable program code to respectively implement the functions of the foregoing processor 1001, thereby implementing the task processing method. That is, the memory 1106 stores instructions for executing the task processing method.

[0109] The communication interface 1108 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the storage device 1100 and other devices or a communication network.

[0110] The embodiments of the present application further provide a storage device cluster. The storage device cluster includes at least one storage device. The storage device may be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the storage device may also be a terminal device such as a desktop computer, a laptop computer, or a smart phone. The storage device cluster includes at least one storage device 1100. The memory 1106 in one or more storage devices 1100 in the storage device cluster may store the same instructions for executing the task processing method.

[0111] The embodiments of the present application further provide a computer program product containing instructions. The computer program product may be a software or program product containing instructions that can run on a computing device or be stored in any available medium. When the computer program product runs on at least one computing device, at least one computing device is caused to execute the task processing method.

[0112] The embodiments of the present application further provide a computer-readable storage medium. The computer-readable storage medium may be any available medium that a computing device can store or a data storage device such as a data center containing one or more available mediums. The available medium may be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid state drive), etc. The computer-readable storage medium includes instructions that instruct a computing device to execute the task processing method, or instruct a computing device to execute the task processing method.

[0113] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the protection scope of the technical solutions of the embodiments of the present invention.

Claims

1. A task processing method, characterized in that, The method is applied to a storage device, which includes at least one processing core, and the method includes: Create a main coroutine and a backup coroutine in the first processing core; Use the main coroutine to poll the task queue of the first processing core, and the backup coroutine is used to upgrade to the main coroutine when the main coroutine is blocked.

2. The method according to claim 1, wherein The method further includes: Divide the IO tasks obtained by the first processing core into multiple groups of IO subtasks; Add at least one group of the IO subtasks to the first task queue of the first processing core, and the first task queue is a lock-free queue.

3. The method according to claim 1, wherein The method further includes: Obtain multiple IO subtasks through the second processing core of the storage device; Lock the second task queue of the first processing core; Add the first IO subtask stored in the second processing core to the second task queue; Wake up the main coroutine when the main coroutine is in a waiting state; Use the second IO subtask stored in the second processing core as the to-be-processed IO subtask; Add the to-be-processed IO subtask to the second task queue; Wake up the main coroutine when the main coroutine is in a waiting state; When the main coroutine is in a running state and the to-be-processed IO subtask is not the last IO subtask, update the to-be-processed IO subtask to the next IO subtask of the to-be-processed IO subtask, and trigger the step of adding the to-be-processed IO subtask to the second task queue; When the main coroutine is in a running state and the to-be-processed IO subtask is the last IO subtask, unlock the second task queue.

4. The method according to any one of claims 1 to 3, characterized in that, The method further includes: When the main coroutine is blocked, modify the main coroutine to an ordinary coroutine and modify the backup coroutine to the main coroutine; Create a coroutine and set the created coroutine as the backup coroutine; Use the modified main coroutine to poll the task queue.

5. The method according to claim 4, wherein The method further includes: When the task polled by the main coroutine is a read task and the read task misses the cache, block the main coroutine.

6. A storage device, characterized in that, It includes: A processor for creating a main coroutine and a backup coroutine in the first processing core, where the main coroutine is used to poll the task queue of the first processing core; the backup coroutine is used to upgrade to the main coroutine when the main coroutine is blocked.

7. The storage device according to claim 6, wherein The processor is further used to divide the IO tasks obtained by the first processing core into multiple groups of IO subtasks; add at least one group of the IO subtasks to the first task queue of the first processing core, and the first task queue is a lock-free queue.

8. The storage device according to claim 6, wherein, The processor is further used to obtain multiple IO subtasks through the second processing core; lock the second task queue of the first processing core; add the first IO subtask stored in the second processing core to the second task queue; wake up the main coroutine when the main coroutine is in a waiting state; Use the second IO subtask stored in the second processing core as the to-be-processed IO subtask; Add the to-be-processed IO subtask to the second task queue; Wake up the main coroutine when the main coroutine is in a waiting state; When the main coroutine is in a running state and the to-be-processed IO sub-task is not the last IO sub-task, update the to-be-processed IO sub-task to the next IO sub-task of the to-be-processed IO sub-task, and trigger the step of adding the to-be-processed IO sub-task to the second task queue; When the main coroutine is in a running state and the to-be-processed IO sub-task is the last IO sub-task, unlock the second task queue.

9. The storage device according to any one of claims 6 to 8, characterized in that, The processor is further configured to, when the main coroutine is blocked, modify the main coroutine to a normal coroutine and modify the standby coroutine to the main coroutine; create a coroutine and set the created coroutine as the standby coroutine; use the modified main coroutine to poll the task queue.

10. The storage device according to claim 9, wherein The processor is further configured to block the main coroutine when the task polled by the main coroutine is a read task and the read task misses the cache.

11. A storage device, characterized in that, It includes a processor and a memory, and the processor is configured to execute the instructions stored in the memory so that the storage device executes the method according to any one of claims 1 to 5.

12. A computer-readable storage medium, characterized in that, It includes computer program instructions, and when the computer program instructions are executed by a computing device, the computing device executes the method according to any one of claims 1 to 5.

13. A computer program product comprising instructions, characterized in that, When the instruction is run by a computing device, it causes the computing device to execute the method according to any one of claims 1 to 5.