Task processing method, device and system and related equipment

By converting data migration tasks at the logical storage unit level into migration requests at the physical address level that the storage device can recognize, the problem of low execution efficiency of CPUs and heterogeneous processors in storage systems is solved, achieving more efficient data migration and performance improvement.

CN120832078APending Publication Date: 2025-10-24HUAWEI TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202410501154.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-04-24
Publication Date
2025-10-24

AI Technical Summary

Technical Problem

In storage systems, when the CPU handles data migration tasks for east-west traffic, it needs to frequently execute the I/O stack and network stack, resulting in low performance. Even after offloading the tasks to heterogeneous processors, low efficiency still exists.

Method used

The task processing unit converts data migration tasks at the logical storage unit level into migration requests at the physical address level that the storage device can recognize. The data migration tasks are then executed directly by the storage device, reducing reliance on the CPU and heterogeneous processors and optimizing the data transmission path.

Benefits of technology

It improves the execution efficiency of data migration tasks, reduces the number of IO stacks and network stacks, and frees up CPU computing power to improve storage system performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120832078A_ABST
    Figure CN120832078A_ABST
Patent Text Reader

Abstract

The invention discloses a task processing method, device and system and related equipment, and relates to the technical field of storage. The task processing device receives a data migration task used for migrating target data indicated by the source logic address to a target logic address, wherein the data migration task comprises the source logic address and the target logic address; converting the source logic address into a first physical address, converting the destination logic address into a second physical address, and generating a migration request carrying the first physical address and the second physical address; and sending a migration request to the first storage device to indicate the first storage device to migrate the target data from the first physical address to the second physical address. Therefore, the storage device executes the process of transmitting and processing the data among different storage devices, so that the number of IO stacks and network stacks needing to be executed can be effectively reduced, and the task execution efficiency is improved. And moreover, the CPU does not need to intervene, so that the CPU can improve the performance of the storage system based on more computing power.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of storage, and in particular, to a task processing method, device, system and related equipment. BACKGROUND

[0002] With the development of storage technology, storage media such as non-volatile memory express (NVMe) are widely used in storage systems. In a storage system, there is usually east-west traffic, that is, after a central processing unit (CPU) in the storage system reads data on a storage medium, the CPU will write the data (such as updated data) back to the storage medium. In this process, the CPU needs to frequently read the metadata corresponding to multiple data blocks on the storage medium, and perform data reading and data writing on the storage medium according to the read metadata, which makes the CPU execute more input / output (IO) stacks (for requesting metadata or data) and network stacks (for transmitting data), thereby occupying more computing power of the CPU, resulting in lower performance of the storage system.

[0003] Currently, the CPU processing data migration tasks of east-west traffic are usually offloaded to a heterogeneous processor of the storage system, such as the CPU can instruct a data processing unit (DPU) in the storage system to process the data migration tasks of east-west traffic.

[0004] However, although this way of offloading data migration tasks to a heterogeneous processor releases the computing power of the CPU, the heterogeneous processor still needs to execute more IO stacks and network stacks to complete the data migration tasks, which will result in that the execution efficiency of the data migration tasks of east-west traffic is still low. SUMMARY

[0005] The present application provides a task processing method to improve the efficiency of processing data migration tasks of east-west traffic. In addition, the present application also provides a corresponding task processing device, a storage system, a computing device, a computer readable storage medium and a computer program product.

[0006] In a first aspect, the present application provides a task processing method, which is applied to a storage system. The storage system includes a plurality of storage devices, each of which can be a solid state disk or the like, and the plurality of storage devices are used to persistently store data. The plurality of storage devices can be located in the same storage server (or referred to as the same storage node), or the plurality of storage devices can be located in different storage servers. The method can be executed by a corresponding task processing apparatus. Specifically, the task processing apparatus receives a data migration task, which includes a source logical address and a destination logical address, and is used to migrate target data indicated by the source logical address to the destination logical address. The source logical address belongs to a first logical storage unit, and the destination logical address belongs to a second logical storage unit. The number of source logical addresses can be one or more, and the number of destination logical addresses can also be one or more. When the source logical address is multiple, the multiple source logical addresses can belong to the same first logical storage unit, or belong to multiple different first logical storage units. Similarly, when the destination logical address is multiple, the multiple destination logical addresses can belong to the same second logical storage unit, or belong to multiple different second logical storage units. The first logical storage unit corresponds to a first physical storage space in the plurality of storage devices, and the second logical storage unit corresponds to a second physical storage space in the plurality of storage devices. The first physical storage space and the second physical storage space can be physical storage spaces in the same storage device, or can be physical storage spaces in different storage devices. After receiving the data migration task, the task processing apparatus converts the source logical address into a first physical address in the first physical storage space, converts the destination logical address into a second physical address in the second physical storage space, and generates a migration request carrying the first physical address and the second physical address, so as to convert the data migration task including the logical address into the migration request including the physical address. Finally, the task processing apparatus sends the migration request to a first storage device in the plurality of storage devices, so as to instruct the first storage device to migrate the target data from the first physical address to the second physical address by using the migration request.

[0007] Thus, in the process of executing the data migration task, the task processing apparatus can convert the data migration task into a migration request including the first physical address and the second physical address recognizable by the storage device, that is, convert the data migration task of the logical storage unit level semantics into a task of the storage device level semantics understandable by the storage device, which enables the storage system to perform the process of transferring and processing data between different storage devices by the storage device, that is, to complete the data migration task. In this process, the CPU or the heterogeneous processor in the storage system can not need to intervene, which can effectively reduce the number of IO stacks and network stacks required to be executed by the storage system, thereby improving the execution efficiency of the data migration task. Moreover, since the CPU does not need to intervene in the process of executing the data migration task, the computing power of the CPU can be released, so that the CPU can serve the user IO of the upper layer based on more computing power, thereby effectively improving the performance of the storage system.

[0008] In a possible implementation, the data migration task is specifically used to indicate migration of target data from a source logical address to a destination logical address; or the data migration task is specifically used to indicate backup of the target data to the destination logical address; or the data migration task is used to indicate reconstruction of first data from the target data based on an error correction code (EC) technology and storage of the first data to a second physical address.

[0009] In a possible implementation, the task processing apparatus can further determine a service type corresponding to the data migration task; when the service type is a first type, the migration request is used to instruct the first storage device to migrate the target data of the first physical address to the second physical address; when the service type is a second type, the migration request is used to instruct the first storage device to backup the target data to the second physical address; and when the service type is a third type, the migration request is used to instruct the first storage device to reconstruct data based on the EC technology according to the target data of the first physical address. Thus, the task processing apparatus generates different types of migration requests according to different data migration tasks to instruct the first storage device to perform different types of migration operations.

[0010] In a possible implementation, the target data is located in the first storage device, and the task processing apparatus can further generate a data view corresponding to the first storage device according to a storage location of the target data in the first storage device, where the data view is used to describe a location distribution of the target data in the first storage device. Thus, when generating the migration request, the task processing apparatus can specifically generate the migration request according to the data view corresponding to the first storage device. In this way, the task processing apparatus can plan and control the first storage device to migrate data indicated by different first physical addresses according to the data view, so as to improve the efficiency of migrating data in different physical storage spaces of the first storage device, that is, improve the execution efficiency of the data migration task, for example, to avoid converged read traffic in the EC backup scenario, or to reduce the number of IOs generated in the data backup process.

[0011] In a possible implementation, the data migration task is used to indicate migration of the target data from a source logical address to a destination logical address, and the migration request generated by the task processing apparatus is used to indicate migration of the target data from a storage location indicated by the first physical address to a storage location indicated by the second physical address. At this time, the task processing apparatus can control the first storage device to migrate data from one storage location to another storage location by using the migration request, and accordingly, the storage location indicated by the first physical address does not store data after the data migration is completed.

[0012] In a possible implementation, the target data is stored in the storage system based on the EC technology, and the migration request generated by the task processing apparatus is specifically used to instruct the first storage device to determine m pieces of data from the target data, where m is an integer greater than 1; generate n pieces of check data based on the EC technology, where n is a positive integer; and split different data in a stripe composed of the m pieces of data and the n pieces of check data to different storage devices for storage. Because the first storage device generates check data for the m pieces of data locally, and splits and stores the m pieces of data and the check data in other storage devices, the target data after migration can be ensured to be accurate and reliable in the storage system by using the newly generated check data. Meanwhile, the first storage device does not need to read the target data from other storage devices in the process of migrating the target data, so as to avoid converged read traffic.

[0013] In a possible implementation, the data migration task is used to instruct to backup the target data to the destination logical address; then, the task processing apparatus can specifically determine, according to the data view corresponding to the first storage device, a plurality of second data in the target data that are migrated to the same physical address in the second physical storage space when generating the migration request according to the data view corresponding to the first storage device; and generate, for the plurality of second data, the migration request used to instruct to backup the plurality of second data in batches. In this way, for the plurality of second data, the task processing apparatus can control the first storage device to backup the plurality of second data in batches to reduce the number of IOs generated in the data migration process and improve the data migration efficiency.

[0014] In a possible implementation, the target data is stored in the storage system based on the EC technology, and the migration request is used to instruct the first storage device to reconstruct the first data from the target data and store the first data in the second physical address in the first storage device. In this way, the process of obtaining data on other storage devices and performing data reconstruction can be performed by the first storage device, without the intervention of the CPU or the heterogeneous processor in the storage system, which can effectively reduce the number of IO stacks and network stacks required to be executed by the storage system, thereby improving the execution efficiency of data reconstruction. Moreover, the CPU can serve the user IO of the upper layer based on more computing power, thereby effectively improving the performance of the storage system.

[0015] In a possible implementation, the plurality of storage devices interact with each other based on an interconnection bus; or the plurality of storage devices interact with each other based on a network card, an intelligent network card or a data processing unit (DPU). In this way, the plurality of storage devices can be connected to each other based on the bus, the network card, the intelligent network card or the DPU. Further, some or each of the storage devices can be configured with hardware that can support the storage device to perform corresponding processing operations according to the received migration request.

[0016] In a second aspect, the present application provides a task processing apparatus applied to a storage system, the storage system comprising a plurality of storage devices for persistently storing data, the task processing apparatus comprising: an obtaining module configured to receive a data migration task, the data migration task comprising a source logical address and a destination logical address, the data migration task being configured to migrate target data indicated by the source logical address to the destination logical address, the source logical address belonging to a first logical storage unit, the destination logical address belonging to a second logical storage unit, the first logical storage unit corresponding to a first physical storage space in the plurality of storage devices, and the second logical storage unit corresponding to a second physical storage space in the plurality of storage devices; a converting module configured to convert the source logical address into a first physical address in the first physical storage space, and convert the destination logical address into a second physical address in the second physical storage space; a generating module configured to generate a migration request, the migration request carrying the first physical address and the second physical address; and a communication module configured to send the migration request to a first storage device in the plurality of storage devices, the migration request being configured to instruct the first storage device to migrate the target data from the first physical address to the second physical address.

[0017] In a possible implementation, the data migration task is configured to instruct migration of the target data from the source logical address to the destination logical address, or the data migration task is configured to instruct backup of the target data to the destination logical address, or the data migration task is configured to instruct reconstruction of first data from the target data based on an EC (Error Correction Code) technology and storage of the first data to the second physical address.

[0018] In a possible implementation, the generating module is further configured to determine a service type corresponding to the data migration task, wherein when the service type is a first type, the migration request is configured to instruct the first storage device to migrate the target data at the first physical address to the second physical address; when the service type is a second type, the migration request is configured to instruct the first storage device to backup the target data to the second physical address; and when the service type is a third type, the migration request is configured to instruct the first storage device to reconstruct data based on the EC technology according to the target data at the first physical address.

[0019] In a possible implementation, the target data is located in the first storage device, and the generating module is further configured to generate a data view corresponding to the first storage device according to a storage location of the target data in the first storage device, the data view being configured to describe a location distribution of the target data in the first storage device, and the generating module is configured to generate the migration request according to the data view corresponding to the first storage device.

[0020] In a possible implementation, the data migration task is used to indicate migration of the target data from a source logical address to a destination logical address, and the migration request is used to indicate migration of the target data from a storage location indicated by the first physical address to a storage location indicated by the second physical address.

[0021] In a possible implementation, the target data is stored in the storage system based on an EC technology; and the migration request is specifically used to indicate: determining m data from the target data, m being an integer greater than 1; generating n check data based on the EC technology, n being a positive integer; and splitting different data in a stripe composed of the m data and the n check data to different storage devices for storage.

[0022] In a possible implementation, the data migration task is used to indicate backup of the target data to the destination logical address; and the generation module is specifically used to, when generating the migration request, determine, according to the data view corresponding to the first storage device, a plurality of second data in the target data that are to be migrated to the same physical address in the second physical storage space; and generate, for the plurality of second data, the migration request, the migration request being used to indicate batch backup of the plurality of second data.

[0023] In a possible implementation, the target data is stored in the storage system based on an EC technology, and the migration request is used to instruct the first storage device to reconstruct the first data from the target data and store the first data in the second physical address in the first storage device.

[0024] In a possible implementation, the plurality of storage devices interact with each other based on an interconnection bus; or the plurality of storage devices interact with each other based on a network card, an intelligent network card, or a DPU (data processing unit).

[0025] In a third aspect, the present application provides a storage system, comprising a task processing apparatus and a plurality of storage devices, the plurality of storage devices being used to persistently store data, and the plurality of storage devices comprising a first storage device; wherein the task processing apparatus is used to execute the task processing method in the first aspect or any implementation manner of the first aspect, and the first storage device is further used to execute the migration request to migrate the target data from the first physical address to the second physical address.

[0026] In a fourth aspect, the present application provides a computing device, comprising a processor and a memory. The processor and the memory are in communication with each other. The processor is configured to execute instructions stored in the memory, so that the computing device performs the task processing method in the first aspect or any implementation manner of the first aspect. It should be noted that the memory can be integrated into the processor or independent of the processor. The computing device can further comprise a bus. The processor is connected to the memory through the bus. The memory can include a readable memory and a random access memory.

[0027] In a fifth aspect, the present application provides a computer readable storage medium, which stores instructions. When the instructions are executed on a computing device, the computing device performs the operation steps of the task processing method in the first aspect or any implementation manner of the first aspect.

[0028] In a sixth aspect, the present application provides a computer program product comprising instructions, which, when executed on a computing device, cause the computing device to perform the operation steps of the task processing method in the first aspect or any implementation manner of the first aspect.

[0029] On the basis of the implementation manners of the above aspects, the present application can be further combined to provide more implementation manners. BRIEF DESCRIPTION OF DRAWINGS

[0030] Figure 1 Structure diagram of an example storage system provided by the present application;

[0031] Figure 2a Schematic diagram of performing garbage collection tasks for heterogeneous processors;

[0032] Figure 2b Schematic diagram of reducing the number of IOs and the number of data transmissions;

[0033] Figure 3 Flowchart of a task processing method provided by the present application;

[0034] Figure 4 Schematic diagram of various data migration tasks provided by the present application;

[0035] Figure 5 Schematic diagram of calculating a physical address according to a logical address;

[0036] Figure 6 Schematic diagram of a generated data view provided by the present application;

[0037] Figure 7 Schematic diagram of generating a migration request and batch backup data according to a data view in a replica scenario;

[0038] Figure 8 An example diagram for generating a migration request according to a data view and splitting stripe data in an EC scenario;

[0039] Figure 9 An example diagram of a task processing apparatus provided in the present application;

[0040] Figure 10 An example diagram of a hardware structure of a computing device provided in the present application. DETAILED DESCRIPTION

[0041] The terms “first”, “second”, etc. in the description and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily indicate a specific order or sequence. It should be understood that the terms thus used can be interchanged under appropriate circumstances, and are merely a distinguishing way used in the description of the embodiments of the present application to describe the objects of the same attribute.

[0042] The technical solutions in the present application will be described below in conjunction with the accompanying drawings provided in the present application.

[0043] Referring to Figure 1 , an example diagram of a storage system 10 is shown. As Figure 1 shown, the storage system 10 includes a plurality of storage servers, each of which is configured with one or more storage devices, Figure 1 For example, the storage system 10 includes two storage servers and N storage devices, i.e., storage server 101, storage server 102, and storage device 1 to storage device N. N is a positive integer greater than 1.

[0044] As Figure 1 shown, the storage server 101 includes a CPU 1011, storage device 1 to storage device M, and M is a positive integer less than N. The CPU 1011 can run a task processing apparatus 200. The task processing apparatus 200 can be implemented by software, for example, the task processing apparatus 200 can be implemented by at least one of an engine, a virtual machine, a component, etc. At this time, the task processing apparatus 200 can also be referred to as a disk move engine (DME). In actual application, the storage server 101 can also include other necessary components, such as memory, cache, etc. (not shown in Figure 1 Further, the CPU 1011 can also run a cache subsystem 201, an index subsystem 202, a persistence subsystem 203, etc. Figure 1As shown, each subsystem can be used to generate a data migration task. The structure of the storage server 102 is similar to that of the storage server 101, which can be seen from Figure 1 As shown, no further elaboration is made.

[0045] The storage devices 1 to N are used to persistently store data. Exemplarily, each storage device can be a solid state drive (SSD), a hard disk drive (HDD), a disk, etc. And the different storage devices can communicate through wired or wireless means. For example, the different storage devices can interact data based on an interconnection bus such as a universal bus (UB) or a compute express link (CXL) bus, etc. For example, each storage device can be configured with a network card, a smart network interface card (SNIC), or a DPU, so that the different storage devices can interact data through the network card, the SNIC, or the DPU, etc., which is not limited in this regard.

[0046] And the different storage devices located in the same storage server can communicate through an internal bus in the storage server, while the storage devices located in different storage servers can communicate through a wired network or a wireless network, i.e., cross-node communication.

[0047] Generally, the cache subsystem 201, the index subsystem 202, and the persistence subsystem 203 can all generate data migration tasks based on logical storage units at the business logic layer. The logical storage unit, which can also be referred to as an object, is a basic unit for the business logic layer to provide data storage in the storage system 10, and specifically refers to a piece or multiple pieces of storage area in the storage device being abstracted into a logical entity capable of storing data for the business logic layer. For example, a data migration task can include the address of the data to be migrated in a logical storage unit A, a data migration operation, and the address in a logical storage unit B after data migration, which is used to indicate that the data in the logical storage unit A is migrated to the logical storage unit B. Since the data migration task includes the related information of the logical storage unit, the data migration task belongs to a task at the logical storage unit level, or can be referred to as an object semantic task. Thus, at the business logic layer, data can be read, written, and managed in units of logical storage units, and accordingly, at the underlying storage device, the storage system 10 will perform corresponding processing on the data in the physical address mapped by the logical storage unit according to the processing operation for the logical storage unit.

[0048] Each logical storage unit can have a unique identifier, which can be a unique number assigned to the logical storage unit in the storage system 10, for example, and each logical storage unit corresponds to one or more segments of storage area on a storage device. In actual applications, different logical storage units can correspond to storage areas of the same size in the storage device. For example, a logical storage unit, such as a logical unit number (LUN), includes a plurality of Plogs, and each Plog can be used to store 64 megabytes (MB) of data, mapping a 64 MB storage space in the storage device, so that the 64 MB of data can be stored on the storage device. In addition, each Plog in the LUN has a unique identifier for identification.

[0049] For example, the data migration task generated by the cache subsystem 201 can be to persist the data in the cache of the storage system 10. When the data is backed up in the storage system 10, the data migration task generated by the cache subsystem 201 instructs to store the data in the cache to a storage device, and to back up the newly stored data in the storage device to other storage devices, thereby generating east-west traffic.

[0050] The data migration task generated by the index subsystem 202, for example, can be a garbage collection task, which is used to automatically identify and reclaim the storage space occupied by objects not used by programs, thereby improving the utilization of storage space. In this process, the valid data to be reclaimed is first read from the storage device, and then written back to the storage device, thereby generating east-west traffic.

[0051] The data migration task generated by the persistence subsystem 203, for example, can be a data migration task for the scenario of expanding storage capacity (i.e., capacity expansion), which is used to migrate data on part of the storage device to the storage device as capacity expansion, thereby generating east-west traffic.

[0052] For the data migration task corresponding to the east-west traffic, if the CPU in the storage system 10 executes the data migration task, it will occupy too much computing power of the CPU, so that more computing power of the CPU cannot serve the user's IO in the storage system 10, affecting the performance of the storage system 10. If a heterogeneous processor (such as an acceleration card) is configured in the storage system 10 and the heterogeneous processor is used to execute the data migration task corresponding to the east-west traffic, although the computing power of the CPU can be released, the heterogeneous processor still needs to execute more IO stacks and network layers to execute the data migration task, thereby resulting in low efficiency of the data migration task.

[0053] Take the garbage collection task as an example, as shown in Figure 2a Similar to the CPU executing the garbage collection task, the heterogeneous processor sends an address request (executes an IO stack once) to the persistent layer (such as the persistent subsystem 203) based on the garbage collection task generated by the index layer, the address request including a logical address in the logical storage unit corresponding to the valid data, for requesting a physical address on the storage device (the physical address is used to persistently store the valid data on the storage device) mapped by the logical address. Wherein, the persistent layer records the mapping relationship between each logical address in the logical storage unit and the physical address on the storage device. Then, the heterogeneous processor accesses the storage device according to the obtained physical address (executes an IO stack again) to obtain the valid data participating in the garbage collection. In this process, each block performs at least 2 IO stacks, and the valid data is transmitted from the persistent layer to the index layer, and there is a data transmission, that is, at least 1 network stack is executed. Next, the heterogeneous processor accesses the persistent layer based on a similar manner to realize writing the recycled valid data into the storage device again, which can be the original storage device (another storage area in the storage device where the valid data is stored) or a new storage device. In this way, at least 2 IO stacks and at least 1 network stack are executed again, as shown in Figure 2b Therefore, in the entire garbage collection process, the heterogeneous processor needs to execute at least 4 IO stacks and at least 2 network stacks, as shown in Figure 2b In actual application, the heterogeneous processor usually performs garbage collection on the data in multiple data blocks, then the heterogeneous processor executes a large number of IO stacks and network stacks, which leads to a relatively low overall efficiency of executing the garbage collection task.

[0054] Based on this, the present application provides a task processing method, aiming to improve the execution efficiency of data migration tasks for east-west traffic. In specific implementation, the task processing device 200 obtains a data migration task, which may be generated by the cache subsystem 201, the index subsystem 202, or the persistent subsystem 203. At this time, the obtained data migration task belongs to a logical storage unit level semantic task, that is, the data migration task includes a source logical address in the logical storage unit 1, a destination logical address in the logical storage unit 2, and is used to indicate that the target data indicated by the source logical address is migrated to the destination logical address. Then, instead of directly executing the data migration task, the task processing device 200 converts the data migration task into a task that can be understood by the storage device, or can be called a storage device level semantic task (such as a disk level semantic task). Specifically, the task processing device 200 converts the source logical address into a physical address 2 in the physical storage space 1 mapped by the logical storage unit 1, and converts the destination logical address into a physical address 2 in the physical storage space 2 mapped by the logical storage unit 2. Then, the task processing device 200 generates a migration request carrying the physical address 1 and the physical address 2, and sends the migration request to the first storage device in the N storage devices included in the storage system 10, which can refer to one or more storage devices. At this time, since the addresses included in the migration request are all physical addresses on the storage device, the first storage device can generally recognize and execute the migration request (that is, can recognize the disk level semantic task). In this way, the first storage device can execute the migration request to migrate the target data at the physical address 1 from the physical address 1 to the physical address 2, so as to realize the migration of the target data indicated by the source logical address to the destination logical address.

[0055] In this way, in the process of executing the data migration task of the east-west traffic, the task processing device 200 can convert the data migration task into a migration request including the physical address 1 (source physical address) and the physical address 2 (destination physical address) that can be recognized by the storage device, that is, convert the logical storage unit level semantic data migration task into a storage device level semantic task that can be understood by the storage device. This makes it possible for the storage device to execute the process of transmitting and processing data between different storage devices in the storage system 10, that is, to complete the data migration task. In this process, the CPU or heterogeneous processor in the storage system 10 does not need to intervene, which can effectively reduce the number of IO stacks and network stacks required to be executed by the storage system 10, thereby improving the execution efficiency of the data migration task.

[0056] Moreover, since no CPU intervention is required in the process of executing the data migration task, the CPU computing power can be released, so that the CPU can serve the upper-layer user IO based on more computing power, thereby effectively improving the performance of the storage system 10.

[0057] Still taking the data migration task as an example of the garbage collection task, the task processing apparatus 200 can access the metadata in the persistent layer, and according to the metadata, convert the source logical address in the garbage collection task to the physical address 1 of the valid data to be collected on the storage device P, and convert the destination logical address in the garbage collection task to the physical address 2 of the data written back after collection on the storage device Q, so as to convert the garbage collection task into a migration request including the physical address 1 and the physical address 2. Then, the task processing apparatus 200 can send the migration request to the storage device P. In this way, the storage device P can identify and execute the migration request, perform the garbage collection operation on the storage device P, and send the valid data reserved after the garbage collection to the storage device Q according to the indication information of the physical address 2 through the bus or communication network between the storage devices, and save the valid data by the storage device Q. In this process, for the data in the plurality of data blocks on the storage device P, the task processing apparatus 200 can only need to perform 1 IO stack (the IO stack is used to send the migration request, and can not include the data in the data block), and can transmit the data between the storage devices once, as shown in FIG. 8. At the same time, directly transmitting the data between the storage devices can also shorten the data transmission path. In this way, the execution efficiency of the garbage collection task is improved. Figure 2b

[0058] It is worth noting that the above Figure 1 ​The storage system 10 shown is only an exemplary illustration and is not intended to be limiting. For example, in actual application scenarios, the storage system 10 can include a larger number of storage servers. Alternatively, the storage server 101 can also include a heterogeneous processor, so that the task processing apparatus 200 can be deployed as software in the heterogeneous processor. Alternatively, the task processing apparatus 200 can be implemented by hardware deployed in the storage server. For example, the task processing apparatus 200 can be implemented by a processor, which can be any one of an application-specific integrated circuit (ASIC), a programmable logic device (PLD), a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), a system on chip (SoC), a software-defined infrastructure (SDI) chip, an artificial intelligence (AI) chip, a data processing unit (DPU), or any combination thereof. Moreover, the number of processors included in the task processing apparatus 200 can be any number, and the types of processors included can be one or more, which can be set according to the business requirements of actual applications, and the present application does not limit the number and types of processors.

[0059] For ease of understanding, the embodiments of the task processing method provided by the present application are described below with reference to the accompanying drawings.

[0060] Referring to Figure 3 , Figure 3 A flowchart of a task processing method provided by an embodiment of the present application is shown, which can be applied to Figure 1 the storage system 10 or other applicable storage systems. For ease of illustration, an exemplary illustration is provided in the embodiment by taking the storage system 10 shown as an example. Figure 1

[0061] Among them, Figure 3 The task processing method shown can specifically include:

[0062] ​S301: The task processing apparatus 200 receives a data migration task, which includes a source logical address and a destination logical address, and is used to instruct migration of target data indicated by the source logical address to the destination logical address, the source logical address belonging to a first logical storage unit, the destination logical address belonging to a second logical storage unit, the first logical storage unit corresponding to a first physical storage space in N storage devices, and the second logical storage unit corresponding to a second physical storage space in the N storage devices.

[0063] In this embodiment, the data migration task refers to a task that generates east-west traffic, i.e., a task of writing data on some storage devices to other storage devices. For example, the data migration task is used to instruct migration or backup of data on some storage devices to other storage devices, or to instruct processing of data on some storage devices and writing of generated new data to other storage devices, etc.

[0064] In the first implementation example, when the data migration task is divided according to the generation manner of the task, the data migration task can be a task generated by different business logic layers, which can include, for example, the cache subsystem 201, the index subsystem 202, or the persistence subsystem 203, etc.

[0065] The data migration task can be a task generated by the cache subsystem 201 to store data in the cache to multiple storage devices (backup storage).

[0066] Alternatively, the data migration task can be a garbage collection task or a performance tier down / up task generated by the index subsystem 202. The performance tier down / up task refers to migration of cold data to a storage device with lower read-write performance and migration of hot data to a storage device with higher read-write performance according to the coldness and hotness of data (e.g., determined according to the frequency of data access).

[0067] Alternatively, the data migration task can be a capacity expansion, capacity reduction (reduction), reconstruction task, or data ownership migration task generated by the persistence subsystem 203. The reconstruction task refers to reconstruction of another part of data stored based on an EC technology by using part of data stored based on the EC technology under an EC mechanism. For example, assuming that 6 storage devices are used to store data based on the EC mechanism in the storage system, when data on one of the storage devices is lost, the data on the storage device can be reconstructed by using data on the remaining storage devices to ensure data storage reliability.

[0068] In the second implementation example, when the data migration task is divided according to the specification of data movement, the data migration task can be a partial migration task, a whole migration task, or a cooperative migration task. The partial migration task refers to migrating part of data in the plurality of data on the storage device indicated by the first logical storage unit to other storage devices. For example, as shown in FIG. 8, part of the data blocks in a chunk (the first logical storage unit) on disk 0 are migrated to disk 1. The whole migration task refers to migrating all data in the plurality of data on the storage device indicated by the first logical storage unit to other storage devices. For example, as shown in FIG. 9, all data in a chunk on disk 0 are migrated to disk 1. The cooperative migration task refers to reconstructing data based on the EC technology using the data on the storage device indicated by the first logical storage unit, and writing the reconstructed new data to other storage devices. For example, as shown in FIG. 10, data on disks 0 to 4 is reconstructed based on the EC technology, and the reconstructed data is written to disk 5. Figure 4 Figure 4 Figure 4

[0069] Generally, the data migration task obtained by the task processing apparatus 200 can include a source logical address in the first logical storage unit and a destination logical address in the second logical storage unit, for indicating migration of the target data indicated by the source logical address to the destination logical address. When data is migrated between the source logical address and the destination logical address, a plurality of processing operations can be performed on the data. For example, the processing operation can be a data migration operation, indicating migration of the data indicated by the source logical address to the destination logical address, so that after the data migration is completed, the source logical address does not indicate the data (the original target data is deleted). For another example, the processing operation can be a data backup operation, indicating backup of the data indicated by the source logical address to the destination logical address, so that after the data migration is completed, the source logical address still indicates the target data. For another example, the processing operation can be a data reconstruction operation, indicating reconstruction of the data indicated by the source logical address and writing the obtained data to the destination logical address. For another example, the processing operation can be an EC (error correction code) calculation operation, indicating EC calculation of the data indicated by the source logical address to obtain a parity, and splitting and writing the data and the parity to a plurality of different positions indicated by the destination logical address. Based on this, the migration request can further include a processing operation, which can be a data migration operation, a data backup operation, a data reconstruction operation, or an EC calculation operation, or can be another type of operation, which is not limited herein.

[0070] ​​​In a possible implementation, the task processing apparatus 200 can provide different interfaces for data migration tasks of different service types, and each interface can indicate a type of processing operation for data. Then, after generating a data migration task, the cache subsystem 201, the index subsystem 202, or the persistent subsystem 203 can invoke a corresponding interface to provide the data migration task to the task processing apparatus 200 according to the service type to which the data migration task corresponds. In this way, the task processing apparatus 200 can determine the service type to which the data migration task corresponds according to the interface used to receive the data migration task, and further determine the processing operation required to be performed on the data when migrating the data between different storage devices according to the service type.

[0071] For example, when the data migration task is specifically a garbage collection task or a performance tier down / up task, the index subsystem 202 can invoke interface 1 to send the data migration task to the task processing apparatus 200. Correspondingly, the task processing apparatus 200 can determine, according to interface 1, that the service type to which the data migration task corresponds is a first type (e.g., a local move service type), and determine to migrate the target data indicated by the source logical address from one storage device to another storage device.

[0072] For another example, when the data migration task is specifically a persistent storage task, the cache subsystem 201 can invoke interface 2 to send the data migration task to the task processing apparatus 200. Correspondingly, the task processing apparatus 200 can determine, according to interface 2, that the service type to which the data migration task corresponds is a second type (e.g., a whole move service type), and determine to backup the target data indicated by the source logical address to a plurality of other storage devices.

[0073] For another example, when the data migration task is specifically a reconstruction task, the persistent subsystem 203 can invoke interface 3 to send the data migration task to the task processing apparatus 200. Correspondingly, the task processing apparatus 200 can determine, according to interface 3, that the service type to which the data migration task corresponds is a third type (e.g., a cooperative move service type), and determine to perform a data reconstruction operation on the target data indicated by the source logical address (the target data is dispersed on a plurality of storage devices) based on an EC technology.

[0074] The first logical storage unit and the second logical storage unit may, for example, be a Plog, a chunk, or a file, etc. The chunk refers to a logical storage unit that maps a piece of writable space of one or more storage devices to the storage system 10. When creating a logical storage unit, the storage system 10 abstracts a logical entity from a piece or pieces of physical storage space on the storage device, and assigns a unique identifier to the logical entity, such as a chunk number, etc. The mapping relationship between the logical storage unit and the physical storage space can be saved in the persistent layer. The processing operation performed may, for example, be a data reconstruction operation, or a data migration operation, etc., which is not limited in this regard. By way of example, the source logical address may, for example, be indicated by the identifier of the first logical storage unit, the offset of the target data in the first logical storage unit, and the length of the target data. Similarly, the destination logical address may, for example, be indicated by the identifier of the second logical storage unit, the offset, and the data length; or the address 2 may, for example, only include the identifier of the second logical storage unit, such as the user specifying that the data is saved to the second logical storage unit, but not specifying the specific location of the data in the second logical storage unit.

[0075] In this embodiment, the first logical storage unit (similar to the second logical storage unit) may, for example, refer to one logical storage unit, or may, for example, refer to a plurality of logical storage units, such as a plurality of chunks, etc., which is not limited in this regard.

[0076] S302: The task processing apparatus 200 converts the source logical address to a first physical address in the first physical storage space, and converts the destination logical address to a second physical address in the second physical storage space.

[0077] In this embodiment, the underlying storage device is generally difficult to directly recognize the data migration task including the source logical address and the destination logical address. Therefore, the task processing apparatus 200 may, for example, first convert the data migration task into a migration request including a physical address that can be recognized by the storage device, so as to subsequently perform the process of migrating the data between different storage devices by the storage device.

[0078] It can be understood that, since each logical storage unit in the storage system 10 generally maps different physical storage spaces on the storage device, the processing operation performed on the target data indicated by the source logical address is the processing operation performed on the data in the storage device indicated by the physical address mapped to the source logical address. Therefore, the task processing apparatus 200 may, for example, convert each logical address in the data migration task to the corresponding physical address.

[0079] In one possible implementation, the persistent subsystem 203 in the storage system 10 can store metadata, and the metadata can be used to describe the mapping relationship between the respective logical storage units and the physical storage space on the storage devices. In this way, the task processing apparatus 200 can obtain the metadata corresponding to the first logical storage unit and the second logical storage unit according to accessing the persistent subsystem 203, so as to convert the logical address in the data migration task according to the metadata.

[0080] Specifically, as shown in the figure, Figure 1 the task processing apparatus 200 can be configured with a unified data processing framework (UDPF), and after obtaining the data migration task, the UDFP can access the persistent subsystem 203 to obtain the first metadata corresponding to the first logical storage unit and the second metadata corresponding to the second logical storage unit. The first metadata includes the first address of the physical storage space in the storage device corresponding to the first logical storage unit, and the second metadata includes the first address of the physical storage space in the storage device corresponding to the second logical storage unit. The physical storage space mapped by the first logical storage unit and the physical storage space mapped by the second logical storage unit can be located in different storage devices. For example, the source logical address can include the identifier of the first logical storage unit, the offset of the target data in the first logical storage unit, and the data length, etc. Then, the task processing apparatus 200 can access the persistent subsystem 203 to obtain the first metadata corresponding to the first logical storage unit according to the identifier of the first logical storage unit. Similarly, the task processing apparatus 200 can obtain the second metadata corresponding to the second logical storage unit according to the identifier of the second logical storage unit included in the address 2.

[0081] Then, the UDFP can calculate the first physical address corresponding to the target data in the storage device according to the first metadata and the source logical address, that is, the first physical address mapped by the source logical address. For example, the UDFP can calculate the starting storage position of the target data stored in the storage device according to the first address of the physical storage space in the storage device corresponding to the first logical storage unit and the offset included in the source logical address. For example, the starting storage position can be identified by a logical block address (LBA). Then, the UDFP can determine the physical storage space of the target data in the storage device according to the data length included in the source logical address and the starting storage position. Then, the address used to indicate the physical storage space, that is, the first physical address of the target data in the storage device, can include the LBA and the data length, etc.

[0082] For example, asFigure 5 As shown, the first logical storage unit is specifically a chunk, the target data is part of the data in the chunk, and the source logical address can include a chunk identifier, an offset of 1 MB (megabyte), and a data length of 4 KB (kilobyte). The first address of the physical storage space in the storage device corresponding to the first logical storage unit is 100 MB. The task processing apparatus 200 calculates the starting position of the target data stored in the storage device as 101 MB according to the offset in the address 1 and the first address, and determines the first physical address of the target data in the storage device as 101 MB to (101 MB + 4 KB) according to the data length, as shown in Figure 5 .

[0083] Further, the task processing apparatus 200 can also calculate the second physical address in the storage device mapped by the destination logical address according to the second metadata and the destination logical address in the second logical storage unit. It should be noted that the target data has not been processed at this time, and thus the calculated second physical address can be the physical address of the data generated by the subsequent processing operation on the target data stored in the storage device.

[0084] S303: The task processing apparatus 200 generates a migration request carrying the first physical address and the second physical address.

[0085] In this embodiment, the task processing apparatus 200 can generate a migration request including the first physical address and the second physical address after converting the source logical address into the first physical address and converting the destination logical address into the second physical address, so as to convert the data migration task at the logical storage unit level into a task (i.e., the migration request) at the disk level recognized by the storage device, so that the storage device can execute the process of migrating the target data indicated by the source logical address to the destination logical address in the subsequent process.

[0086] In actual application, since the storage system 10 includes a plurality of storage devices, the first physical address and the second physical address determined by the task processing apparatus 200 can also be associated with the identifiers of the storage devices, respectively, so as to determine the storage device to which the first physical address belongs and the storage device to which the second physical address belongs by using the associated identifiers.

[0087] Further, the migration request can also carry a processing operation (such as an operation code included in the migration request), so as to instruct the storage device to process the target data by using the processing operation.

[0088] S304: The task processing apparatus 200 sends the migration request to the first storage device in the N storage devices.

[0089] S305: The first storage device executes the migration request to migrate the target data from the first physical address to the second physical address.

[0090] The first storage device can be one or more of the N storage devices. In the data migration scenario, the first storage device can be the storage device to which the physical storage space indicated by the first physical address belongs. In the EC reconstruction scenario, the first storage device can be the storage device to which the physical storage space indicated by the second physical address belongs.

[0091] In the process of executing the migration request, the first storage device can access the target data on the corresponding storage device according to the first physical address in the migration request, and perform the corresponding processing operation (in the migration request) on the target data, and then write the data obtained by performing the processing operation to the storage location indicated by the second physical address according to the second physical address in the migration request. In actual application scenarios, the first storage device can be configured with hardware for parsing tasks, and the first storage device can use the hardware to parse the first physical address, the processing operation, and the second physical address and other information from the migration request, and perform the processing operation on the target data indicated by the first physical address. Alternatively, the first storage device can be configured with a software program for parsing tasks, and the first storage device can implement parsing the first physical address, the processing operation, and the second physical address and other information from the migration request by running the software program, and performing the processing operation on the target data indicated by the first physical address.

[0092] When the physical storage space indicated by the first physical address and the physical storage space indicated by the second physical address are different physical storage spaces in the same storage device, the first storage device can internally perform data migration, avoiding data transmission across storage devices in the communication network, which can effectively improve the efficiency of performing the data migration task.

[0093] When the physical storage space indicated by the first physical address and the physical storage space indicated by the second physical address are physical storage spaces in different storage devices, the first storage device can transmit the data obtained by performing the processing operation on the target data from the first storage device to the storage device to which the physical storage space indicated by the second physical address belongs, or read the target data in other storage devices into the first storage device to perform the processing operation, and save the obtained data in the first storage device, which can effectively shorten the data transmission path when performing the data migration task (without the need to transmit data from the storage device to the CPU / heterogeneous processor and then back to the storage device), so as to improve the efficiency of performing the data migration task.

[0094] In actual application scenarios, the number of the second physical addresses included in the migration request can be one or multiple. When the number of the second physical addresses is one, the first storage device can write the data obtained by performing the processing operation to a physical storage space on one storage device indicated by the second physical address. When the number of the second physical addresses is multiple, the first storage device can write the data obtained by performing the processing operation to physical storage spaces in different storage devices or to different physical storage spaces in the same storage device. For example, the data migration task can be a local move task (such as a write-back cache task), and the indexing subsystem 202 can generate a local move task to indicate that the target data is written to multiple storage devices in the storage system 10. At this time, the migration request generated based on the local move task can include multiple second physical addresses of the storage devices used to back up the target data, so that the first storage device can transmit the target data to the corresponding multiple storage devices for saving according to the multiple second physical addresses when executing the migration request.

[0095] In addition, the number of the first physical addresses included in the migration request can be one or multiple.

[0096] 1. The number of the first physical addresses is one.

[0097] At this time, the data migration task can be, for example, a local move task or a whole move task. The first storage device can perform a data backup operation on the data indicated by the first physical address and save the obtained data to the storage location indicated by the second physical address.

[0098] 2. The number of the first physical addresses is multiple.

[0099] At this time, the data migration task can be, for example, a cooperative move task or a local move task.

[0100] For example, in an EC reconstruction scenario (the data migration task is a cooperative move task), the number of the first physical addresses can be multiple, and at this time, different first physical addresses can indicate physical storage spaces on different storage devices.

[0101] By way of example, assume that in the storage system 10, data is stored based on K+1 storage devices using the EC technology, K is an integer greater than 1. When data loss occurs on one of the storage devices, the persistent subsystem 203 can generate the cooperative migration task to instruct the use of data on the remaining K storage devices to recover the lost data. At this time, in the migration request generated based on the cooperative migration task, the first physical addresses of the remaining data when stored on the K storage devices respectively can be included. In this way, when the first storage device executes the migration request, the first storage device can acquire target data from the corresponding K storage devices respectively according to the K first physical addresses carried in the migration request, and reconstruct the data to be recovered (i.e., reconstruct the lost data) based on the target data acquired from the K storage devices, and further save the reconstructed data on the first storage device.

[0102] For another example, in the scenario of migrating data stored based on the EC technology (the data migration task can be a partial migration task), the number of first physical addresses can be multiple, at this time, different first physical addresses can indicate multiple different physical storage spaces on the same storage device, and the data in each physical storage space is the data in the storage device in the multiple stripes included by the first logical storage unit.

[0103] Further, when the processing operation in the migration request is a migration operation, since the migration request is converted from the data migration task based on the logical storage unit level semantics, when the first storage device executes the migration request, the first storage device still migrates the target data indicated by the first physical address according to the data processing logic in the logical storage unit dimension. When the number of first physical addresses is multiple, the first storage device usually migrates the data in the physical storage space indicated by each first physical address one by one, which can result in low migration efficiency of the data in different physical storage spaces, and thus result in low execution efficiency of the data migration task. Moreover, when storage is based on the EC technology, if the data is migrated according to the data processing logic in the logical storage unit dimension, the first storage device usually reads the data on each storage device participating in the EC storage first, and then splits the data and the check data to different storage devices for storage after EC calculation, which can result in large aggregation read traffic (i.e., data on other storage devices is aggregated to the first storage device).

[0104] To this end, as shown in FIG. 2, the persistent subsystem 203 can further include a migration request conversion module 205. The migration request conversion module 205 can be configured to convert the data migration task based on the logical storage unit level semantics into a migration request based on the physical storage space level semantics, and the migration request conversion module 205 can be configured to convert the data migration task based on the logical storage unit level semantics into a migration request based on the physical storage space level semantics. Figure 1As shown, the task processing apparatus 200 can further include a data-aware task orchestration (DATO) module, and the DATO module can generate a data view corresponding to the first storage device according to the storage location of the target data on the first storage device, the data view being used to describe the location distribution of the target data on the first storage device, and generate a migration request capable of providing task execution efficiency according to the data view corresponding to the first storage device. Thus, the DATO module can provide the migration request to the first storage device, so that the first storage device responds to the migration request to implement processing of the target data. In this way, the DATO module plans and controls the first storage device to migrate data processes indicated by different first physical addresses according to the data view, which can optimize the efficiency of the first storage device in migrating data in different physical storage spaces, that is, improve the execution efficiency of the data migration task. For example, assuming that the first logical storage unit is a chunk, the information of the logical address in the data migration task and the information of the physical address in the migration request are as follows Figure 6 As shown, the data view can be used to describe the storage location of the target data on the storage device. Then, the DATO module can generate a data view as shown in Figure 6 According to the information of the logical address in the data migration task and the converted physical address information, and when migrating data stored based on the EC technology, the first storage device migrates data belonging to the first logical storage unit on the first storage device according to the data view, which can also avoid generating a large amount of aggregation read traffic.

[0105] The following will be exemplarily described in combination with a specific scenario of data migration.

[0106] Scenario one, the data is backup stored in the storage system 10, that is, a copy of the same data can be stored on multiple storage devices. When the target data belonging to the first logical storage unit on the first storage device needs to be migrated to other multiple storage devices, the DATO module in the task processing apparatus 200 can determine data respectively indicated by multiple first physical addresses in the same storage device according to the data view corresponding to the first storage device (describing the location distribution of the target data on the first storage device), generate migration requests for the data respectively indicated by the multiple first physical addresses, and send the migration requests to the first storage device, so as to instruct the first storage device to batch backup the data respectively indicated by the multiple first physical addresses by using the migration requests.

[0107] For example, as shown in Figure 7As shown, assuming that the first storage device is the storage device 10, the target data belonging to the first logical storage unit includes data in a plurality of first physical addresses corresponding to chunks 8, 9 and 10 (the physical storage space is discontinuous), wherein the target data in chunks 8 and 10 needs to be backed up in storage devices 1, 3 and 5, and the target data in chunk 9 needs to be backed up in storage devices 2, 4 and 6 (for example, different data is backed up in different storage devices based on a load balancing strategy). Then, after the migration request is converted, the task processing apparatus 200 can generate a data view as shown in Figure 7 based on the location distribution of the target data on the storage device 10, and generate a migration request 1 (or can be referred to as a control command) and a migration request 2 as shown in Figure 7 based on the data view. The migration request 1 is used to instruct the storage device 10 to migrate 4 KB of data in chunk 9 to storage devices 2, 4 and 6. The migration request 2 is used to instruct the storage device 10 to batch backup 2 target data of 4 KB in chunk 8 and 8 KB of target data in chunk 10 to storage devices 1, 3 and 5, as shown in Figure 7 In this way, for the target data indicated by the three physical addresses, the task processing apparatus 200 can control the storage device 10 to batch backup the target data to reduce the number of IOs generated in the data migration process and improve the data migration efficiency.

[0108] In the second scenario, the target data in the first logical storage unit can include multiple stripes of data, and each stripe of data is stored in multiple storage devices, i.e., each storage device can store part of the data of the stripe. The data view corresponding to each storage device is the data view generated according to the location distribution of the target data on the storage device. Assuming that the data in the storage system 10 is stored based on the EC technology through (m+n) storage devices, m is an integer greater than 1, and n is a positive integer. When the target data needs to be migrated to other storage devices for storage, for each storage device participating in storing the target data, the task processing apparatus 200 (the DATO module therein) can determine m data (or can be referred to as user data / valid data) from the target data according to the data view corresponding to the first storage device (describing the location distribution of the data of the multiple stripes of the target data on the first storage device), generate a migration request for the m data, and send the migration request to the first storage device. In this way, the first storage device can generate n check data for the m data based on the EC technology according to the migration request, specifically according to the processing operation carried in the migration request. In actual application scenarios, the first storage device can be configured with hardware for EC calculation, and the first storage device can use the hardware to perform operations such as generating n check data for the m data by performing EC calculation. Alternatively, the first storage device can be configured with a software program for implementing the EC technology, so that the first storage device can implement operations such as generating n check data for the m data by performing EC calculation by running the software program. Then, the first storage device splits different data (including user data and check data) in the stripe to different storage devices for storage. In this way, the first storage device can avoid aggregating stripe data on multiple storage devices (storage devices before data migration) and then dispersing the stripe data to different storage devices (storage devices after data migration) during data migration, i.e., avoid aggregating read traffic, improve the efficiency of data migration, i.e., improve the efficiency of executing the data migration task, while also reducing the required resource overhead. For other data in the target data in addition to the m data, the first storage device can also perform data migration according to the migration request in a similar manner as described above; and when the number of the remaining data is less than m, m data can be constructed by padding 0, and then data migration can be performed in a similar manner as described above.

[0109] For example, as Figure 8As shown, assuming that in the storage system 10, storage devices 0, 1, 2, and 3 are used to store target data, and storage devices 4 and 5 are used to store the check data corresponding to the target data (i.e., a 4+2 EC storage ratio is adopted), the data view corresponding to the target data stored on storage device 0 is as follows: Figure 8 Then, for storage device 0, the task processing apparatus 200 can determine the data in the four storage areas from the target data on storage device 0 according to the data view, such as Figure 8 As shown, a migration request is generated for the data in the four storage areas, for example, Figure 8 The migration request shown is sent to storage device 0. Storage device 0 can generate two storage area size verification data for the data in the four storage areas based on EC technology according to the received migration request (the specifications of each storage area are the same, such as 4KB in size). In this way, the task processing device 200 can generate a new stripe based on the data in the four storage areas and the two storage area size verification data, and split the different parts of the data in the new stripe into the other six storage devices (such as storage device 6 to storage device 11) for storage, such as Figure 8 As shown. Each storage device can store the same amount of data. For each storage device that stores target data, the task processing device 200 can generate a corresponding migration request for the storage device in accordance with the above method, and by sending the migration request to the storage device, instruct the storage device to disperse the target data stored thereon to different storage devices for storage. In this process, since the storage device will generate verification data for the target data locally and save the verification data on other storage devices, after migrating the target data, the newly generated verification data can be used to ensure the accuracy and reliability of the migrated target data in the storage system 10. At the same time, in the process of migrating the target data, there is no need to execute the process of reading the target data from storage device 0 to storage device 3, thereby avoiding the aggregation of read traffic. In actual application scenarios, the read traffic that can be saved based on the above process can reach 50%. It should be noted that in this embodiment, the storage device 4 and storage device 5 are used as an example to store verification data separately. In actual application, a single storage device can not only store valid data / user data in part of the stripe, but also store verification data of other stripes. At this time, the storage device may also refer to the above method, and after piecing together the data of multiple stripes on the storage device into one stripe based on EC technology according to the data view, split the stripe onto multiple different storage devices for storage.

[0110] Further, for the storage devices 4 and 5 that save the check data corresponding to the target data, the task processing apparatus 200 can refer to the above manner to instruct the storage devices 4 and 5 to also transmit the initially generated check data corresponding to the target data to other storage devices for saving. Or, in the case where the accuracy and reliability of the target data have been ensured by using the newly generated check data, the task processing apparatus 200 can also not instruct the storage devices 4 and 5 to save the initially generated check data saved thereon to other storage devices, such as the task processing apparatus 200 can instruct the storage devices 4 and 5 to delete the check data saved thereon and release the storage space occupied by the check data, etc.

[0111] In the embodiment, the task processing apparatus 200 is taken as an example to perform the processing operation on the target data in the first logical storage unit, and in actual application scenarios, when there is a task in the storage system 10 that generates the east-west traffic for the data in other logical storage units, the task processing apparatus 200 can refer to the above manner to convert the task of the logical storage unit level semantics into the migration request of the storage device level semantics including the physical address, and send the migration request to the second storage device for execution, so as to improve the efficiency of executing the task in the storage system 10.

[0112] It is worth noting that other reasonable step combinations that can be thought of by those skilled in the art according to the above description also belong to the protection scope of the present application. Secondly, those skilled in the art should also be familiar that the embodiments described in the specification all belong to preferred embodiments, and the actions involved are not necessarily required by the present application.

[0113] The above is combined with Figures 1 to 8 The task processing method provided by the embodiment of the present application is introduced, and then the structure of the task processing apparatus and the computing device provided by the embodiment of the present application is introduced in combination with the drawings.

[0114] Referring to Figure 9 , a structural schematic diagram of a task processing apparatus is shown. Figure 9 The task processing apparatus 900 shown is applied to a storage system, and the storage system includes the task processing apparatus 900 and a plurality of storage devices for persistently storing data.

[0115] As Figure 9 shown, the task processing apparatus 900 includes:

[0116] The acquisition module 901 is configured to receive a data migration task, the data migration task comprising a source logical address and a destination logical address, the data migration task being used to migrate target data indicated by the source logical address to the destination logical address, the source logical address belonging to a first logical storage unit, the destination logical address belonging to a second logical storage unit, the first logical storage unit corresponding to a first physical storage space in a plurality of storage devices, and the second logical storage unit corresponding to a second physical storage space in the plurality of storage devices;

[0117] The conversion module 902 is configured to convert the source logical address into a first physical address in the first physical storage space, and convert the destination logical address into a second physical address in the second physical storage space.

[0118] The generation module 903 is configured to generate a migration request, the migration request carrying the first physical address and the second physical address.

[0119] The communication module 904 is configured to send the migration request to a first storage device in the plurality of storage devices, the migration request being used to instruct the first storage device to migrate the target data from the first physical address to the second physical address.

[0120] In a possible implementation, the data migration task is used to instruct to migrate the target data from the source logical address to the destination logical address.

[0121] Alternatively, the data migration task is used to instruct to backup the target data to the destination logical address.

[0122] Alternatively, the data migration task is used to instruct to reconstruct first data from the target data based on an EC (Error Correction Code) technology, and store the first data to the second physical address.

[0123] In a possible implementation, the generation module 903 is further configured to:

[0124] determine a service type corresponding to the data migration task;

[0125] When the service type is a first type, the migration request is used to instruct the first storage device to migrate the target data at the first physical address to the second physical address.

[0126] When the service type is a second type, the migration request is used to instruct the first storage device to backup the target data to the second physical address.

[0127] When the service type is a third type, the migration request is used to instruct the first storage device to reconstruct data based on the EC technology according to the target data at the first physical address.

[0128] In a possible implementation, the target data is located in the first storage device, and the generation module 903 is further configured to generate a data view corresponding to the first storage device according to a storage location of the target data in the first storage device, where the data view is used to describe a location distribution of the target data in the first storage device.

[0129] Therefore, when generating the migration request, the generation module 903 is specifically configured to generate the migration request according to the data view corresponding to the first storage device.

[0130] In a possible implementation, the data migration task is used to indicate migration of the target data from a source logical address to a destination logical address, and the migration request is used to indicate migration of the target data from a storage location indicated by the first physical address to a storage location indicated by the second physical address.

[0131] In a possible implementation, the target data is stored in the storage system based on an EC technology.

[0132] The migration request is specifically used to indicate:

[0133] m data is determined from the target data, where m is an integer greater than 1;

[0134] n check data is generated for the m data based on the EC technology, where n is a positive integer;

[0135] Different data in a strip composed of the m data and the n check data is split to different storage devices for storage.

[0136] In a possible implementation, the data migration task is used to indicate backup of the target data to the destination logical address.

[0137] Therefore, when generating the migration request, the generation module 903 is specifically configured to:

[0138] According to the data view corresponding to the first storage device, a plurality of second data in the target data that are migrated to the same physical address in the second physical storage space is determined.

[0139] For the plurality of second data, a migration request is generated, where the migration request is used to indicate batch backup of the plurality of second data.

[0140] In a possible implementation, the target data is stored in the storage system based on an EC technology, and the migration request is used to instruct the first storage device to reconstruct first data from the target data and store the first data in the second physical address in the first storage device.

[0141] In a possible implementation, the plurality of storage devices interact with each other based on an interconnection bus.

[0142] Alternatively, the plurality of storage devices perform data interaction based on a network card, an intelligent network card, or a DPU (Data Processing Unit).

[0143] As Figure 9 The task processing apparatus 900 shown in FIG. 9 corresponds to the task processing apparatus 200 shown in FIG. 2, and thus Figure 3 The task processing apparatus 200 shown in the embodiment of FIG. 2, and thus Figure 9 The specific implementation mode of the task processing apparatus 900 shown in FIG. 9 and the technical effects thereof are described above with reference to the specific implementation mode of the task processing apparatus 200 shown in FIG. 2 and the technical effects thereof, and thus are not described herein again. Figure 3

[0144] A hardware structure schematic diagram of a computing device 1000 is provided in the present application, which may, for example, implement the task processing apparatus 200 and the like shown in the embodiment of FIG. 2. Figure 10 Figure 3 As

[0145] As Figure 10 The computing device 1000 includes a processor 1001, a memory 1002, and a communication interface 1003. The processor 1001, the memory 1002, and the communication interface 1003 communicate through a bus 1004, and can also communicate through wireless transmission and other means. The memory 1002 is used to store instructions, and the processor 1001 is used to execute the instructions stored in the memory 1002. Further, the computing device 1000 can also include a memory unit 1005, and the memory unit 1005 can be connected to the processor 1001, the storage medium 1002, and the communication interface 1003 through the bus 1004. The memory 1002 stores program code, and the processor 1001 can execute the following operations by using the program code stored in the memory 1002.

[0146] Receive a data migration task, the data migration task including a source logical address and a destination logical address, the data migration task being used to migrate target data indicated by the source logical address to the destination logical address, the source logical address belonging to a first logical storage unit, the destination logical address belonging to a second logical storage unit, the first logical storage unit corresponding to a first physical storage space in a plurality of storage devices, the second logical storage unit corresponding to a second physical storage space in the plurality of storage devices, the plurality of storage devices being used to persistently store data;

[0147] Convert the source logical address into a first physical address in the first physical storage space, and convert the destination logical address into a second physical address in the second physical storage space;

[0148] Generate a migration request, the migration request carrying the first physical address and the second physical address; ​

[0149] sending the migration request to a first storage device of the plurality of storage devices, the migration request to instruct the first storage device to migrate the target data from the first physical address to the second physical address.

[0150] It should be appreciated that in this embodiment, the processor 1001 can be a CPU, and the processor 1001 can also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor, etc.

[0151] The memory 1002 can include a read-only memory and a random access memory, and provide instructions and data to the processor 1001. The memory 1002 can also include a non-volatile random access memory.

[0152] The memory 1002 can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically EPROM (EEPROM) or a flash memory. The volatile memory can be a random access memory (RAM) used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM) and direct memory bus random access memory (DRAM).

[0153] The communication interface 1003 is configured to communicate with other devices connected to the computing device 1000. The bus 1004 can include, in addition to the data bus, a power bus, a control bus, and a state signal bus, etc. However, for the sake of clarity, all the buses are marked as the bus 1004 in the figure.

[0154] It should be understood that the computing device 1000 according to the embodiments of the present application can correspond to the task processing apparatus 200 in the embodiments of the present application, and can correspond to the method performed by the task processing apparatus 200 in the embodiments of the present application. Figure 3 The above and other operations and / or functions implemented by the computing device 1000 are respectively for implementing the flow of the corresponding method in the embodiments of the present application, and for brevity, will not be repeated here. Figure 3 The above and other operations and / or functions implemented by the computing device 1000 are respectively for implementing the flow of the corresponding method in the embodiments of the present application, and for brevity, will not be repeated here.

[0155] The embodiments of the present application also provide a computer readable storage medium. The computer readable storage medium can be any available medium or data storage device that can be used to store data that can be accessed by a computing device, or a data center containing one or more available media. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid state disk), etc. The computer readable storage medium includes instructions that instruct the computing device to perform the above task processing method.

[0156] The embodiments of the present application also provide a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computing device, the flow or function according to the embodiments of the present application is generated in whole or in part.

[0157] The computer instructions can be stored in a computer readable storage medium, or transmitted from one computer readable storage medium to another computer readable storage medium, for example, the computer instructions can be transmitted from one website, computer or data center to another website, computer or data center through wired (for example, coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (for example, infrared, wireless, microwave, etc.) mode.

[0158] The computer program product can be a software installation package, and in the case of needing to use any method of the above task processing method, the computer program product can be downloaded and executed on the computing device.

[0159] The above embodiments can be implemented in whole or in part by software, hardware, firmware or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded or executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that contains one or more available media sets. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state drive.

[0160] The terms used in the above embodiments are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in the specification of this application and the appended claims, the singular expressions "a", "an", "said", "above", "the" and "this" are intended to also include expressions such as "one or more", unless the context clearly indicates otherwise. It should also be understood that in the embodiments of the present application, "one or more" refers to one, two or more; the character " / " generally indicates that the objects before and after are in an "or" relationship. In the embodiments of the present application. "Simultaneously" refers to the same time period, including the situation at the same moment.

[0161] References to "one embodiment" or "some embodiments" in this specification mean that a particular feature, structure, or characteristic described in conjunction with that embodiment is included in one or more embodiments of the present application. Thus, phrases such as "in one embodiment," "in some embodiments," "in other embodiments," and "in yet other embodiments" appearing in various places in this specification do not necessarily refer to the same embodiment, but rather mean "one or more but not all embodiments," unless otherwise specifically emphasized. The terms "including," "comprising," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.

[0162] The above merely provides the specific implementation of the present application, but the protection scope of the present application is not limited to this. Any person skilled in the art can easily think of various equivalent modifications or replacements within the technical range disclosed by the present application, and these modifications or replacements should be covered in the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A task processing method characterized by, The method is applied to a storage system including a plurality of storage devices for persistently storing data, and the method includes: receiving a data migration task including a source logical address and a destination logical address, the data migration task being used to migrate target data indicated by the source logical address to the destination logical address, the source logical address belonging to a first logical storage unit, the destination logical address belonging to a second logical storage unit, the first logical storage unit corresponding to a first physical storage space in the plurality of storage devices, and the second logical storage unit corresponding to a second physical storage space in the plurality of storage devices; converting the source logical address into a first physical address in the first physical storage space and converting the destination logical address into a second physical address in the second physical storage space; generating a migration request carrying the first physical address and the second physical address; sending the migration request to a first storage device in the plurality of storage devices, the migration request being used to instruct the first storage device to migrate the target data from the first physical address to the second physical address.

2. The method of claim 1, wherein: the data migration task is used to instruct to migrate the target data from the source logical address to the destination logical address; or, the data migration task is used to instruct to backup the target data to the destination logical address; or, the data migration task is used to instruct to reconstruct data based on an error correction code (EC) technology according to the target data.

3. The method of claim 2, wherein, The method further includes: determining a service type corresponding to the data migration task; wherein, when the service type is a first type, the migration request is used to instruct the first storage device to migrate target data of the first physical address to the second physical address; when the service type is a second type, the migration request is used to instruct the first storage device to backup the target data to the second physical address; when the service type is a third type, the migration request is used to instruct the first storage device to reconstruct first data based on an EC technology according to target data of the first physical address, and store the first data to the second physical address.

4. The method according to claim 2 or 3, characterized in that, The target data is located in the first storage device, and the method further includes: generating a data view corresponding to the first storage device according to a storage location of the target data on the first storage device, the data view being used to describe a location distribution of the target data on the first storage device; then, the generating of the migration request includes: generating the migration request according to the data view corresponding to the first storage device.

5. The method of claim 4, wherein, The target data is stored in the storage system based on the EC technology; then, the migration request is specifically used to instruct: determining m data from the target data, m being an integer greater than 1; generating n check data based on the EC technology for the m data, n being a positive integer; Different data in a stripe composed of the m data and the n check data is split to different storage devices for storage.

6. The method of claim 4, wherein, The data migration task is used to instruct to backup the target data to the destination logical address; The migration request is generated according to the data view corresponding to the first storage device, including: According to the data view corresponding to the first storage device, determine a plurality of second data in the target data migrated to the same physical address in the second physical storage space; For the plurality of second data, the migration request is generated, and the migration request is used to instruct to backup the plurality of second data in batches.

7. The method according to any one of claims 1 to 6, characterized in that, The plurality of storage devices interact with each other based on an interconnection bus. Alternatively, the plurality of storage devices interact with each other based on a network card, an intelligent network card, or a data processing unit (DPU).

8. A task processing apparatus characterized by comprising: The task processing apparatus is applied to a storage system, the storage system includes a plurality of storage devices, the plurality of storage devices are used to persistently store data, and the task processing apparatus includes: The acquisition module is used to receive a data migration task, the data migration task includes a source logical address and a destination logical address, the data migration task is used to migrate target data indicated by the source logical address to the destination logical address, the source logical address belongs to a first logical storage unit, the destination logical address belongs to a second logical storage unit, the first logical storage unit corresponds to a first physical storage space in the plurality of storage devices, and the second logical storage unit corresponds to a second physical storage space in the plurality of storage devices; The conversion module is used to convert the source logical address into a first physical address in the first physical storage space, and convert the destination logical address into a second physical address in the second physical storage space; The generation module is used to generate a migration request, and the migration request carries the first physical address and the second physical address; The communication module is used to send the migration request to a first storage device in the plurality of storage devices, and the migration request is used to instruct the first storage device to migrate the target data from the first physical address to the second physical address.

9. The apparatus of claim 8, wherein: The data migration task is used to instruct to migrate the target data from the source logical address to the destination logical address; Alternatively, the data migration task is used to instruct to backup the target data to the destination logical address; Alternatively, the data migration task is used to instruct to reconstruct data based on an error correction code (EC) technology according to the target data.

10. The apparatus of claim 9, wherein, The generation module is further used to: Determine a service type corresponding to the data migration task; When the service type is a first type, the migration request is used to instruct the first storage device to migrate the target data of the first physical address to the second physical address; When the service type is a second type, the migration request is used to instruct the first storage device to backup the target data to the second physical address; When the service type is the third type, the migration request is used to instruct the first storage device to reconstruct first data from target data of the first physical address based on an EC technology, and store the first data to the second physical address.

11. The apparatus of claim 9 or 10, wherein, The target data is located in the first storage device, and the generation module is further used to generate a data view corresponding to the first storage device according to a storage location of the target data on the first storage device, the data view being used to describe a location distribution of the target data on the first storage device. Then, the generation module is specifically used to generate the migration request according to the data view corresponding to the first storage device when generating the migration request.

12. The apparatus of claim 11, wherein, The target data is stored in the storage system based on the EC technology. The migration request is specifically used to instruct: to determine m data from the target data, m being an integer greater than 1; to generate n check data for the m data based on the EC technology, n being a positive integer; to split different data in a stripe composed of the m data and the n check data to different storage devices for storage.

13. The apparatus of claim 11, wherein, The data migration task is used to instruct to backup the target data to the destination logical address. Then, the generation module is specifically used to: determine, according to the data view corresponding to the first storage device, a plurality of second data in the target data that are migrated to the same physical address in the second physical storage space; and generate the migration request for the plurality of second data, the migration request being used to instruct to backup the plurality of second data in batches.

14. The apparatus of any one of claims 8 to 13, wherein, The plurality of storage devices interact with each other based on an interconnection bus. Alternatively, the plurality of storage devices interact with each other based on a network card, an intelligent network card, or a data processing unit (DPU).

15. A storage system, characterized by The storage system includes a task processing device and a plurality of storage devices, the plurality of storage devices being used to persistently store data, and the plurality of storage devices including a first storage device. The task processing device is used to execute the method according to any one of claims 1 to 7. The first storage device is further used to execute a migration request to migrate target data from a first physical address to a second physical address.

16. A computing device, comprising: comprising a processor, a memory; The processor is used to execute instructions stored in the memory, so that the computing device executes steps of the method according to any one of claims 1 to 7.

17. A computer-readable storage medium, characterized in that, comprising instructions that, when executed on a computing device, cause the computing device to execute steps of the method according to any one of claims 1 to 7.

18. A computer program product comprising instructions, characterized in that, when executed on at least one computing device, cause the at least one computing device to execute the method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Intelligent rapid planning system and method based on multi-agent cooperation

    CN121957823A