Host task processing method and device and medium

By executing subtasks in parallel in the host task processing device, the problem of low parallelism in subtask execution in traditional hardware solutions is solved, and efficient task processing and system performance improvement is achieved.

CN120179181APending Publication Date: 2025-06-20SHANDONG YUNHAI GUOCHUANG CLOUD COMPUTING EQUIP IND INNOVATION CENT CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510376681.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

In traditional hardware solutions, the processing steps of PRPEntry are executed serially, resulting in low parallelism of subtask execution and difficult to meet high performance requirements.

Method used

The target task is obtained through the task shard controller, and the target task is split into multiple subtasks according to the cache enabled state and disk read and write configuration. The subtasks are executed in parallel by using multiple control units in the task executor to realize parallel reading and writing data of the page address of the physical area.

Benefits of technology

It greatly improves task processing efficiency, comprehensively improves the overall performance of the system, optimizes resource utilization, enhances system stability and reliability, and reduces bus usage and data transmission overhead.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120179181A_ABST
    Figure CN120179181A_ABST
Patent Text Reader

Abstract

The invention discloses a host task processing method and device and a medium, and relates to the field of computers, the host task processing method and device are applied to the host task processing device, the host task processing device comprises a task fragmentation controller, a task executor, a state manager and an entry sending controller, and the method comprises the steps that a target task is obtained through the task fragmentation controller, splitting the target task into a plurality of sub-tasks; according to task information of subtasks, physical area page addresses corresponding to the subtasks are dispatched through a task executor, and the subtasks are distributed to a plurality of control units in the task executor, so that parallel read-write operation on data of the physical area page addresses is achieved; updating the task execution state through a state manager according to the execution result of each subtask; and aggregating the submission queue entries of the plurality of subtasks through the entry sending controller. In conclusion, the subtasks can be executed in parallel, the task processing efficiency is greatly improved, and the overall performance of the system is comprehensively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technologies, and particularly to a host task processing method, device, and medium. Background Art

[0002] With the gradual enhancement of the performance of Solid State Drives (SSDs), SSDs implemented based on the Peripheral Component Interconnect Express (PCIe) interface have gradually become the mainstream of development due to their high performance, low latency, and other characteristics.

[0003] In the data interaction between an SSD and a host, Physical Region Range Entry (PRP Entry) is an indispensable key factor. Traditional Non-Volatile Memory Host Controller Interface Specification (NVMe) SSD controllers mostly use software to obtain PRP Entries, and this method has a large delay.

[0004] To solve the problem of the large delay in obtaining PRP Entries by the software method, a hardware solution is proposed. However, the processing steps of PRP Entries in this hardware solution are executed serially. As Figure 1 shown, after reading the PRP Entry of the current subtask is completed, the write operation of the PRP Entry of this subtask can be sequentially executed, and then the PRP Entry of the next subtask can be read. At the same time, the update operation of the corresponding Submission Queue Entry (SQE) of the current subtask can be performed. After the SQE update is completed, the doorbell register can be updated. That is to say, the PRP operation of one subtask must be completely completed before the next subtask can start to execute. It can be seen that in the traditional hardware solution, the parallelism of subtask execution is very low, and it is difficult to meet the actual requirements in scenarios with high performance requirements. Summary of the Invention

[0005] In view of this, the purpose of the present invention is to provide a host task processing method, device, and medium, which can execute subtasks in parallel, greatly improve the task processing efficiency, and comprehensively improve the overall performance of the system. The specific solutions are as follows:

[0006] In a first aspect, the present application discloses a method for processing host tasks, which is applied to a host task processing device. The host task processing device includes a task sharding controller, a task executor, a status manager, and an entry sending controller. The method includes:

[0007] Obtain a target task through the task sharding controller, and split the target task into multiple subtasks according to the cache enable status and disk read / write configuration to obtain the task information of each subtask;

[0008] Through the task executor, according to the task information of the subtasks, schedule the physical area page addresses corresponding to the subtasks, and allocate the subtasks to multiple control units in the task executor, so that the multiple control units can execute the subtasks in parallel, realizing the parallel read / write operation of the data at the physical area page addresses, and obtaining the execution results of each subtask;

[0009] Update the task execution status through the status manager according to the execution results of each subtask;

[0010] Aggregate the submission queue entries of the multiple subtasks through the entry sending controller, and when the aggregation condition is met, write the aggregated submission queue entries into a preset accessible space on the disk.

[0011] Optionally, through the task executor, according to the task information of the subtasks, schedule the physical area page addresses corresponding to the subtasks, and allocate the subtasks to multiple control units in the task executor, so that the multiple control units can execute the subtasks in parallel, realizing the parallel read / write operation of the data at the physical area page addresses, including:

[0012] Obtain the first current remaining space and the first existing task execution situation fed back by each input control unit in the task executor through the input scheduling unit in the task executor, and according to the first current remaining space and the first existing task execution situation, send the task information of the subtasks to the appropriate input control unit;

[0013] Through the appropriate input control unit, schedule the physical area page address corresponding to the subtask according to the task information of the subtask;

[0014] Obtain the second current remaining space and the second existing task execution situation fed back by each output control unit in the task executor through the output scheduling unit in the task executor, and according to the second current remaining space and the second existing task execution situation, send the physical area page address corresponding to the subtask to the appropriate output control unit;

[0015] Execute the subtask through the appropriate output control unit to realize the parallel read / write operation of the data at the physical area page addresses.

[0016] Optionally, the host task processing method further includes:

[0017] Setting the number of parallel combinations according to the first feature information of the target task; wherein each parallel combination includes an input control unit and an output control unit;

[0018] Sending the number to the task executor so that the task executor selects corresponding input control units and output control units to work in parallel according to the number.

[0019] Optionally, the host task processing method further includes:

[0020] Through the cache control unit in the task executor, based on the second feature information of the subtask, saving the physical area page address corresponding to the subtask to the corresponding storage location, and recording the usage status and information index of each storage location.

[0021] Optionally, after meeting the aggregation condition, writing the aggregated submission queue entries to the preset accessible space on the disk, including:

[0022] Judging whether the number of aggregated submission queue entries reaches a preset threshold;

[0023] If the number of aggregated submission queue entries reaches the preset threshold, writing the aggregated submission queue entries to the preset accessible space on the disk through the target Advanced Extensible Interface operation.

[0024] Optionally, splitting the target task into multiple subtasks according to the cache enable status and disk read / write configuration to obtain the task information of each subtask, including:

[0025] If the target cache is enabled, splitting the target task into multiple subtasks according to the data distribution of the target cache and the maximum number of read / write requests that the disk can receive to obtain the task information of each subtask; wherein the task information includes the physical area page address, data size, read / write type, and cache hit flag.

[0026] Optionally, splitting the target task into multiple subtasks according to the cache enable status and disk read / write configuration to obtain the task information of each subtask, including:

[0027] If the target cache is not enabled, splitting the target task into multiple subtasks according to the maximum number of read / write requests that the disk can receive to obtain the task information of each subtask; the task information includes the physical area page address, data size, read / write type, and target retry times.

[0028] In a second aspect, the present application discloses a host task processing device, including:

[0029] A task sharding controller, which is used to obtain a target task and split the target task into multiple subtasks according to the cache enable status and disk read / write configuration, so as to obtain the task information of each subtask;

[0030] A task executor, which is used to schedule the physical area page address corresponding to the subtask according to the task information of the subtask, and allocate the subtask to multiple control units in the task executor, so that the multiple control units can execute the subtask in parallel, realize the parallel read / write operation of the data of the physical area page address, and obtain the execution results of each subtask;

[0031] A status manager, which is used to update the task execution status according to the execution results of each subtask;

[0032] An entry sending controller, which is used to aggregate the submission queue entries of multiple subtasks, and when the aggregation condition is met, write the aggregated submission queue entries into a preset accessible space on the disk.

[0033] Optionally, the task executor includes:

[0034] An input scheduling unit, which is used to obtain the first current remaining space and the first existing task execution situation fed back by each input control unit in the task executor, and issue the task information of the subtask to the adapted input control unit according to the first current remaining space and the first existing task execution situation;

[0035] The adapted input control unit is used to schedule the physical area page address corresponding to the subtask according to the task information of the subtask;

[0036] An output scheduling unit, which is used to obtain the second current remaining space and the second existing task execution situation fed back by each output control unit in the task executor, and send the physical area page address corresponding to the subtask to the adapted output control unit according to the second current remaining space and the second existing task execution situation;

[0037] The adapted output control unit is used to execute the subtask and realize the parallel read / write operation of the data of the physical area page address.

[0038] In a third aspect, the present application discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, the foregoing disclosed host task processing method is implemented.

[0039] As can be seen, the present application proposes a host task processing method, which is applied to a host task processing device. The device includes a task sharding controller, a task executor, a status manager, and an entry sending controller. The method includes: obtaining a target task through the task sharding controller, and splitting the target task into multiple subtasks according to the cache enabling status and disk read / write configuration to obtain the task information of each subtask; scheduling the physical area page addresses corresponding to the subtasks through the task executor according to the task information of the subtasks, and allocating the subtasks to multiple control units in the task executor so that the multiple control units can execute the subtasks in parallel to implement parallel read / write operations on the data of the physical area page addresses, and obtaining the execution results of each subtask; updating the task execution status through the status manager according to the execution results of each subtask; aggregating the submission queue entries of the multiple subtasks through the entry sending controller, and when the aggregation condition is met, writing the aggregated submission queue entries into a preset accessible space on the disk. It can be seen that the multiple control units in the task executor can execute the subtasks in parallel to achieve parallel read / write of the data of the physical area page addresses, greatly improving the task processing efficiency. Further, the task sharding controller splits the task according to the cache enabling status and disk read / write configuration, optimizing resource utilization and avoiding resource waste or overoccupation. At the same time, the status manager updates the task status according to the execution results of the subtasks, enhancing the system stability and reliability, and can handle abnormal situations in task execution in a timely manner. Finally, the entry sending controller aggregates the submission queue entries and then writes them into the preset accessible space on the disk, reducing the bus occupancy and data transmission overhead, improving the bus resource utilization rate, and comprehensively improving the overall performance of the system. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.

[0041] Figure 1 Schematic diagram of a traditional host task processing method;

[0042] Figure 2 Flowchart of a host task processing method disclosed in the present application;

[0043] Figure 3 Schematic diagram of an NVMe PRP concurrent processing hardware disclosed in the present application;

[0044] Figure 4 Schematic diagram of a PRP data workflow disclosed in the present application;

[0045] Figure 5 A timing diagram of PRP entry processing disclosed in the present application;

[0046] Figure 6 A schematic structural diagram of a host task processing device disclosed in the present application. Specific embodiments

[0047] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0048] To solve the problem of large delay in obtaining PRPEntry in software, a hardware solution is proposed. However, the processing steps of PRPEntry in this hardware solution are executed serially. After reading the PRPEntry of the current subtask, the write operation of the PRPEntry of this subtask can be executed in sequence, and then the PRPEntry of the next subtask can be read. At the same time, the update operation of the SQE corresponding to the current subtask can be performed. After the SQE update is completed, the doorbell register can be updated. That is to say, the PRP operation of one subtask must be completed before the next subtask can start to execute. In the traditional hardware solution, the parallelism of subtask execution is very low, and it is difficult to meet the actual requirements in scenarios with high performance requirements.

[0049] For this reason, the embodiments of the present application propose a host task processing solution that can execute subtasks in parallel, greatly improving the task processing efficiency and comprehensively improving the overall performance of the system.

[0050] The embodiments of the present application disclose a host task processing method, which is applied to a host task processing device. The host task processing device includes a task sharding controller, a task executor, a status manager, and an entry sending controller. See Figure 2 As shown, the method includes:

[0051] Step S11: Obtain a target task through the task sharding controller, and split the target task into multiple subtasks according to the cache enable status and disk read / write configuration to obtain the task information of each subtask.

[0052] Aiming at the performance deficiency problem existing in the current command acceleration processing solution of the Non-Volatile Memory Host Controller Interface Specification (NVMe), this application specifically proposes a performance optimization solution for Physical Region Range Entry (PRP) data processing. By parallelizing the PRP data processing, the data processing speed can be significantly improved. Moreover, aiming at the problem that the issuance of a single Submission Queue Entry (SQE) occupies the bus, resulting in low bus resource utilization and limited overall system performance, this application proposes a SQE aggregation solution, which can integrate multiple SQEs and then issue them, reducing the number of bus occupations, thereby improving the performance of the overall device system, effectively alleviating the bus pressure, and enhancing the efficiency of the system in data transmission and task processing. The following is a specific expansion:

[0053] First, obtain the target task through the task sharding controller, and split the target task into multiple subtasks according to the cache enabling status and disk read / write configuration to obtain the task information of each subtask.

[0054] On the one hand, if the target cache is enabled, the target task is split into multiple subtasks according to the data distribution of the target cache and the maximum number of read / write requests that the disk can receive, and the task information of each subtask is obtained; wherein, the task information includes the physical area page address, data size, read / write type, and cache hit flag. Suppose there is a data backup task in this embodiment, and a large number of files need to be backed up from the host storage to the disk array. At this time, the target cache is enabled, and the data in the target cache is relatively scattered, with some frequently accessed files concentrated in specific areas, and the maximum number of read / write requests that the disk can receive is 100. The task sharding controller will split the task according to this information. Exemplarily, there is a 10GB large file that needs to be backed up, and the file contains multiple data blocks of different types, some of which are frequently accessed and already cached in the cache. The task sharding controller will first analyze the data distribution of the target cache and find that 3GB of data is already in the cache and distributed in several consecutive areas. According to the maximum number of read / write requests that the disk can receive, the remaining 7GB of data is split into multiple subtasks according to certain rules. Suppose the average data size of each subtask is 500MB (not exceeding the processing capacity corresponding to the maximum number of read / write requests of the disk), so 14 subtasks are obtained. The task information of each subtask includes the physical area page address (pointing to the corresponding storage location on the disk), data size (500MB), read / write type (here it is a write operation), and cache hit flag (for the 3GB data part that is already in the cache, the cache hit flag of the corresponding subtask is "yes", and the rest is "no"). In this way, for the data that is already in the cache, it can be directly read from the cache in subsequent processing, greatly reducing the time to read from the disk and improving the data processing speed. Secondly, splitting the task reasonably according to the maximum number of read / write requests of the disk avoids sending too many requests to the disk at one time, which may cause the disk processing capacity to be overloaded, ensures that the disk can efficiently and stably process the data write operation, and improves the execution efficiency of the overall data backup task.

[0055] On the other hand, if the target cache is not enabled, the target task is split into multiple subtasks according to the maximum number of read / write requests that the disk can receive, and the task information of each subtask is obtained; the task information includes the physical area page address, data size, read / write type, and target retry count. Exemplarily, taking the data backup task as an example again, assume that the total amount of data to be backed up is 8 GB, and the maximum number of read / write requests that the disk can receive is 80. The task sharding controller will split the task according to the maximum number of read / write requests that the disk can receive. The 8 GB of data is split into multiple subtasks, and the data size of each subtask is 100 MB (8 GB ÷ 80 = 100 MB). In this way, 80 subtasks are obtained. The task information of each subtask includes the physical area page address, data size (100 MB), read / write type (write operation), and target retry count (assume it is set to 3 times, which is to cope with possible data write failures). In this way, since the cache is not enabled and cache acceleration for data reading cannot be utilized, by reasonably controlling the data size of the subtasks to match the maximum number of read / write requests of the disk, it is ensured that the data can be written to the disk orderly and efficiently. At the same time, setting the target retry count allows for automatic re-writing attempts when data writing fails, improving the reliability of data backup. Even when facing occasional instability of the disk, it can ensure that the data is fully backed up to the disk to the greatest extent.

[0056] Step S12: According to the task information of the subtask, the task executor schedules the physical area page address corresponding to the subtask and assigns the subtask to multiple control units in the task executor, so that the multiple control units can execute the subtask in parallel, realizing the parallel read / write operation of the data at the physical area page address, and obtaining the execution results of each subtask.

[0057] In this embodiment, through the input scheduling unit in the task executor, the first current remaining space and the first existing task execution situation feedback by each input control unit in the task executor are obtained. According to the first current remaining space and the first existing task execution situation, the task information of the subtask is sent to the appropriate input control unit; through the appropriate input control unit, the physical area page address corresponding to the subtask is scheduled according to the task information of the subtask; through the output scheduling unit in the task executor, the second current remaining space and the second existing task execution situation feedback by each output control unit in the task executor are obtained. According to the second current remaining space and the second existing task execution situation, the physical area page address corresponding to the subtask is sent to the appropriate output control unit; through the appropriate output control unit, the subtask is executed to realize the parallel read / write of the data at the physical area page address.

[0058] Suppose there are 3 input control units in the task executor, namely A, B, and C. The first current remaining space of input control unit A is 600 MB, and it is executing the task of reading 200 MB of data from the physical area page address P, which is expected to be completed in 15 seconds; the remaining space of input control unit B is 850 MB, and there is no task being executed; the remaining space of input control unit C is 300 MB, and it is executing the task of reading 250 MB of data with a progress of 70%, which is expected to be completed in 8 seconds. There is a new subtask that needs to read 450 MB of data with the address Q. Since the remaining space of input control unit B is sufficient and idle, the input scheduling unit sends the task information to B.

[0059] Suppose there are 3 output control units in the task executor, namely 1, 2, and 3. The second current remaining space of output control unit 1 is 500 MB, and it is executing the task of writing 350 MB of data to the physical area page address R with a progress of 90%, which is expected to be completed in 5 seconds; the remaining space of output control unit 2 is 750 MB, and there is no task being executed; the remaining space of output control unit 3 is 200 MB, and it is executing the task of writing 180 MB of data with a progress of 50%, which is expected to be completed in 10 seconds. There is a processed subtask that needs to write 400 MB of data to the address S. Based on the above principles, the output scheduling unit sends the task information to output control unit 2.

[0060] In this embodiment, the number of parallel combinations is set according to the first feature information of the target task; each parallel combination includes an input control unit and an output control unit; the number is sent to the task executor so that the task executor can select the corresponding input control unit and output control unit to work in parallel according to the number. The first feature information, as the basis for setting the number of parallel combinations, covers many key elements. It can include the size of the task data volume. There are significant differences in the parallel processing requirements between tasks with a large amount of data and tasks with a small amount of data. The task type is also an important feature. For example, for compute-intensive tasks and data transfer tasks, their focuses and resource requirements for parallel processing are different. The urgency of the task is also crucial. Urgent tasks require quick responses and have higher requirements for the speed and efficiency of parallel processing, while ordinary tasks can pay more attention to balance in resource utilization. In addition, the distribution characteristics of the data, such as whether the data is stored centrally in a continuous area or scattered in different storage locations, will also affect the setting of parallel combinations. As the task is executed, the first feature information may change. For example, the data volume may suddenly increase or the task priority may change during the task process. At this time, the number of parallel combinations can be dynamically adjusted to achieve dynamic changes in the parallel granularity. For example, in a data mining task, the initial estimated data volume is small, and fine-grained parallelism is used to quickly start the task and process the initial data. As the mining progresses, the data volume increases sharply. The system reduces the number of parallel combinations according to the changed first feature information to improve the resource utilization efficiency for processing large-scale data. This dynamic adjustment feature enables the system to flexibly adapt to task requirements, maintain high performance at different task stages, and effectively improve the system's adaptability to complex and changing tasks.

[0061] Furthermore, the task executor in this embodiment also includes a cache control unit. Through the cache control unit in the task executor, based on the second feature information of the subtask, the physical area page address corresponding to the subtask is saved to the corresponding storage location, and the usage status and information index of each storage location are recorded. It should be noted that the second feature information can include the data access frequency, that is, the number of times the subtask accesses a specific physical area page address within a period of time; the data correlation, which reflects the degree of tightness of the data correlation between different subtasks, etc. If the second feature information shows that a certain physical area page address is frequently accessed, the cache control unit will save it to a high-speed storage location. In this way, when the subsequent input control unit reads data, it does not need to wait for disk addressing and can directly obtain the data from the cache, greatly shortening the data reading time and improving the task processing speed. In addition, during the task execution, if a failure or data anomaly occurs in a certain storage location, the original physical area page address can be quickly located through the information index to re-obtain or repair the data, reducing the risk of system crashes caused by data problems and comprehensively improving the system performance.

[0062] Step S13: Update the task execution status according to the execution results of each subtask through the status manager.

[0063] In this embodiment, for the subtasks split from the same task, the status manager manages their execution status in a specific order and sequentially sends the information of each subtask to the entry sending controller.

[0064] Step S14: Aggregate the submission queue entries of multiple subtasks through the entry sending controller, and when the aggregation condition is met, write the aggregated submission queue entries into the preset accessible space on the disk.

[0065] Judge whether the number of aggregated submission queue entries reaches a preset threshold; if the number of aggregated submission queue entries reaches the preset threshold, write the aggregated submission queue entries into the preset accessible space on the disk through the target Advanced Extensible Interface operation. Assume that the preset threshold is 3 submission queue entries. During the data processing, the system continuously monitors the number of aggregated submission queue entries. For example, an entry generated by a subtask is aggregated, and then an entry generated by another subtask is also aggregated. At this time, the number is 2, which does not reach the threshold. When the entry generated by the third subtask is aggregated, the number of aggregated submission queue entries reaches 3, reaching the preset threshold. At this time, the system will write the aggregated submission queue entries containing the above three entries into the preset accessible space on the disk through the target Advanced Extensible Interface operation. It should be noted that the subtasks participating in the aggregation come from the same target task.

[0066] See Figure 3 as shown in Figure 3It is a schematic diagram of NVMe PRP concurrent processing hardware. This application is applicable to scenarios where cache is enabled or disabled, the command sizes between the host and the disk do not match, and tasks need to be split. Taking the case of supporting up to 4 parallel data input caches and data output caches as an example. This application proposes to use hardware concurrent acceleration to process the PRP data input and output cache process, and supports multiple PRP data inputs or outputs simultaneously. The overall hardware structure is divided into a task processing plane and an IO command processing plane. The task processing plane mainly parses, addresses, splits into subtasks (each subtask corresponds to an IO), and performs task response operations on the tasks received by the hardware device; the IO processing plane mainly processes information such as PRP, SQE, and doorbell of the subtasks. The detailed functions of each hardware processing unit are as follows: Task receiving unit: Receives and caches the original task information; Task parsing unit: Sequentially reads the complete task information, parses the key fields, and obtains information such as the task type and the number of processed data; Data address scheduling unit: Schedules the required PRP address information according to the address list pointer and quantity; Task sharding control unit: Splits the task based on whether the cache is enabled and the maximum number of IOs that the disk can receive configured by the user, and obtains the relevant configuration information of the subtasks; PRP data input scheduling unit (i.e., the aforementioned input scheduling unit): Reads the subtasks from the cache of the task sharding control unit, schedules the PRP data input, and issues information; PRP data input cache control unit (a type of the aforementioned cache control unit): Implements the PRP data input cache to improve the system efficiency; PRP data output scheduling unit (i.e., the aforementioned output scheduling unit): Issues the PRP data and subtask information according to the data reading order, the status of the lower-level module, and the priority weight; PRP data output cache control unit (another type of the aforementioned cache control unit): Implements the PRP data output cache to improve the system efficiency; Status manager: This unit reorganizes and outputs the PRP update status and command information of the completed subtasks according to the upper-level instruction rules; Entry sending controller: Writes one or more SQEs to the disk accessible space when the conditions are met according to the SQE aggregation configuration; Doorbell register sending controller: Notifies the disk that a new SQE has been written; Command response processing unit: Receives the disk command response information and manages the execution status of a single command; Task response processing unit: Manages the completion status of multiple subtasks corresponding to a single task, and feeds back the total task execution status to the system task scheduling engine.

[0067] See Figure 4 As shown, taking 3 tasks as an example, Task 1 is split into 5 subtasks, denoted as ; Task 2 is split into 4 subtasks, denoted as ; Task 3 is split into 3 subtasks, denoted as The following is the description of PRP data processing: The outstanding ability of AXI read and write operations is 4, which enables 4 input control components and 4 output control components to issue AXI read and write operations simultaneously, and only one AXI (Advanced eXtensible Interface) response time is required to complete the corresponding read and write. The SQE aggregation count is set to 3, that is, once every three SQEs of subtasks are prepared, an AXI operation is initiated. And the SQEs of subtasks between different tasks are not aggregated. For example, once the SQEs of the 4th and 5th subtasks of Task 1 are prepared, an AXI operation of the SQE will be initiated.

[0068] Taking 3 tasks as an example for analysis Figure 1 and Figure 5 . Ignoring the influence of factors such as the arbitration handshake time of AXI read and write commands, the read and write operation length, and the response gap, assuming that the response time of the AXI read and write command response is t for both, the total time consumed for reading and writing 1 subtask of PRP is 2t, and the time consumed for writing 1 subtask of SQE to completion is t. In Figure 1 , the longest time consumed for the PRP read and write in this level of pipeline is 10t; while in Figure 5 , the longest time consumed for this level of pipeline is only 3t. After calculation, the optimization amplitude of this application compared with the general design reaches (10t - 3t) / 10t = 70%. For multiple tasks (set as a), the maximum number of subtask splits is bmax. In the PRP read and write in this level of pipeline, Figure 1 the longest time consumed is 2bmaxt, Figure 5 the longest time consumed is . When bmax > 1, Figure 1 the single-level pipeline time consumption is always greater than Figure 5 . Thus, it can be seen that the proposed PRP read and write parallel scheme in this application significantly improves the performance of the pipeline level.

[0069] This embodiment can also design an intelligent fault detection and recovery mechanism. Using the interaction information between each processing unit and the task execution status data, a fault monitoring model is constructed. For example, by analyzing the data transmission situation between the PRP data input scheduling unit and the input control module, and the cooperation status of the PRP data output scheduling unit and the output control component, it can monitor in real time whether faults such as data transmission errors and abnormal task processing occur in the system. Once a fault is detected, the system automatically starts the recovery program, and according to the fault type and location, quickly switches to the standby processing path or performs data repair operations. For example, when a certain input control module fails, the system can temporarily allocate its tasks to other idle input control modules to ensure the continuity of task processing, greatly improving the stability and reliability of the system.

[0070] As can be seen, the present application proposes a host task processing method, which is applied to a host task processing device. The device includes a task sharding controller, a task executor, a status manager, and an entry sending controller. The method includes: obtaining a target task through the task sharding controller, and splitting the target task into multiple subtasks according to the cache enabling status and disk read / write configuration to obtain the task information of each subtask; scheduling the physical area page address corresponding to the subtask according to the task information of the subtask through the task executor, and allocating the subtask to multiple control units in the task executor, so that the multiple control units can execute the subtasks in parallel to implement parallel read / write operations on the data of the physical area page address, and obtaining the execution results of each subtask; updating the task execution status according to the execution results of each subtask through the status manager; aggregating the submission queue entries of the multiple subtasks through the entry sending controller, and writing the aggregated submission queue entries into a preset accessible space on the disk when the aggregation condition is met. As can be seen, the multiple control units in the task executor can execute the subtasks in parallel to achieve parallel read / write of the data of the physical area page address, greatly improving the task processing efficiency. Further, the task sharding controller splits the task according to the cache enabling status and disk read / write configuration, optimizing resource utilization and avoiding resource waste or overoccupation. At the same time, the status manager updates the task status according to the execution results of the subtasks, enhancing the system stability and reliability and enabling timely handling of abnormal situations during task execution. Finally, the entry sending controller aggregates the submission queue entries and then writes them into the preset accessible space on the disk, reducing bus occupancy and data transmission overhead, improving the bus resource utilization rate, and comprehensively enhancing the overall performance of the system.

[0071] Correspondingly, the embodiment of the present application also discloses a host task processing device. Refer to Figure 6 As shown, the device includes:

[0072] A task sharding controller 11, configured to obtain a target task, and split the target task into multiple subtasks according to the cache enabling status and disk read / write configuration to obtain the task information of each subtask;

[0073] A task executor 12, configured to schedule the physical area page address corresponding to the subtask according to the task information of the subtask, and allocate the subtask to multiple control units in the task executor, so that the multiple control units can execute the subtasks in parallel to implement parallel read / write operations on the data of the physical area page address, and obtain the execution results of each subtask;

[0074] A status manager 13, configured to update the task execution status according to the execution results of each subtask;

[0075] An entry sending controller 14, which is configured to aggregate submission queue entries of multiple subtasks, and write the aggregated submission queue entries into a preset accessible space on a disk after the aggregation condition is met.

[0076] Among them, for the more specific working processes of the above-mentioned various modules, reference can be made to the corresponding content disclosed in the foregoing embodiments, and details will not be elaborated herein.

[0077] It can be seen that the present application proposes a host task processing method, which is applied to a host task processing device. The device includes a task sharding controller, a task executor, a status manager, and an entry sending controller. The method includes: obtaining a target task through the task sharding controller, and splitting the target task into multiple subtasks according to the cache enabling status and disk read / write configuration to obtain the task information of each subtask; scheduling the physical area page addresses corresponding to the subtasks through the task executor according to the task information of the subtasks, and allocating the subtasks to multiple control units in the task executor, so that the multiple control units can execute the subtasks in parallel to implement parallel read / write operations on the data of the physical area page addresses, and obtaining the execution results of each subtask; updating the task execution status through the status manager according to the execution results of each subtask; aggregating the submission queue entries of multiple subtasks through the entry sending controller, and writing the aggregated submission queue entries into a preset accessible space on a disk after the aggregation condition is met. It can be seen that the multiple control units in the task executor can execute the subtasks in parallel to implement parallel read / write of the data of the physical area page addresses, greatly improving the task processing efficiency. Further, the task sharding controller splits the task according to the cache enabling status and disk read / write configuration, optimizing resource utilization and avoiding resource waste or overoccupation. At the same time, the status manager updates the task status according to the execution results of the subtasks, enhancing the system stability and reliability, and can handle abnormal situations in task execution in a timely manner. Finally, the entry sending controller aggregates the submission queue entries and then writes them into the preset accessible space on the disk, reducing the bus occupancy and data transmission overhead, improving the bus resource utilization rate, and comprehensively improving the overall performance of the system.

[0078] Further, an embodiment of the present application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, the foregoing disclosed host task processing method is implemented.

[0079] For the specific steps of this method, reference can be made to the corresponding content disclosed in the foregoing embodiments, and details will not be elaborated herein.

[0080] In the present application, the various embodiments are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the embodiments, reference can be made to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple. For the relevant parts, reference can be made to the description in the method section.

[0081] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Skilled professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0082] The steps of the methods or algorithms described in combination with the embodiments disclosed herein can be directly implemented by hardware, software modules executed by a processor, or a combination of the two. The software modules can be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium well-known in the technical field.

[0083] Finally, it should also be noted that in this document, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the existence of additional identical elements in the process, method, article or device comprising the element.

[0084] The above has introduced in detail a host task processing method, device, and storage medium provided by the present application. Specific examples are used in this article to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.

Claims

1. A host task processing method, characterized in that: Applied to a host task processing device, the host task processing device includes a task slice controller, a task executor, a state manager and an entry sending controller, the method includes: Obtaining a target task through the task slicing controller, and splitting the target task into multiple subtasks according to the cache enabling state and the disk read / write configuration, and obtaining task information of each subtask; The task executor schedules the physical area page address corresponding to the subtask according to the task information of the subtask, and allocates the subtask to multiple control units in the task executor, so that the multiple control units execute the subtask in parallel, realize parallel reading and writing operations on the data of the physical area page address, and obtain the execution results of each subtask; Updating the task execution status according to the execution results of each subtask by the status manager; The submission queue entries of the plurality of subtasks are aggregated through the entry sending controller, and when an aggregation condition is met, the aggregated submission queue entries are written into a preset accessible space of a disk.

2. The host task processing method according to claim 1, characterized in that: The task executor schedules the physical area page address corresponding to the subtask according to the task information of the subtask, and allocates the subtask to multiple control units in the task executor so that the multiple control units execute the subtask in parallel, thereby realizing parallel reading and writing operations on the data of the physical area page address, including: Obtaining, through the input scheduling unit in the task executor, a first current remaining space and a first existing task execution status fed back by each input control unit in the task executor, and sending the task information of the subtask to the corresponding input control unit according to the first current remaining space and the first existing task execution status; By means of the adapted input control unit, the physical area page address corresponding to the subtask is scheduled according to the task information of the subtask; The output scheduling unit in the task executor obtains the second current remaining space and the second existing task execution status fed back by each output control unit in the task executor, and sends the physical area page address corresponding to the subtask to the corresponding output control unit according to the second current remaining space and the second existing task execution status; The subtask is executed by the corresponding output control unit to realize parallel reading and writing operations on the data of the page address of the physical area.

3. The host task processing method according to claim 2, characterized in that: Also includes: The number of parallel combinations is set according to the first characteristic information of the target task; wherein each of the parallel combinations includes one input control unit and one output control unit; The number is sent to the task executor so that the task executor selects the corresponding input control unit and the output control unit to work in parallel according to the number.

4. The host task processing method according to claim 2, characterized in that: Also includes: Through the cache control unit in the task executor, based on the second characteristic information of the subtask, the physical area page address corresponding to the subtask is saved to the corresponding storage location, and the usage status and information index of each storage location are recorded.

5. The host task processing method according to claim 1, characterized in that: When the aggregation condition is met, the aggregated submission queue entries are written to a preset accessible space on the disk, including: Determine whether the number of aggregated submission queue entries reaches a preset threshold; If the number of aggregated submission queue entries reaches the preset threshold, the aggregated submission queue entries are written to a preset accessible space of the disk through a target advanced extensible interface operation.

6. The host task processing method according to any one of claims 1 to 5, characterized in that: The target task is split into a plurality of subtasks according to the cache enabling state and the disk read / write configuration, and task information of each subtask is obtained, including: If the target cache is turned on, the target task is split into multiple subtasks according to the data distribution of the target cache and the maximum number of read and write requests that the disk can receive, and the task information of each subtask is obtained; wherein the task information includes the physical area page address, data size, read and write type and cache hit flag.

7. The host task processing method according to any one of claims 1 to 5, characterized in that: The target task is split into a plurality of subtasks according to the cache enabling state and the disk read / write configuration, and task information of each subtask is obtained, including: If the target cache is not enabled, the target task is split into multiple subtasks according to the maximum number of read and write requests that the disk can receive, and the task information of each subtask is obtained; the task information includes the physical area page address, data size, read and write type and target retry count.

8. A host task processing device, characterized in that: include: A task slicing controller is used to obtain a target task, and split the target task into multiple subtasks according to the cache enablement state and the disk read / write configuration, and obtain task information of each subtask; A task executor, configured to schedule a physical area page address corresponding to the subtask according to the task information of the subtask, and assign the subtask to a plurality of control units in the task executor, so that the plurality of control units execute the subtask in parallel, thereby realizing parallel reading and writing operations on the data of the physical area page address and obtaining execution results of each subtask; A status manager, used to update the task execution status according to the execution results of each subtask; The entry sending controller is used to aggregate the submission queue entries of the plurality of subtasks, and when an aggregation condition is met, write the aggregated submission queue entries into a preset accessible space of the disk.

9. The host task processing device according to claim 8, characterized in that: The task executor comprises: An input scheduling unit, used for obtaining the first current remaining space and the execution status of the first existing task fed back by each input control unit in the task executor, and sending the task information of the subtask to the corresponding input control unit according to the first current remaining space and the execution status of the first existing task; The adapted input control unit is used to schedule the physical area page address corresponding to the subtask according to the task information of the subtask; an output scheduling unit, configured to obtain the second current remaining space and the second existing task execution status fed back by each output control unit in the task executor, and send the physical area page address corresponding to the subtask to the corresponding output control unit according to the second current remaining space and the second existing task execution status; The corresponding output control unit is used to execute the subtask to realize parallel reading and writing operations on the data of the page address of the physical area.

10. A computer-readable storage medium, characterized in that: Used to store computer programs; wherein, when the computer program is executed by a processor, the host task processing method as described in any one of claims 1 to 7 is implemented.