Task distribution method and device, equipment and medium
By monitoring the parameter values of the hardware acceleration unit and determining the target hardware acceleration unit according to different distribution modes, the problem of insufficient resource utilization caused by the binding of the hardware acceleration unit and IO tasks is solved, and more efficient and balanced task distribution and resource utilization are achieved.
Patent Information
- Application Number
- CN202510031089.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-08
- Publication Date
- 2025-05-09
AI Technical Summary
In the prior art, there is a binding relationship between the hardware acceleration unit and the IO task, resulting in insufficient resource utilization and the inability to effectively utilize the resources of all hardware acceleration units. Especially when the IO tasks issued by the host are only for a few functions, the load load of the hardware acceleration unit is unbalanced.
By monitoring the current parameter values of the hardware acceleration unit, including the number of pending tasks and the delay required to complete the task, the target hardware acceleration unit is determined according to different distribution modes (such as monitoring based on weights, task count, weighting and task number mixing, weighting and task count and delay required to complete the task), and the task is assigned to the most suitable unit.
It realizes more efficient and balanced task distribution, maximizes the use of resources of hardware acceleration units, improves the overall performance and flexibility of the system to handle IO tasks, reduces task congestion and delay, improves work efficiency and avoids resource waste.
Smart Images

Figure CN119960982A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of control system task dispatching, and in particular to a task dispatching method, device, equipment and medium. Background Art
[0002] In the era of big data, data growth is explosive. Although solid-state drives (SSDs) are increasingly occupying the personal consumer market and small data center market due to factors such as increased performance, lower demand and lower prices, more and more companies are investing in the research and development of storage management chips (RAID cards).
[0003] In order to meet higher performance requirements, many manufacturers will design corresponding hardware acceleration units in the chip when designing RAID cards. This hardware acceleration unit integrates the algorithm acceleration engine, input and output (IO) collaborative processing engine, disk management engine and other engines to process the IO requests sent from the host to the disk side.
[0004] In order to increase the performance of IO processing, the number of acceleration units is usually increased in design. However, when the number is increased, a problem will arise: how to allocate the issued IO queues. Common processing methods are:
[0005] Each hardware acceleration unit is directly connected to a virtualized physical machine, that is, the number of acceleration units is the same as the number of virtualized physical machines and is independent of each other. A single acceleration unit is only responsible for processing IO tasks assigned to the specific functions of this unit; or, one virtualized physical machine corresponds to several hardware acceleration units, and the IO tasks issued by the host are assigned to several designated hardware acceleration units for processing through RAID stripes.
[0006] However, the existing methods all have a binding relationship between the hardware acceleration unit and the IO tasks that need to be processed, which has poor flexibility. If the IO tasks issued by the host at the same time are only for certain functions, several of the hardware acceleration units will be highly loaded, while other hardware acceleration units will have no tasks to process, and the resources of all hardware acceleration units cannot be utilized.
[0007] Therefore, how to reasonably distribute IO tasks to hardware acceleration units, how to maximize the use of system on-chip resources, and how to design a control system that can uniformly receive IOs distributed by the host to all PFs and distribute them to various hardware units in a reasonable manner have become bottlenecks in the performance and stability of the entire system. Summary of the invention
[0008] Based on the above technical problems, the embodiments of the present application provide a task dispatching method, apparatus, device and medium, aiming at how to efficiently and reasonably dispatch tasks so that the resources of each hardware acceleration unit can be fully utilized.
[0009] A first aspect of an embodiment of the present application provides a task dispatching method, which is applied to a task dispatching system, wherein the task dispatching system is communicatively connected with a plurality of hardware acceleration units, and the method comprises:
[0010] Monitoring a current parameter value of each hardware acceleration unit in the plurality of hardware acceleration units, wherein the current parameter value includes at least one of the following: the number of tasks to be processed by each hardware acceleration unit within a unit time, and the delay required for each hardware acceleration unit to complete a task within a unit time;
[0011] When a task to be dispatched is detected, a target hardware acceleration unit is determined according to a target dispatch mode based on the current parameter value of each hardware acceleration unit;
[0012] Allocate the task to be dispatched to the target hardware acceleration unit.
[0013] Optionally, from the register operation sequence, the current parameter value further includes: a weight of each hardware acceleration unit, the target dispatch mode is a weight-based dispatch mode, the weight of the hardware acceleration unit is used to express the ability to process tasks, the weight includes a queue weight and a current weight, the queue weight is used to express the inherent hardware performance of the hardware acceleration unit, and the current weight is used to express the status of tasks currently being processed and to be processed by the hardware acceleration unit; when a task to be dispatched is detected, based on the current parameter value of each hardware acceleration unit, the target hardware acceleration unit is determined according to the target dispatch mode, and the method further includes:
[0014] When the queue weights are the same, the hardware acceleration units are sequentially determined as target hardware acceleration units;
[0015] When the queue weights are different, the target weight of each queue is calculated according to the sum of the queue weight and the current weight, and the hardware acceleration unit with the largest target weight is determined as the target hardware acceleration unit.
[0016] Optionally, the current parameter value includes: the number of tasks to be processed by each hardware acceleration unit in a unit time, the target dispatch mode is a task number monitoring mode, based on the current parameter value of each hardware acceleration unit, according to the target dispatch mode, determining the target hardware acceleration unit, the method includes:
[0017] The hardware acceleration unit with the least number of tasks to be processed is determined as the target hardware acceleration unit.
[0018] Optionally, the current parameter value includes: a weight of each hardware acceleration unit, and the number of tasks to be processed by each hardware acceleration unit in a unit time, wherein the weight of the hardware acceleration unit is used to express the ability to process tasks, and the weight includes a queue weight and a current weight, the queue weight is used to express the inherent hardware performance of the hardware acceleration unit, and the current weight is used to express the tasks currently being processed and to be processed by the hardware acceleration unit; the target dispatch mode is a weighted and task number mixed monitoring mode; based on the current parameter value of each hardware acceleration unit, the target hardware acceleration unit is determined according to the target dispatch mode, and the method further includes:
[0019] When the queue weights are the same, the hardware acceleration units are sequentially determined as target hardware acceleration units;
[0020] When the queue weights are different, the target weight of each queue is calculated according to the sum of the queue weight and the current weight, and the hardware acceleration unit with the largest target weight is determined as the first target hardware acceleration unit;
[0021] The first target hardware acceleration unit with the least number of tasks to be processed is determined as the target hardware acceleration unit.
[0022] Optionally, the current parameter value includes: a weight of each hardware acceleration unit, and the number of tasks to be processed by each hardware acceleration unit in a unit time, and the delay required for each hardware acceleration unit to complete a task in a unit time, wherein the weight of the hardware acceleration unit is used to express the ability to process tasks, and the weight includes a queue weight and a current weight, the queue weight is used to express the inherent hardware performance of the hardware acceleration unit, and the current weight is used to express the tasks currently being processed and to be processed by the hardware acceleration unit; the target dispatch mode is a mixed monitoring mode of weighted and task number and delay required to complete the task; based on the current parameter value of each hardware acceleration unit, the target hardware acceleration unit is determined according to the target dispatch mode, and the method further includes:
[0023] When the queue weights are the same, the hardware acceleration units are sequentially determined as target hardware acceleration units;
[0024] When the queue weights are different, the target weight of each queue is calculated according to the sum of the queue weight and the current weight, and the hardware acceleration unit with the largest target weight is determined as the first target hardware acceleration unit;
[0025] Determine the first target hardware acceleration unit with the least number of tasks to be processed as the second target hardware acceleration unit;
[0026] The second target hardware acceleration unit with the lowest delay required to complete the task is determined as the target hardware acceleration unit.
[0027] Optionally, a health monitoring threshold is configured for the hardware acceleration unit, where the health monitoring threshold is used to indicate whether the hardware acceleration unit can complete a task normally. The method includes:
[0028] The task dispatching system configures the health monitoring threshold for each hardware acceleration unit and detects the health monitoring threshold of each hardware acceleration unit during operation;
[0029] When it is detected that the health monitoring threshold of the hardware acceleration unit is lower than the target threshold, the hardware acceleration unit is determined to be a sub-healthy hardware acceleration unit, and a detection instruction is issued to the sub-healthy hardware acceleration unit, wherein the detection instruction is used to determine whether the sub-healthy hardware acceleration unit has recovered to a healthy state;
[0030] When the health monitoring threshold of the sub-healthy hardware acceleration unit is higher than the target threshold, it is determined that the sub-healthy hardware acceleration unit is restored to a healthy hardware acceleration unit, and the healthy hardware acceleration unit is initialized.
[0031] Optionally, the method further comprises:
[0032] According to the number of hardware acceleration units and hardware performance, configure the number of queues dispatched to the hardware acceleration units and the number of tasks contained in each queue;
[0033] The queues are bound to the hardware acceleration units one by one, and each hardware acceleration unit executes the tasks in the corresponding queue.
[0034] A second aspect of an embodiment of the present application provides a task dispatching device, the device comprising:
[0035] A monitoring module, configured to monitor a current parameter value of each hardware acceleration unit in the plurality of hardware acceleration units, wherein the current parameter value includes at least one of the following: the number of tasks to be processed by each hardware acceleration unit in a unit time, and the delay required for each hardware acceleration unit to complete a task in a unit time;
[0036] A target acceleration unit determination module is used to determine the target hardware acceleration unit according to the target dispatch mode based on the current parameter value of each hardware acceleration unit when a task to be dispatched is detected;
[0037] The dispatching module is used to assign the task to be dispatched to the target hardware acceleration unit.
[0038] Optionally, the target acceleration unit determination module further includes:
[0039] A first weight determination submodule, configured to determine the hardware acceleration units as target hardware acceleration units in sequence when the queue weights are the same;
[0040] The second weight determination submodule is used to calculate the target weight of each queue according to the sum of the queue weight and the current weight when the queue weights are different, and determine the hardware acceleration unit with the largest target weight as the target hardware acceleration unit.
[0041] Optionally, the target acceleration unit determination module further includes:
[0042] The pending task determination submodule is used to determine the hardware acceleration unit with the least number of pending tasks as the target hardware acceleration unit.
[0043] Optionally, the target acceleration unit determination module further includes:
[0044] A first weight determination submodule, configured to determine the hardware acceleration units as target hardware acceleration units in sequence when the queue weights are the same;
[0045] A second weight determination submodule, configured to calculate a target weight for each queue according to the sum of the queue weight and the current weight when the queue weights are different, and determine the hardware acceleration unit with the largest target weight as the first target hardware acceleration unit;
[0046] The to-be-processed task determination submodule is used to determine the first target hardware acceleration unit with the least number of to-be-processed tasks as the target hardware acceleration unit.
[0047] Optionally, the target acceleration unit determination module includes:
[0048] A first weight determination submodule, configured to determine the hardware acceleration units as target hardware acceleration units in sequence when the queue weights are the same;
[0049] A second weight determination submodule, configured to calculate a target weight for each queue according to the sum of the queue weight and the current weight when the queue weights are different, and determine the hardware acceleration unit with the largest target weight as the first target hardware acceleration unit;
[0050] A pending task determination submodule, used for determining the first target hardware acceleration unit having the least number of pending tasks as the second target hardware acceleration unit;
[0051] The delay determination submodule is used to determine the second target hardware acceleration unit with the lowest delay required to complete the task as the target hardware acceleration unit.
[0052] Optionally, the monitoring module includes:
[0053] The health monitoring threshold detection submodule is used for the task dispatching system to configure the health monitoring threshold for each hardware acceleration unit, and to detect the health monitoring threshold of each hardware acceleration unit during operation;
[0054] A sub-healthy hardware acceleration unit determination submodule is used to determine that the hardware acceleration unit is a sub-healthy hardware acceleration unit when it is detected that the health monitoring threshold of the hardware acceleration unit is lower than the target threshold, and to issue a detection instruction to the sub-healthy hardware acceleration unit, wherein the detection instruction is used to determine whether the sub-healthy hardware acceleration unit has recovered to a healthy state;
[0055] The healthy hardware acceleration unit determination submodule is used to determine that the subhealthy hardware acceleration unit is restored to a healthy hardware acceleration unit and initialize the healthy hardware acceleration unit when the health monitoring threshold of the subhealthy hardware acceleration unit is higher than the target threshold.
[0056] Optionally, the monitoring module further includes:
[0057] A configuration submodule, used to configure the number of queues dispatched to the hardware acceleration units and the number of tasks contained in each queue according to the number of hardware acceleration units and hardware performance;
[0058] The queue binding submodule is used to bind the queues to the hardware acceleration units one by one, and each hardware acceleration unit executes the tasks in the corresponding queue.
[0059] A third aspect of an embodiment of the present application provides an electronic device, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the computer program is executed by the processor, the task dispatching method of the first aspect of the embodiment of the present application is implemented.
[0060] A fourth aspect of an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the task dispatching method of the first aspect of the embodiment of the present application is implemented.
[0061] Through the task dispatching method of the embodiment of the present application, the current parameter value of each hardware acceleration unit in the multiple hardware acceleration units, that is, the number of tasks to be processed by each hardware acceleration unit per unit time and the delay required for each hardware acceleration unit to complete a task per unit time, is monitored, and when a task to be dispatched is detected, based on the current parameter value of each hardware acceleration unit and in accordance with the target dispatching mode, a target hardware acceleration unit is determined, and the task to be dispatched is assigned to the target hardware acceleration unit.
[0062] In the present application, the binding situation of the hardware acceleration unit and the virtualized physical machine in the prior art is abandoned. By monitoring three parameters, namely the number of tasks to be processed by each hardware acceleration unit in a unit time and the delay required for each hardware acceleration unit to complete a task in a unit time, the amount of tasks performed by each hardware acceleration unit in the process of executing tasks can be judged in real time. When a new task dispatched by the host is received, the hardware acceleration unit can be reasonably selected to be dispatched according to the dispatch mode, so as to maximize the utilization of system resources and improve the overall performance of the system in processing IO tasks, enhance the flexibility and scalability of the system in processing IO tasks, and effectively reduce the congestion and delay that may occur when processing a large number of IO tasks of the same type, which not only greatly improves work efficiency, but also avoids waste of resources, saves costs, and achieves cost reduction and efficiency improvement in the field of task dispatching. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the description of the embodiments of the present application will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0064] Figure 1 is a flow chart of a task dispatching method proposed in one embodiment of the present application;
[0065] Figure 2 This is a schematic diagram of a task dispatch architecture proposed in an embodiment of the present application;
[0066] Figure 3 This is a configuration flow chart of a task dispatch provided by an embodiment of the present application;
[0067] Figure 4 is a structural block diagram of task dispatching provided by an embodiment of the present application;
[0068] Figure 5 It is a schematic diagram of an electronic device shown in an embodiment of the present application. DETAILED DESCRIPTION
[0069] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0070] In the drawings, the size of the constituent elements, the thickness of the layer or the area may be exaggerated for the sake of clarity. Therefore, any implementation of the present disclosure is not necessarily limited to the size shown in the drawings, and the shapes and sizes of the components in the drawings do not reflect the true proportions. In addition, the drawings schematically show ideal examples, and any implementation of the present disclosure is not limited to the shapes or values shown in the drawings.
[0071] Please refer to Figure 1 , Figure 1 This is a flowchart of a task dispatching method of the present application. Figure 1 As shown, the method may include steps S101 to S103:
[0072] Step S101: monitoring a current parameter value of each hardware acceleration unit in the plurality of hardware acceleration units, wherein the current parameter value includes at least one of the following: the number of tasks to be processed by each hardware acceleration unit in a unit time, and the delay required for each hardware acceleration unit to complete a task in a unit time;
[0073] Step S102: when a task to be dispatched is detected, based on the current parameter value of each hardware acceleration unit, a target hardware acceleration unit is determined according to a target dispatching mode;
[0074] Step S103: Allocate the task to be dispatched to the target hardware acceleration unit.
[0075] In the prior art, in the process of processing IO tasks, a hardware acceleration unit is usually bound to a virtualized physical machine, and the IO tasks issued by the host are all assigned to a specific hardware acceleration unit. However, this method has obvious disadvantages, such as waste of resource utilization, and it is easy to cause task congestion when receiving a large number of IO tasks, which greatly slows down the processing speed.
[0076] Therefore, a method for overall control and IO task allocation is proposed in the present application. The task system proposed in the present application is connected to multiple hardware acceleration units in communication, and monitors each hardware acceleration unit in real time. After obtaining the parameter values of each hardware acceleration unit through monitoring, when receiving a new IO task issued by the host, it can calculate according to the parameter values obtained through monitoring in the set dispatching mode, calculate the target hardware acceleration unit that is most suitable for processing this IO task, and then dispatch the task to it. In this way, the dispatching system uniformly allocates the IO tasks issued by the host, and performs overall control of each hardware acceleration unit according to the parameter values, avoiding the hardware acceleration unit being able to process only a single task, thereby causing task congestion, and can distribute IO tasks more evenly, thereby improving the overall processing efficiency and resource utilization of the system.
[0077] Step S101: monitoring a current parameter value of each hardware acceleration unit in the plurality of hardware acceleration units, wherein the current parameter value includes at least one of the following: the number of tasks to be processed by each hardware acceleration unit per unit time, and the delay required for each hardware acceleration unit to complete a task per unit time.
[0078] In the embodiment of the present application, the task dispatching system is connected to multiple hardware acceleration units in communication, and needs to monitor the parameter values in each hardware acceleration unit to determine the state of the hardware acceleration unit as the input part when dispatching tasks. Among them, the number of pending tasks for each hardware acceleration unit per unit time indicates the number of tasks that have been dispatched but not yet processed by each hardware acceleration unit, which is used to characterize the expected working state of the hardware acceleration unit; the delay required for each hardware acceleration unit to complete a task per unit time indicates the time required for the hardware acceleration unit to complete a task. Due to the different task amounts of each IO task, the speed at which each hardware acceleration unit processes tasks is also different. There is also time affected by transmission interaction and other reasons. The delay in completing a task is used to characterize the real-time work efficiency of the hardware acceleration unit.
[0079] Step S102: when a task to be dispatched is detected, a target hardware acceleration unit is determined based on the current parameter value of each hardware acceleration unit and in accordance with a target dispatching mode.
[0080] In the embodiment of the present application, when the task dispatching system detects the IO task issued by the host, it will comprehensively consider the current parameter values of each hardware acceleration unit monitored in step S101 and the target dispatching mode determined by the task dispatching system to select the hardware acceleration unit that is most suitable for executing the current IO task. The dispatching principle is to ensure that each hardware acceleration unit works as much as possible while avoiding too many tasks being dispatched to a certain hardware acceleration unit.
[0081] Step S103: Allocate the task to be dispatched to the target hardware acceleration unit.
[0082] In the embodiment of the present application, the task dispatching system also includes a task space to be dispatched for temporarily storing IO tasks issued by the host. When the host issues multiple IO tasks at the same time, the task dispatching system needs to dispatch each task to the hardware acceleration unit in turn, and the tasks that have not yet been dispatched are temporarily stored in the task space to be dispatched. After the task dispatching system determines which hardware acceleration unit the IO task issued by the current host will be dispatched to, the hardware acceleration unit is set as the target hardware acceleration unit, and the current IO task is dispatched to the target hardware acceleration unit.
[0083] For example, the specific implementation process is as follows Figure 2 As shown, Figure 2This is an architectural diagram of a task dispatching proposed in an embodiment of the present application. The host sends IO tasks to the task dispatching system. The task dispatching system first temporarily stores the IO tasks in the task space to be dispatched. Then, the configured task dispatching system determines which hardware acceleration unit to dispatch the IO tasks to based on the parameter values of each hardware acceleration unit monitored and the current dispatching mode. After determining the target acceleration unit, the dispatching module dispatches the IO tasks in the task space to be dispatched to the corresponding target acceleration unit.
[0084] In combination with the above embodiments, in one implementation, the present application further provides a task dispatching method, wherein the current parameter value further includes: a weight of each hardware acceleration unit, the target dispatching mode is a weight-based dispatching mode, the weight of the hardware acceleration unit is used to express the ability to process tasks, the weight includes a queue weight and a current weight, the queue weight is used to express the inherent hardware performance of the hardware acceleration unit, and the current weight is used to express the status of tasks currently being processed and to be processed by the hardware acceleration unit; when a task to be dispatched is detected, based on the current parameter value of each hardware acceleration unit, the target hardware acceleration unit is determined according to the target dispatching mode, specifically including the following contents:
[0085] In an embodiment of the present application, the task dispatching system monitors the parameter values of the hardware acceleration unit, which also includes the weight of the hardware acceleration unit. In hardware design, the design resources occupied by each hardware acceleration unit are not the same, so the processing capabilities of the hardware acceleration unit are also different. Therefore, the present application proposes to use the weight of the hardware acceleration unit to characterize the ability of the hardware acceleration unit to process tasks.
[0086] In actual applications, hardware acceleration units that occupy more resources usually process a single task faster than hardware acceleration units that occupy fewer resources. However, if tasks are only assigned to hardware acceleration units that occupy more resources, it will still cause uneven task distribution and some hardware acceleration units to be idle. Therefore, the weight of the hardware acceleration unit in this application includes the queue weight and the current weight, where the queue refers to the IO task queue to be executed by the hardware acceleration unit, and the queue weight is used to express the inherent hardware performance of the hardware acceleration unit, that is, the ability of the hardware acceleration unit corresponding to the queue to process tasks per unit time; the current weight represents the cumulative state of the hardware acceleration unit processing multiple IO tasks, which is used to express the status of the tasks being processed and the tasks to be processed.
[0087] First, when the queue weights are the same, the hardware acceleration units are sequentially determined as target hardware acceleration units.
[0088] In the embodiment of the present application, if the hardware resources of each hardware acceleration unit are evenly distributed, the queue weights corresponding to each hardware acceleration unit are consistent, that is, for the same IO task, the execution speed of each hardware acceleration unit should be consistent in theory. Therefore, when the queue weights are the same, each hardware acceleration unit is determined as the target hardware acceleration unit in sequence, and the IO tasks dispatched by the host are dispatched to each hardware acceleration unit in sequence.
[0089] However, usually the queue weights are different. When the queue weights are different, the target weight of each queue is calculated according to the sum of the queue weight and the current weight, and the hardware acceleration unit with the largest target weight is determined as the target hardware acceleration unit.
[0090] In an embodiment of the present application, if tasks are dispatched only according to queue weights, all tasks will be dispatched to the hardware acceleration unit with the highest queue weight. Therefore, in order to avoid this extremely unbalanced situation, it is necessary to consider the current weight of the hardware acceleration unit when allocating tasks according to the weights, and calculate the target weight of each queue by the sum of the queue weight and the current weight, and determine the hardware acceleration unit with the largest target weight as the target hardware acceleration unit.
[0091] Among them, the current weights of each hardware acceleration unit are not fixed. After the target acceleration unit is determined, the current weights of the hardware acceleration units that are not assigned to tasks will increase, while the weights of the hardware acceleration units assigned to tasks will decrease accordingly. In this way, the firmware situation of the hardware acceleration unit in this application and the tasks currently required to be performed by each hardware acceleration unit are combined, and comprehensive consideration is made to make IO task distribution more balanced.
[0092] For the above distribution mode based on weights, the present application gives a simple example here, which is only used as an illustration of the above embodiment. The specific situation should be based on the actual operation. The example is as follows:
[0093] The task dispatching system is connected to the three hardware acceleration units A, B, and C. The queue weights of the three hardware acceleration units are 3, 2, and 1, respectively, and the initial current weights are all 0. At this time, the task dispatching system receives 6 IO tasks sent by the host and temporarily stores them in the task space to be dispatched. In the first task dispatch, the target weight of each hardware acceleration unit is the sum of the queue weight and the initial current weight, which are 3, 2, and 1, respectively. Therefore, the first task is dispatched to queue A. The target weight at this time is not only used as the basis for the current task dispatch, but also inherited as the new current weight to participate in the calculation before the next task dispatch. At this time, the current weights of the three queues are 3, 2, and 1, respectively.
[0094] After the dispatch is completed, the current weight of queue A needs to be limited to avoid allocating all tasks to queue A. In this embodiment, the current weight minus the sum of the weights of all queues is used as the new current weight of queue A. The result obtained after calculation is -3. The whole process of dispatching a task is now completed. At this time, the current weights of the three queues are -3, 2, and 1 respectively.
[0095] The process of dispatching subsequent tasks is similar to that of the first task, which will not be described here. For details, please refer to Table 1.
[0096] Table 1 Distribution by weight
[0097]
[0098] Each allocation allocates an IO task to be processed. The current weight and tasks in the queue are the status of each queue after each task is completed. ①~⑥ are used to represent the number of the IO task.
[0099] As shown in Table 1, in the process of task allocation, tasks are evenly distributed to the three hardware acceleration units A, B, and C. Moreover, the hardware acceleration unit A with the highest queue weight is assigned the most tasks in the same time period, and the hardware acceleration unit C with the lowest queue weight is assigned the least tasks in the same time period. This weight-based distribution mode has stronger subjective ability than the previous distribution mode, and can take into account both balance and the overall progress of task processing, thus avoiding the waste of hardware acceleration unit resources.
[0100] In combination with the above embodiments, in one implementation, the present application further provides a task dispatching method, wherein the current parameter value includes: the number of tasks to be processed by each hardware acceleration unit in a unit time, the target dispatching mode is a task number monitoring mode, and based on the current parameter value of each hardware acceleration unit, the target hardware acceleration unit is determined according to the target dispatching mode, specifically including the following contents:
[0101] The hardware acceleration unit with the least number of tasks to be processed is determined as the target hardware acceleration unit.
[0102] In an embodiment of the present application, a dispatching mode of a task number monitoring mode is also proposed, that is, the number of tasks to be processed in each hardware acceleration unit is monitored, and by comparing the number of tasks to be processed in each hardware acceleration unit, the newly dispatched IO tasks of the host are distributed to the hardware acceleration unit with the least number of tasks to be processed.
[0103] Here we still take the content shown in Table 1 as an example. In the task number monitoring mode, if the host sends 6 IO tasks to 3 hardware acceleration units at the same time, the tasks will be sent in the order of hardware acceleration unit A, hardware acceleration unit A, hardware acceleration unit B, hardware acceleration unit B, hardware acceleration unit C, and hardware acceleration unit C according to the number of tasks to be processed. This method is more suitable for the situation where the queue weights of each hardware acceleration unit are consistent, that is, the inherent hardware performance of each hardware acceleration unit is relatively consistent, and the allocation is based on the number of tasks to be processed by each hardware acceleration unit, so that each hardware acceleration unit can be kept in working state.
[0104] Furthermore, although the task distribution can be relatively balanced according to the task number monitoring mode, when there are differences in the queue weights of the hardware acceleration units A, B, and C, there may be an imbalance in that the hardware acceleration units with high queue weights complete the assigned tasks, while the hardware acceleration units with low queue weights still have unfinished tasks.
[0105] Therefore, on this basis, combined with the above embodiments, in one implementation, the present application further provides a task dispatching method, wherein the current parameter value includes: the weight of each hardware acceleration unit, and the number of tasks to be processed by each hardware acceleration unit in unit time, wherein the weight of the hardware acceleration unit is used to express the ability to process tasks, and the weight includes a queue weight and a current weight, wherein the queue weight is used to express the inherent hardware performance of the hardware acceleration unit, and the current weight is used to express the tasks currently being processed and to be processed by the hardware acceleration unit; the target dispatching mode is a weighted and task number mixed monitoring mode; based on the current parameter value of each hardware acceleration unit, according to the target dispatching mode, the target hardware acceleration unit is determined, specifically including the following contents:
[0106] When the queue weights are the same, the hardware acceleration units are sequentially determined as target hardware acceleration units;
[0107] When the queue weights are different, the target weight of each queue is calculated according to the sum of the queue weight and the current weight, and the hardware acceleration unit with the largest target weight is determined as the first target hardware acceleration unit;
[0108] The first target hardware acceleration unit with the least number of tasks to be processed is determined as the target hardware acceleration unit.
[0109] In this embodiment, a target dispatching mode is also proposed as a weighted and task number mixed monitoring mode. This monitoring mode is based on the above two monitoring modes, that is, the target weight of each hardware acceleration unit is first judged, and the judgment method is as described above. If there are two queues with the same target weight, the judgment will continue according to the number of tasks to be processed in the queue, and the IO task will be dispatched to the hardware acceleration unit with fewer tasks.
[0110] Here we still use the above situation as an example. After the host sends the IO task, the target weight of each hardware acceleration unit should be the sum of the current weight and the queue weight. After comparison, if the target weights are the same, the number of tasks to be executed in the queues of each hardware acceleration unit is compared, and the new IO task is assigned to the hardware acceleration unit with fewer tasks to be executed.
[0111] In combination with the above embodiments, in one implementation, the present application further provides a task dispatching method, wherein the current parameter value includes: the weight of each hardware acceleration unit, and the number of tasks to be processed by each hardware acceleration unit in a unit time, and the delay required for each hardware acceleration unit to complete a task in a unit time, wherein the weight of the hardware acceleration unit is used to express the ability to process tasks, and the weight includes a queue weight and a current weight, wherein the queue weight is used to express the inherent hardware performance of the hardware acceleration unit, and the current weight is used to express the status of tasks currently being processed and to be processed by the hardware acceleration unit; the target dispatching mode is a mixed monitoring mode of weighted and number of tasks and delay required to complete the task; based on the current parameter value of each hardware acceleration unit, the target hardware acceleration unit is determined according to the target dispatching mode, which specifically includes the following contents:
[0112] When the queue weights are the same, the hardware acceleration units are sequentially determined as target hardware acceleration units;
[0113] When the queue weights are different, the target weight of each queue is calculated according to the sum of the queue weight and the current weight, and the hardware acceleration unit with the largest target weight is determined as the first target hardware acceleration unit;
[0114] Determine the first target hardware acceleration unit with the least number of tasks to be processed as the second target hardware acceleration unit;
[0115] The second target hardware acceleration unit with the lowest delay required to complete the task is determined as the target hardware acceleration unit.
[0116] In an embodiment of the present application, as described in the above step S101, the task dispatching system will also monitor the delay required for each hardware acceleration unit to complete a task. The delay required for each hardware acceleration unit to complete a task in unit time represents the time required for the hardware acceleration unit to complete a task. Since the task volume of each IO task is different, the speed at which each hardware acceleration unit processes tasks is also different. There is also time affected by transmission interaction and other reasons. The delay in completing a task is used to characterize the real-time work efficiency of the hardware acceleration unit.
[0117] Therefore, the present application also proposes that when the dispatching mode is a mixed monitoring mode of weighted and number of tasks and delay required to complete the task, and the target weight of the hardware acceleration unit and the number of tasks to be processed are consistent, it is necessary to compare the delay required for each hardware acceleration unit to complete the task, and dispatch the IO task to the hardware acceleration unit with the lowest delay. In this way, the task dispatching system can dispatch tasks to each hardware acceleration unit more evenly, and can ensure the efficiency of completing the IO tasks issued to the host.
[0118] To sum up, in the embodiments of the present application, a variety of target dispatch modes are proposed, such as a weight-based dispatch mode, a task number monitoring mode, a weighted and task number mixed monitoring mode, and a weighted and task number and delay required to complete the task mixed monitoring mode. In the actual application process, a more suitable dispatch mode should be selected to dispatch IO tasks according to the actual situation. It should be noted that the several task dispatch modes proposed here are only methods proposed based on parameter values such as the weight of the hardware acceleration unit, the number of tasks to be processed, and the delay in processing IO tasks. In actual applications, according to different parameter settings of the monitoring hardware acceleration unit, there should be more task dispatching methods, all of which are within the protection scope of the present application.
[0119] In combination with the above embodiments, in one implementation, the present application further provides a task dispatching method, in which a health monitoring threshold is configured for a hardware acceleration unit, and the health monitoring threshold is used to characterize whether the hardware acceleration unit can complete the task normally, and specifically includes the following contents:
[0120] First, the task dispatching system configures a health monitoring threshold for each hardware acceleration unit, and detects the health monitoring threshold of each hardware acceleration unit during operation.
[0121] In an embodiment of the present application, the task dispatching system's monitoring of each hardware acceleration unit also includes monitoring its health status. First, it is necessary to configure a health threshold for each hardware acceleration unit. The health threshold is used to describe the health status of the hardware acceleration unit. Hardware acceleration units above the health threshold are regarded as healthy hardware acceleration units, and hardware acceleration units below the health threshold are regarded as sub-healthy hardware acceleration units. Among them, healthy hardware acceleration units can normally perform the work of processing IO tasks, while sub-healthy hardware acceleration units are regarded as having a fault or suspending work, etc., and cannot normally perform the work of processing IO tasks. During operation, each hardware acceleration unit is monitored in real time, which helps the task dispatching system to judge the health level of the hardware acceleration unit and prevent the hardware acceleration unit from still dispatching tasks to it when it fails, resulting in the inability to complete the task.
[0122] Then, when it is detected that the health monitoring threshold of the hardware acceleration unit is lower than the target threshold, the hardware acceleration unit is determined to be a sub-healthy hardware acceleration unit, and a detection instruction is issued to the sub-healthy hardware acceleration unit, wherein the detection instruction is used to determine whether the sub-healthy hardware acceleration unit has recovered to a healthy state.
[0123] In an embodiment of the present application, when it is detected that the hardware acceleration unit is in a sub-healthy state, the task dispatching system will no longer dispatch IO tasks to it, that is, the hardware acceleration unit in the sub-healthy state will not be used as the target acceleration unit. In addition, the task dispatching system will issue a detection instruction to the sub-healthy hardware acceleration unit, monitor its health status in real time, and detect whether it has recovered to a healthy state.
[0124] Finally, when the health monitoring threshold of the sub-healthy hardware acceleration unit is higher than the target threshold, it is determined that the sub-healthy hardware acceleration unit is restored to a healthy hardware acceleration unit, and the healthy hardware acceleration unit is initialized.
[0125] In an embodiment of the present application, after monitoring that a sub-healthy hardware acceleration unit has recovered to a healthy hardware acceleration unit, the healthy hardware acceleration unit will be initialized. After the initialization is completed, the task dispatching system will dispatch IO tasks to the hardware acceleration unit that has recovered to a healthy state when necessary in accordance with the monitoring mode described above.
[0126] In combination with the above embodiments, in one implementation, the present application further provides a task dispatching method, in which the method specifically includes the following steps:
[0127] First, according to the number of hardware acceleration units and hardware performance, configure the number of queues dispatched to the hardware acceleration units and the number of tasks contained in each queue;
[0128] Then, the queues are bound to the hardware acceleration units one by one, and each hardware acceleration unit executes the tasks in the corresponding queue.
[0129] In the embodiment of the present application, before the task dispatching system dispatches tasks, a series of initialization configurations need to be performed, such as Figure 2 The configuration module shown configures the number of queues dispatched to the hardware acceleration units according to the number of hardware acceleration units. The number of queues corresponds to the number of hardware acceleration units one by one, and each hardware acceleration unit executes the tasks in its corresponding queue. The number of tasks that can be accommodated in each queue is configured according to the hardware performance of each hardware acceleration unit to avoid dispatching too many tasks to hardware acceleration units with poor performance, which will cause task congestion.
[0130] Furthermore, the configuration module in this application should also include the configuration of the queue weights and target dispatch modes of each hardware acceleration unit as described above, as well as the configuration of the monitoring sliding window when the task dispatch system monitors. The monitoring sliding window is used to express the sliding window value and reserved value when the task dispatch system monitors each hardware acceleration unit. In the actual implementation process, the monitoring sliding window can be configured by configuring the register. Setting the sliding window value reduces the resource consumption caused by real-time monitoring of the hardware acceleration unit.
[0131] For example, the specific implementation process is as follows Figure 3 As shown, Figure 3 This is a configuration flow chart of task dispatch proposed in one embodiment of the present application. During the initialization process, the task dispatch system needs to configure the monitoring sliding window, the dispatch queue, the binding relationship between the queue and the hardware acceleration unit, the healthy monitoring threshold, and the target dispatch mode, and the configured parameters can be adjusted.
[0132] Based on the same design concept, an embodiment of the present application provides a task dispatching device. Figure 4 , Figure 4 This is a structural diagram of task dispatching provided by an embodiment of the present application. Figure 4 As shown, the device comprises:
[0133] A monitoring module, configured to monitor a current parameter value of each hardware acceleration unit in the plurality of hardware acceleration units, wherein the current parameter value includes at least one of the following: the number of tasks to be processed by each hardware acceleration unit in a unit time, and the delay required for each hardware acceleration unit to complete a task in a unit time;
[0134] A target acceleration unit determination module is used to determine the target hardware acceleration unit according to the target dispatch mode based on the current parameter value of each hardware acceleration unit when a task to be dispatched is detected;
[0135] The dispatching module is used to assign the task to be dispatched to the target hardware acceleration unit.
[0136] Optionally, the target acceleration unit determination module further includes:
[0137] A first weight determination submodule, configured to determine the hardware acceleration units as target hardware acceleration units in sequence when the queue weights are the same;
[0138] The second weight determination submodule is used to calculate the target weight of each queue according to the sum of the queue weight and the current weight when the queue weights are different, and determine the hardware acceleration unit with the largest target weight as the target hardware acceleration unit.
[0139] Optionally, the target acceleration unit determination module further includes:
[0140] The pending task determination submodule is used to determine the hardware acceleration unit with the least number of pending tasks as the target hardware acceleration unit.
[0141] Optionally, the target acceleration unit determination module further includes:
[0142] A first weight determination submodule, configured to determine the hardware acceleration units as target hardware acceleration units in sequence when the queue weights are the same;
[0143] A second weight determination submodule, configured to calculate a target weight for each queue according to the sum of the queue weight and the current weight when the queue weights are different, and determine the hardware acceleration unit with the largest target weight as the first target hardware acceleration unit;
[0144] The to-be-processed task determination submodule is used to determine the first target hardware acceleration unit with the least number of to-be-processed tasks as the target hardware acceleration unit.
[0145] Optionally, the target acceleration unit determination module includes:
[0146] A first weight determination submodule, configured to determine the hardware acceleration units as target hardware acceleration units in sequence when the queue weights are the same;
[0147] A second weight determination submodule, configured to calculate a target weight for each queue according to the sum of the queue weight and the current weight when the queue weights are different, and determine the hardware acceleration unit with the largest target weight as the first target hardware acceleration unit;
[0148] A pending task determination submodule, used for determining the first target hardware acceleration unit having the least number of pending tasks as the second target hardware acceleration unit;
[0149] The delay determination submodule is used to determine the second target hardware acceleration unit with the lowest delay required to complete the task as the target hardware acceleration unit.
[0150] Optionally, the monitoring module includes:
[0151] The health monitoring threshold detection submodule is used for the task dispatching system to configure the health monitoring threshold for each hardware acceleration unit, and to detect the health monitoring threshold of each hardware acceleration unit during operation;
[0152] A sub-healthy hardware acceleration unit determination submodule is used to determine that the hardware acceleration unit is a sub-healthy hardware acceleration unit when it is detected that the health monitoring threshold of the hardware acceleration unit is lower than the target threshold, and to issue a detection instruction to the sub-healthy hardware acceleration unit, wherein the detection instruction is used to determine whether the sub-healthy hardware acceleration unit has recovered to a healthy state;
[0153] The healthy hardware acceleration unit determination submodule is used to determine that the subhealthy hardware acceleration unit is restored to a healthy hardware acceleration unit and initialize the healthy hardware acceleration unit when the health monitoring threshold of the subhealthy hardware acceleration unit is higher than the target threshold.
[0154] Optionally, the monitoring module further includes:
[0155] A configuration submodule, used to configure the number of queues dispatched to the hardware acceleration units and the number of tasks contained in each queue according to the number of hardware acceleration units and hardware performance;
[0156] The queue binding submodule is used to bind the queues to the hardware acceleration units one by one, and each hardware acceleration unit executes the tasks in the corresponding queue.
[0157] Based on the same design concept, another embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps in the task dispatching method described in any of the above embodiments of the present application are implemented.
[0158] Based on the same design concept, another embodiment of the present application provides an electronic device, such as Figure 5 shown. Figure 5 1 is a schematic diagram of an electronic device shown in an embodiment of the present application. The electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, and when the processor executes, the steps in the task dispatching method described in any of the above embodiments of the present application are implemented.
[0159] As for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0160] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.
[0161] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, devices, or computer program products. Therefore, the embodiments of the present application may adopt the form of complete hardware embodiments, complete software embodiments, or embodiments in combination with software and hardware. Moreover, the embodiments of the present application may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0162] The embodiments of the present application are described with reference to the flowcharts and / or block diagrams of the methods, terminal devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0163] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing terminal device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured product including an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0164] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device so that a series of operating steps are executed on the computer or other programmable terminal device to produce a computer-implemented process, thereby providing instructions for executing on the computer or other programmable terminal device to implement the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0165] Although the preferred embodiments of the present application have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the embodiments of the present application.
[0166] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or terminal device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or terminal device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or terminal device including the elements.
[0167] The above is a detailed introduction to a task dispatching method, device, equipment and medium provided by the present application. Specific examples are used in this article to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea; at the same time, for general technical personnel in this field, according to the idea of the present application, there will be changes in the specific implementation method and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.
Claims
1. A task dispatching method, characterized in that: Applied to a task dispatching system, the task dispatching system is communicatively connected with a plurality of hardware acceleration units, including: Monitoring a current parameter value of each hardware acceleration unit in the plurality of hardware acceleration units, wherein the current parameter value includes at least one of the following: the number of tasks to be processed by each hardware acceleration unit within a unit time, and the delay required for each hardware acceleration unit to complete a task within a unit time; When a task to be dispatched is detected, a target hardware acceleration unit is determined according to a target dispatch mode based on the current parameter value of each hardware acceleration unit; Allocate the task to be dispatched to the target hardware acceleration unit.
2. A task dispatching method according to claim 1, characterized in that: The current parameter value also includes: the weight of each hardware acceleration unit, the target dispatch mode is a weight-based dispatch mode, the weight of the hardware acceleration unit is used to express the ability to process tasks, the weight includes a queue weight and a current weight, the queue weight is used to express the inherent hardware performance of the hardware acceleration unit, and the current weight is used to express the status of the tasks currently being processed and to be processed by the hardware acceleration unit; when a task to be dispatched is detected, based on the current parameter value of each hardware acceleration unit, the target hardware acceleration unit is determined according to the target dispatch mode, including: When the queue weights are the same, the hardware acceleration units are sequentially determined as target hardware acceleration units; When the queue weights are different, the target weight of each queue is calculated according to the sum of the queue weight and the current weight, and the hardware acceleration unit with the largest target weight is determined as the target hardware acceleration unit.
3. A task dispatching method according to claim 1, characterized in that: The current parameter value includes: the number of tasks to be processed by each hardware acceleration unit in a unit time, the target dispatch mode is the task number monitoring mode, and based on the current parameter value of each hardware acceleration unit, the target hardware acceleration unit is determined according to the target dispatch mode, including: The hardware acceleration unit with the least number of tasks to be processed is determined as the target hardware acceleration unit.
4. A task dispatching method according to claim 1, characterized in that: The current parameter value includes: the weight of each hardware acceleration unit, and the number of tasks to be processed by each hardware acceleration unit in a unit time, wherein the weight of the hardware acceleration unit is used to express the ability to process tasks, and the weight includes a queue weight and a current weight, wherein the queue weight is used to express the inherent hardware performance of the hardware acceleration unit, and the current weight is used to express the tasks currently being processed and to be processed by the hardware acceleration unit; the target dispatch mode is a weighted and task number mixed monitoring mode; based on the current parameter value of each hardware acceleration unit, according to the target dispatch mode, the target hardware acceleration unit is determined, including: When the queue weights are the same, the hardware acceleration units are sequentially determined as target hardware acceleration units; When the queue weights are different, the target weight of each queue is calculated according to the sum of the queue weight and the current weight, and the hardware acceleration unit with the largest target weight is determined as the first target hardware acceleration unit; The first target hardware acceleration unit with the least number of tasks to be processed is determined as the target hardware acceleration unit.
5. A task dispatching method according to claim 1, characterized in that: The current parameter value includes: the weight of each hardware acceleration unit, the number of tasks to be processed by each hardware acceleration unit in a unit time, and the delay required for each hardware acceleration unit to complete a task in a unit time, wherein the weight of the hardware acceleration unit is used to express the ability to process tasks, and the weight includes a queue weight and a current weight, the queue weight is used to express the inherent hardware performance of the hardware acceleration unit, and the current weight is used to express the tasks currently being processed and to be processed by the hardware acceleration unit; the target dispatch mode is a mixed monitoring mode of weighted and task number and delay required to complete the task; based on the current parameter value of each hardware acceleration unit, according to the target dispatch mode, the target hardware acceleration unit is determined, including: When the queue weights are the same, the hardware acceleration units are sequentially determined as target hardware acceleration units; When the queue weights are different, the target weight of each queue is calculated according to the sum of the queue weight and the current weight, and the hardware acceleration unit with the largest target weight is determined as the first target hardware acceleration unit; Determine the first target hardware acceleration unit with the least number of tasks to be processed as the second target hardware acceleration unit; The second target hardware acceleration unit with the lowest delay required to complete the task is determined as the target hardware acceleration unit.
6. A task dispatching method according to claim 1, characterized in that: A health monitoring threshold is configured for the hardware acceleration unit. The health monitoring threshold is used to indicate whether the hardware acceleration unit can complete the task normally, including: The task dispatching system configures the health monitoring threshold for each hardware acceleration unit and detects the health monitoring threshold of each hardware acceleration unit during operation; When it is detected that the health monitoring threshold of the hardware acceleration unit is lower than the target threshold, the hardware acceleration unit is determined to be a sub-healthy hardware acceleration unit, and a detection instruction is issued to the sub-healthy hardware acceleration unit, wherein the detection instruction is used to determine whether the sub-healthy hardware acceleration unit has recovered to a healthy state; When the health monitoring threshold of the sub-healthy hardware acceleration unit is higher than the target threshold, it is determined that the sub-healthy hardware acceleration unit is restored to a healthy hardware acceleration unit, and the healthy hardware acceleration unit is initialized.
7. A task dispatching method according to claim 1, characterized in that: The method further comprises: According to the number of hardware acceleration units and hardware performance, configure the number of queues dispatched to the hardware acceleration units and the number of tasks contained in each queue; The queues are bound to the hardware acceleration units one by one, and each hardware acceleration unit executes the tasks in the corresponding queue.
8. A task dispatching device, characterized in that: The device comprises: A monitoring module, configured to monitor a current parameter value of each hardware acceleration unit in the plurality of hardware acceleration units, wherein the current parameter value includes at least one of the following: the number of tasks to be processed by each hardware acceleration unit in a unit time, and the delay required for each hardware acceleration unit to complete a task in a unit time; A target acceleration unit determination module is used to determine the target hardware acceleration unit according to the target dispatching mode based on the current parameter value of each hardware acceleration unit when a task to be dispatched is detected; The dispatching module is used to assign the task to be dispatched to the target hardware acceleration unit.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the computer program is executed by the processor, the task dispatching method according to any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the task dispatching method according to any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Inference task processing method and device, electronic equipment and storage medium
CN120560867A