A computing kernel dynamic scheduling system and method
Through the dynamic scheduling system of the computing kernel, the computing resource utilization of intelligent computing power chips is optimized, and the problems of waste of computing power and real-time tasks are solved, and the efficiency and throughput of task execution are improved.
Patent Information
- Application Number
- CN202411863320.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-17
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2044-12-17
AI Technical Summary
The intelligent computing power chip has problems of waste of computing power and real-time tasks when scheduling computing cores, which limits the efficiency of task execution.
The dynamic scheduling system of the computing kernel is adopted, including the request management module, the scheduling management module, the monitoring management module and the operation management module. The utilization of computing resources is optimized by obtaining the computing kernel to be run, determining the post-coefficient, collecting operation status data, and distributing the computing kernel.
It improves the efficiency of task execution through intelligent computing power chips, reduces computing power waste and improves the real-time and throughput of tasks.
Smart Images

Figure CN119336510B_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of intelligent computing technology, and in particular, to a computing core dynamic scheduling system and method. Background Art
[0002] With the rapid development of artificial intelligence technology, especially the emergence of large models, the demand for intelligent computing power has increased sharply.
[0003] However, when intelligent computing power chips such as Graphics Processing Unit (GPU) schedule the computing cores of different tasks to run in different computing units, problems such as computing power waste and interference with task real-time performance may occur due to the fixed computing power and black-box scheduling mechanism set in the intelligent computing power chips, which in turn limits the efficiency of task execution through the intelligent computing power chips.
[0004] For example: A GPU chip generally consists of multiple computing units with independent computing capabilities (Streaming Multiprocessor), and these computing units contain multiple computing cores (Streaming Processor), as well as components such as instruction caches, L1 caches, and shared memory. However, the computing power of each computing core in these computing units is fixed, and since the scheduling mechanism inside the GPU chip is a black box for the tasks being executed and is limited by hardware capabilities, the scheduling mechanism used when scheduling the computing cores of the tasks being executed to run in the GPU chip is a simple first-in, first-out scheduling mechanism, resulting in the computing power of the computing cores and computing units in the GPU chip not being fully utilized, thus causing computing power waste.
[0005] For another example: Different computing tasks to be executed by the GPU chip often have different requirements for real-time performance, that is, some tasks have high requirements for real-time performance, while some tasks have no requirements for real-time performance. When the above two types of tasks are executed simultaneously, problems such as resource competition and scheduling strategy conflicts may occur, which in turn leads to a decrease in the real-time performance and throughput of the GPU chip for executing computing tasks.
[0006] Therefore, how to improve the efficiency of task execution through intelligent computing power chips is an urgent problem to be solved. Summary of the Invention
[0007] This specification provides a computing core dynamic scheduling system and method to partially solve the above problems existing in the prior art.
[0008] This specification adopts the following technical solutions:
[0009] This specification provides a computing kernel dynamic scheduling system, which includes a request management module, a scheduling management module, a monitoring management module, and an operation management module;
[0010] The request management module is used to obtain at least one computing kernel to be run, where the computing kernel to be run is a computing business logic code segment of a target task that needs to be executed;
[0011] The scheduling management module is used to, for each computing kernel to be run, determine a post coefficient of the computing kernel to be run according to the execution characteristic data of the computing kernel to be run, and according to the post coefficients of each computing kernel to be run and the running state data of a preset specified chip collected by the monitoring management module, determine a target computing kernel from each computing kernel to be run, and determine a target processing unit from each processing unit of the specified chip, and send the target computing kernel and the target processing unit to the operation management module, where the execution characteristic data is used to reflect the allowed delay size of the computing kernel to be run, including at least one of: the task type of the target task to which the computing kernel to be run belongs, the waiting time of the computing kernel to be run, the estimated completion time of the computing kernel to be run, and the service performance index parameters corresponding to the computing kernel to be run, and the running state data includes: state data used to reflect the state of each processing unit in the specified chip and video memory resource occupancy data;
[0012] The operation management module is used to, after receiving the target computing kernel and the target processing unit, dispatch the target computing kernel to the target processing unit for running.
[0013] Optionally, the request management module is used to obtain a task request, determine a target task that needs to be executed according to the task request, and determine at least one computing kernel to be run that needs to be run through a specified chip when executing the target task. For each computing kernel to be run, according to the task type of the target task to which the computing kernel to be run belongs, store the computing kernel to be run in the waiting queue that matches the computing kernel to be run in each preset waiting queue. The waiting queues include: a real-time task waiting queue, a non-real-time task waiting queue, and an interrupt task waiting queue;
[0014] The scheduling management module is used to determine the post coefficient of each computing kernel to be run in each waiting queue.
[0015] Optionally, the task types of the target tasks include: real-time tasks and non-real-time tasks;
[0016] The scheduling management module is used to, when determining that the task type of the target task is a real-time task, for each computing kernel to be run, determine the post coefficient of the computing kernel to be run according to the service performance index parameter and the estimated completion time corresponding to the computing kernel to be run.
[0017] Optionally, the scheduling management module is used to, when determining that the computing kernel to be run is the target computing kernel according to the post coefficient of the computing kernel to be run, and determining that there is no idle processing unit among the processing units according to the operation state data, determine the target processing unit from the processing units according to the service performance index parameters of the computing kernels being run by each processing unit, and send the target computing kernel and the target processing unit to the operation management module;
[0018] The operation management module is used to, when receiving the target computing kernel and the target processing unit, and determining that there is a computing kernel being run by the target processing unit, stop the computing kernel being run by the target processing unit, and dispatch the target computing kernel to the target processing unit for execution.
[0019] Optionally, the processing unit includes: a computing unit, and computing cores in the computing unit;
[0020] The scheduling management module is used to, when determining that there is a computing kernel with interrupted operation in the target processing unit in the specified chip, and determining that the target processing unit is a computing unit, re-acquire the operation state data of the specified chip, and when determining that there are idle computing cores among the processing units according to the operation state data, use the idle computing cores as interrupted task processing units, and use the computing kernel with interrupted operation as the interrupted computing kernel, and send the interrupted computing kernel and the interrupted task processing unit to the operation management module;
[0021] The operation management module is used to, after receiving the interrupted computing kernel and the interrupted task processing unit, dispatch the interrupted computing kernel to the interrupted task processing unit for running.
[0022] Optionally, the task type of the target task includes: real-time task, non-real-time task;
[0023] The scheduling management module is used to, when determining that the task type of the target task is a non-real-time task, for each computing kernel to be run, determine the post coefficient of the computing kernel to be run according to the waiting time and the estimated completion time corresponding to the computing kernel to be run.
[0024] Optionally, the scheduling management module is configured to, when determining that the to-be-run computing kernel is a target computing kernel according to the post coefficient of the to-be-run computing kernel and determining that there are processing units in an idle state among the processing units according to the operation state data, use the processing units in the idle state as target processing units, and when determining that the number of floating-point operations required to run the to-be-run computing kernel is less than a preset threshold, determine at least some of the to-be-run computing kernels from the to-be-run computing kernels to form a target computing kernel set with the target computing kernel according to a preset target combination strategy, and send the target computing kernel set and the target processing units to the operation management module.
[0025] Optionally, the processing unit includes at least one of a computing unit and a computing core in the computing unit;
[0026] The operation management module is configured to, after receiving the target computing kernel set, determine whether the target processing unit is a computing unit. If so, for each computing kernel included in the target computing kernel set, determine a selected computing core from the computing cores of the target processing unit, and dispatch the computing kernel to the selected computing core for running.
[0027] Optionally, the operation management module is configured to create a control thread in the specified computing core of the target processing unit, and through the control thread, for each computing kernel included in the target computing kernel set, determine a selected computing core from the computing cores of the target processing unit, and dispatch the computing kernel to the selected computing core for running.
[0028] Optionally, the operation management module is configured to monitor the computing cores of the target processing unit through the control thread, so as to, when determining that there are computing cores that have completed running among the computing cores of the target processing unit, use the computing cores that have completed running as idle computing cores, and upload the state data of the idle computing cores and the video memory resource occupancy data after the idle computing cores have completed running to the monitoring management module.
[0029] Optionally, the monitoring management module is configured to call a preset query interface to send a query request for the operation state data of the to-be-query processed unit in the specified chip to the operation management module, and receive the state data of the to-be-query processed unit and the video memory resource occupancy data returned by the operation management module according to the operation state data query request; and,
[0030] be configured to receive the state data of the idle computing cores and the video memory resource occupancy data after the idle computing cores have completed running uploaded by the operation management module.
[0031] Optionally, the scheduling management module is configured to, when determining that the number of floating-point operations required for running the to-be-run computing kernel is less than a preset threshold, determine a target combination policy from preset combination policies according to the service performance metric parameters corresponding to each to-be-run computing kernel, and according to the target combination policy, determine at least some of the to-be-run computing kernels from the to-be-run computing kernels to form a target computing kernel set with the target computing kernel, and send the target computing kernel set and the target processing unit to the operation management module. The combination policies include at least one of a first combination policy and a second combination policy. The first combination policy is used to determine a target computing kernel set according to the video memory resource occupancy data of the specified chip and the video memory resource data required for running each to-be-run computing kernel. The second combination policy is used to determine a target computing kernel set according to the number of floating-point operations required for running each to-be-run computing kernel.
[0032] This specification provides a method for dynamically scheduling computing kernels. The method is applied to a computing kernel dynamic scheduling system, which includes a request management module, a scheduling management module, a monitoring management module, and an operation management module. The method includes:
[0033] Receiving a task request, and obtaining at least one to-be-run computing kernel according to the task request through the request management module, where the to-be-run computing kernel is a computing service logic code segment for executing the target task corresponding to the task request;
[0034] Receiving, by the scheduling management module, for each to-be-run computing kernel, a post coefficient of the to-be-run computing kernel determined according to the execution feature data of the to-be-run computing kernel, and according to the post coefficient of each to-be-run computing kernel, and the running state data of a preset specified chip collected by the monitoring management module, determining a target computing kernel from the to-be-run computing kernels, and determining a target processing unit from the processing units of the specified chip. The execution feature data is used to reflect the allowed delay size of the to-be-run computing kernel, and includes at least one of the task type of the target task to which the to-be-run computing kernel belongs, the waiting time of the to-be-run computing kernel, the estimated completion time of the to-be-run computing kernel, and the service performance metric parameters corresponding to the to-be-run computing kernel. The running state data includes state data for reflecting the state of each processing unit in the specified chip and video memory resource occupancy data;
[0035] Sending the target computing kernel and the target processing unit to the operation management module, so that after receiving the target computing kernel and the target processing unit, the operation management module dispatches the target computing kernel to run in the target processing unit.
[0036] The above at least one technical solution adopted in this specification can achieve the following beneficial effects:
[0037] In the computing core dynamic scheduling system provided in this specification, the computing core dynamic scheduling system includes: a request management module, a scheduling management module, a monitoring management module, and an operation management module. Among them, the request management module is used to obtain at least one computing core to be run, and the computing core to be run here is the computing service logic code segment of the target task to be executed. The scheduling management module is used for each computing core to be run, according to the execution characteristic data of the computing core to be run, to determine the post coefficient of the computing core to be run, and according to the post coefficient of each computing core to be run, and the operation status data of the specified chip collected by the monitoring management module, to determine the target computing core from each computing core to be run, and to determine the target processing unit from each processing unit of the specified chip, and send the target computing core and the target processing unit to the operation management module. Here, the execution characteristic data is used to reflect the allowable delay size of the computing core to be run, including: at least one of the task type of the target task to which the computing core to be run belongs, the waiting time of the computing core to be run, the estimated completion time of the computing core to be run, and the service performance index parameters corresponding to the computing core to be run. The operation status data includes: the status data used to reflect the status of each processing unit in the specified chip and the video memory resource occupancy data. The operation management module is used to dispatch the target computing core to the target processing unit for running after receiving the target computing core and the target processing unit.
[0038] As can be seen from the above method, through the request management module of the computing core dynamic scheduling system, each computing core to be run belonging to different target tasks can be stored in different queues for management, and through the scheduling management module, for each computing core in the queue, according to the allowable degree of delay of each computing core, the post coefficient of each computing core is determined, and then each computing core can be scheduled to the corresponding target computing core for running according to the post coefficient of each computing core, so as to improve the efficiency of running each target task through the specified chip. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] The drawings described herein are used to provide a further understanding of this specification and form a part of this specification. The illustrative embodiments of this specification and their descriptions are used to explain this specification and do not constitute an improper limitation of this specification. In the drawings:
[0040] Figure 1 It is a schematic diagram of a computing core dynamic scheduling system provided in this specification;
[0041] Figure 2It is a schematic diagram of the computing core scheduling process provided in this specification;
[0042] Figure 3 It is a schematic diagram of the computing core monitoring method provided in this specification;
[0043] Figure 4 It is a schematic diagram of the computing core operation control method provided in this specification;
[0044] Figure 5 It is a schematic flow diagram of a computing core dynamic scheduling method provided in this specification;
[0045] Figure 6 It is a schematic diagram of a computing core dynamic scheduling device provided in this specification;
[0046] Figure 7 It is provided in this specification corresponding to Figure 5 the schematic diagram of the electronic device. Specific Embodiments
[0047] To make the objectives, technical solutions, and advantages of this specification clearer, the technical solutions of this specification will be clearly and completely described below in conjunction with specific embodiments of this specification and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this specification, rather than all of the embodiments. Based on the embodiments in this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this specification.
[0048] The following will describe in detail the technical solutions provided in each embodiment of this specification with reference to the accompanying drawings.
[0049] Currently, for intelligent computing power chips such as Graphics Processing Unit (GPU) and Neural network Processing Unit (NPU), the computing power of each processing unit is often fixed. Taking the A100 chip as an example, each computing unit is fixed with 64 computing cores, and the number of floating-point operations that each computing core can execute per second is fixed. At this time, when the Central Processing Unit (CPU) transfers a smaller computing core (e.g., a computing core that only requires 12 computing cores to run) to the A100 chip for execution, it may cause waste of computing power resources of some computing cores in the computing unit allocated for this computing core, and thus the task processing efficiency of the chip is often low.
[0050] In addition, when the intelligent computing power chip is running the computing kernels of a large number of non-real-time tasks, the computing kernels of newly allocated real-time tasks may not be able to obtain sufficient computing power to run, resulting in the failure to meet the real-time requirements of real-time tasks. Or when the intelligent computing power chip is running the computing kernels of a large number of real-time tasks, the computing kernels of non-real-time tasks may be infinitely delayed and thus "starved".
[0051] Based on this, a computing kernel dynamic scheduling system is provided in this specification, specifically as Figure 1 shown.
[0052] Figure 1 It is a schematic diagram of a computing kernel dynamic scheduling system provided in this specification.
[0053] Combined with Figure 1 shown, the computing kernel dynamic scheduling system in this specification may include: a central processing unit and at least one specified chip. Among them, a request management module, a scheduling management module, a monitoring management module, and an operation management module may be set in the central processing unit, so that the central processing unit can schedule each computing kernel required to execute different tasks to each processing unit of the specified chip through these modules to improve the efficiency of task execution through the specified chip. Among them, the above-mentioned specified chip can be set according to actual needs. For example: a graphics processing unit, a neural network processing unit, etc. The above-mentioned each processing unit may refer to the computing unit in the specified chip or the computing core in the computing unit.
[0054] Specifically, when a user needs to execute a task, a task request can be sent to the CPU. After receiving the task request, the CPU can, through the request management module, determine the task to be executed according to the received task request as the target task, and can determine at least one computing kernel that needs to be run through the specified chip when executing the target task as the computing kernel to be run. Then, scheduling management can be performed on each computing kernel to be run of the target task to schedule each computing kernel to be run of the target task to each processing unit of the specified chip for running.
[0055] Among them, there are various methods for the above user to send a task request to the CPU. For example: the user can send a task request to the CPU through an application installed in the device used by the user. For another example: the user can input a control instruction through an input device and send the control instruction to the CPU, so that the CPU can obtain the task request sent by the user according to the received control instruction.
[0056] In the above content, the computing kernel of the target task may refer to the code segment corresponding to the computing business logic of the target task, that is, the computing kernel may refer to the core code part that undertakes the computing business logic in the target task. It can be understood that the computing kernel is the expression operation (such as: multi-linear function operation, etc.) required to execute the target task.
[0057] It should be noted that in the case of different complexities of the target task, the number of computing kernels included in the target task is also different. That is, the target task may only include one computing kernel or may include multiple computing kernels.
[0058] In an actual application scenario, the target tasks that need to be executed by a specified chip at the same time can be one or multiple. Among them, when the to-be-run computing kernels of multiple target tasks need to be scheduled to the specified chip for execution, in order to avoid problems such as resource competition and scheduling policy conflicts between different target tasks due to different real-time requirements of different target tasks. In this specification, the request management module can, for each to-be-run computing kernel of different target tasks, determine the waiting queue that matches the to-be-run computing kernel from the preset waiting queues according to the task type of the target task to which the to-be-run computing kernel belongs, and store the to-be-run computing kernel in the waiting queue that matches the to-be-run computing kernel, so as to facilitate the scheduling management module to schedule each to-be-run computing kernel stored in different waiting queues.
[0059] Among them, the above-mentioned task type of the target task can be used to reflect the demand of the target task for computing power resource allocation, such as: real-time tasks, non-real-time tasks. Here, real-time tasks usually require priority allocation of resources, such as CPU time, memory, computing power resources of the specified chip, etc., to ensure that real-time tasks can be executed in time. Here, non-real-time tasks usually allow delays in resource allocation and require the throughput of the whole composed of multiple non-real-time tasks.
[0060] The above-mentioned waiting queues can include: real-time task waiting queue, non-real-time task waiting queue, interrupt task waiting queue. Among them, the interrupt task waiting queue is used to store the computing kernels whose computing power resources have been preempted by other computing kernels during the running process.
[0061] Furthermore, the scheduling management module can schedule the computing kernels in each waiting queue according to the running state data of the specified chip returned by the monitoring management module, specifically as Figure 2 shown.
[0062] Figure 2 It is a schematic diagram of the computing kernel scheduling process provided in this specification.
[0063] Combined withFigure 2 It can be seen that the scheduling management module can determine the post - coefficient of each to - be - run computing kernel of the target task according to the execution characteristic data of the to - be - run computing kernel.
[0064] Among them, the above - mentioned execution characteristic data can be used to reflect the allowable delay size of the to - be - run computing kernel, including at least one of the following: the task type of the target task to which the to - be - run computing kernel belongs, the waiting time of the to - be - run computing kernel, the estimated completion time of the to - be - run computing kernel, and the service performance index parameters corresponding to the to - be - run computing kernel.
[0065] The lower the post - coefficient of the above - mentioned computing kernel, the more the computing kernel needs to be preferentially scheduled to run with computing power resources.
[0066] The above - mentioned service performance index parameters (Service Level Objective, SLO) can be index parameters set by users for the execution quality of real - time tasks. For example, index parameters such as response time and latency set for the real - time performance when executing real - time tasks.
[0067] For the sake of easy understanding, the following takes the response time as an example of the above - mentioned service performance index parameters to elaborate on the role of the above - mentioned service performance index parameters in detail. For example, when the response time of the target task set by the user is 10 milliseconds, it is required that the specified chip returns the execution result of the target task within 10 milliseconds.
[0068] The above - mentioned estimated completion time of the to - be - run computing kernel can be predicted through a preset neural network model or determined according to the completion times of each computing kernel in historical runs.
[0069] Among them, the method of determining the estimated completion time of the to - be - run computing kernel according to the completion times of each computing kernel in historical runs can be as follows: According to the sizes of different computing kernels, divide each computing kernel in historical runs into different intervals (for example: n intervals, where n can be set according to actual needs, such as: 10). Thus, for each interval, the mean value of the completion times of each computing kernel in historical runs belonging to this interval can be determined as the reference completion time corresponding to this interval. Furthermore, for each to - be - run computing kernel, according to the size of the computing kernel, determine the interval to which the to - be - run computing kernel belongs, and according to the reference completion time of the interval to which the to - be - run computing kernel belongs, determine the estimated completion time required to run the to - be - run computing kernel.
[0070] In this specification, for each processing unit, the operating status data of the processing unit includes: status data for reflecting the status of each processing unit in a specified chip, and video memory resource occupancy data. Among them, for each processing unit, the status data of the processing unit is used to reflect whether the processing unit is in an idle state. The video memory resource occupancy data here can be the size of the used video memory in the specified chip.
[0071] In addition, it can be seen from the above content that the target tasks of different task types often have different requirements for computing resource allocation. Therefore, the scheduling and management module can adopt different post - coefficient determination strategies for the computing kernels of target tasks belonging to different task types to determine their post - coefficients.
[0072] Specifically, for each to - be - run computing kernel in each waiting queue, if the scheduling and management module determines that the task type of the target task to which the to - be - run computing kernel belongs is a real - time task, then the post - coefficient of the to - be - run computing kernel can be determined according to the service performance index parameters and the estimated completion time corresponding to the to - be - run computing kernel. Specifically, it can refer to the following formula:
[0073]
[0074] In the above formula, is the adjustment coefficient, which can be set according to actual needs. For example: 0.5, is the remaining time after subtracting the current waiting time of the to - be - run computing kernel from the running time to complete the running of the to - be - run computing kernel according to the service performance index parameters (that is, the running time of the to - be - run computing kernel under the condition of meeting the service performance index parameters), is the estimated completion time required to run the to - be - run computing kernel.
[0075] It can be seen from the above formula that by adjusting the above - mentioned adjustment coefficient, under normal circumstances (such as when the waiting time of the to - be - run computing kernel of a non - real - time task is lower than a specified value), the post - coefficient of the to - be - run computing kernel of a real - time task determined is lower than the post - coefficient of the to - be - run computing kernel of a non - real - time task determined.
[0076] Furthermore, for each to - be - run computing kernel in each waiting queue, if the scheduling and management module determines that the task type of the target task to which the to - be - run computing kernel belongs is a non - real - time task, then the post - coefficient of the to - be - run computing kernel can be determined according to the waiting time and the estimated completion time corresponding to the to - be - run computing kernel. Specifically, it can refer to the following formula:
[0077]
[0078] In the above formula, The estimated completion time required to run the to-be-run compute kernel is the waiting time corresponding to the to-be-run compute kernel
[0079] As can be seen from the above, when the waiting time of the to-be-run compute kernel of a non-real-time task is higher than the specified value, the post coefficient determined for the to-be-run compute kernel of the non-real-time task starts to be lower than the post coefficient of the to-be-run compute kernel of a real-time task with a shorter waiting time, thereby avoiding the situation where the to-be-run compute kernel of a non-real-time task cannot be run for a long time
[0080] Furthermore, when the scheduling management module determines the to-be-run compute kernel of a real-time task as the target compute kernel according to the post coefficients of each to-be-run compute kernel in each waiting queue, and determines that there is an idle processing unit among each processing unit according to the operation status data of the specified chip, the idle processing unit can be used as the target processing unit, and then the target compute kernel and the target processing unit can be sent to the operation management module, so that the operation management module can, after receiving the target compute kernel and the target processing unit, dispatch the target compute kernel to the target processing unit for running
[0081] In addition, when the scheduling management module determines the compute kernel of a real-time task as the to-be-run compute kernel according to the post coefficients of each compute kernel in each waiting queue, and determines that there is no idle processing unit among each processing unit according to the operation status data of the specified chip, the target processing unit can be determined from each processing unit according to the service performance index parameters of the compute kernel currently running on each processing unit, and the target compute kernel and the target processing unit are sent to the operation management module, so that when the operation management module receives the target compute kernel and the target processing unit, and determines that there is a compute kernel currently running on the target processing unit, it stops the compute kernel currently running on the target processing unit and dispatches the target compute kernel to the target processing unit for execution
[0082] Specifically, the scheduling management module can determine each compute kernel that is currently being run among the compute kernels of each real-time task as each to-be-preempted compute kernel according to the metadata of each currently running compute kernel saved in the preset operation queue, and then, for each to-be-preempted compute kernel, determine the position of the metadata of the to-be-preempted compute kernel in the operation queue according to the remaining running time of the to-be-preempted compute kernel when meeting the service performance index parameters
[0083] Among them, the longer the remaining time of the to-be-preempted compute kernel, the closer the position of the metadata of the to-be-preempted compute kernel in the operation queue is to the queue head
[0084] The metadata of each of the above computing cores is updated by the operation management module. The metadata of each computing core includes: the task type of the target task to which the computing core belongs, the running time of the computing core, the service performance metric parameters corresponding to the computing core, the size of the video memory resources required for the computing core to run, the processing units occupied by the computing core, etc.
[0085] Further, the scheduling management module can start from the head of the run queue and sequentially judge, for the metadata of each computing core, whether the size of the video memory resources required for the computing core to run in the metadata of the computing core is lower than a preset occupation threshold. If so, it determines the processing unit where the computing core corresponding to the metadata of the computing core is located as the target processing unit.
[0086] In addition, if the scheduling management module does not determine the target processing unit after traversing the run queue, it can determine each computing core that is being run among the computing cores of each non-real-time task as each computing core to be preempted, and use the processing unit of the computing core to be preempted with the highest postposition coefficient among the computing cores to be preempted that are being run as the target processing unit.
[0087] Further, when the scheduling management module determines the computing core to be run of the non-real-time task as the target computing core according to the postposition coefficients of each computing core to be run in each waiting queue, the method of scheduling the target computing core to run in each processing unit can include two cases. The following will explain these three cases in detail respectively.
[0088] The first case is that when the scheduling management module determines that there are computing units in an idle state in each processing unit according to the running state data of the specified chip, it can use the processing unit in the idle state as the target processing unit. And when it determines that the number of floating-point operations required to run the computing core to be run is greater than a preset threshold, it can send the target computing core and the target processing unit to the operation management module.
[0089] The second case is that when the scheduling management module determines that there are computing units in an idle state in each processing unit according to the running state data of the specified chip, it can use the processing unit in the idle state as the target processing unit. And when it determines that the number of floating-point operations required to run the computing core to be run is less than a preset threshold, according to a preset target combination strategy, it determines at least some of the computing cores to be run among each computing core to be run to form a target computing core set with the target computing core, and sends the target computing core set and the target processing unit to the operation management module.
[0090] It should be noted that in actual application scenarios, since the video memory resources of a specified chip are usually limited, if the video memory resources are insufficient, it may take additional time to obtain and transfer data, thereby reducing the efficiency of the specified chip in running computing kernels. Therefore, in order to improve the utilization rate of the video memory resources of the specified chip, when the scheduling and management module determines that the remaining time of each to-be-run computing kernel stored in the real-time task waiting queue is greater than a preset remaining time threshold under the condition of meeting the service performance index parameters according to the service performance index parameter of each to-be-run computing kernel stored in the real-time task waiting queue, it can determine a preset first combination strategy as the target combination strategy. Furthermore, according to the target combination strategy, the target computing kernel set can be determined based on the video memory resource occupancy data of the specified chip and the video memory resource data required for each to-be-run computing kernel of the target task.
[0091] Of course, the scheduling and management module can also determine the target computing kernel set according to the target combination strategy, based on the video memory resource occupancy data of the specified chip and the video memory resource data required for each to-be-run computing kernel in the non-real-time task waiting queue.
[0092] Specifically, the scheduling and management module can determine the remaining video memory resource size of the specified chip according to the video memory resource occupancy data of the specified chip, and then, based on the video memory resource size required for each to-be-run computing kernel and the remaining video memory resource size of the specified chip, determine at least some of the to-be-run computing kernels and the target computing kernels to form the target computing kernel set from each to-be-run computing kernel.
[0093] In addition, when the scheduling and management module determines that there is a to-be-run computing kernel in each to-be-run computing kernel stored in the real-time task waiting queue whose remaining time is less than the preset remaining time threshold under the condition of meeting the service performance index parameters according to the service performance index parameter of each to-be-run computing kernel stored in the real-time task waiting queue, it can determine a preset second combination strategy as the target combination strategy, and then, according to the target combination strategy, determine the target computing kernel set based on the number of floating-point operations required for each to-be-run computing kernel during operation and the number of floating-point operations that the target processing unit can process per second.
[0094] As can be seen from the above, when the scheduling management module needs to combine the computing cores of non-real-time tasks and hand them over to the computing units in the specified chip for execution, it can determine whether the remaining time of the computing cores of the current real-time task is sufficient under the condition of meeting the service performance index parameters according to the service performance index parameters of the computing cores of the real-time task. If it is sufficient, it can aim to maximize the utilization rate of the video memory resources of the specified chip, combine the to-be-run computing cores of the non-real-time tasks, and hand them over to the target processing unit for operation, thereby improving the throughput rate of the specified chip for running each computing core. If it is not sufficient, it can aim to maximize the utilization rate of the computing resources of the processing units in the idle state, combine the to-be-run computing cores of the non-real-time tasks, thereby improving the satisfaction rate of the service performance index parameters of the computing cores of the real-time task.
[0095] Furthermore, after receiving the target computing core set and the target processing unit, the operation management module can determine whether the target processing unit is a computing unit. If so, for each computing core included in the target computing core set, it can determine the selected computing core from the computing cores of the target processing unit and dispatch the computing core to the selected computing core for operation.
[0096] Specifically, the operation management module can also create a control thread in the specified computing core of the target processing unit, and through the control thread, for each computing core included in the target computing core set, determine the selected computing core from the computing cores of the target processing unit, dispatch the computing core to the selected computing core for operation, and can monitor the computing cores of the target processing unit through the control thread to, when it is determined that there are computing cores that have completed operation among the computing cores of the target processing unit, use the computing cores that have completed operation as idle computing cores and upload the status data of the idle computing cores and the video memory resource occupancy data after the idle computing cores have completed operation to the monitoring management module.
[0097] Furthermore, there are two methods for the monitoring management module to collect the operation status data of the specified chip, specifically as Figure 3 shown.
[0098] Figure 3 It is a schematic diagram of the computing core monitoring method provided in this specification.
[0099] Combined with Figure 3It can be seen that the first method for the monitoring and management module to collect the operating status data of the specified chip is that the monitoring and management module can receive the status data of the idle computing cores uploaded by the operation management module and the video memory resource occupancy data after the idle computing cores complete their operations, and determine the operating status data of the specified chip according to the received status data of the idle computing cores uploaded by the operation management module and the video memory resource occupancy data after the idle computing cores complete their operations.
[0100] The second method for the monitoring and management module to collect the operating status data of the specified chip is that the monitoring and management module can call a preset query interface to send a query request for the operating status data of the processing unit to be queried in the specified chip to the operation management module, and receive the status data of the processing unit to be queried and the video memory resource occupancy data returned by the operation management module according to the query request for the operating status data, and determine the operating status data of the specified chip according to the received status data of the processing unit to be queried and the video memory resource occupancy data returned by the operation management module according to the query request for the operating status data.
[0101] Specifically, for each processing unit, the above-mentioned monitoring and management module can use the processing unit as the processing unit to be queried, and send a query request for the operating status data of the processing unit to be queried at a specified query period through a preset query interface, so that the operation management module can return the status data of the processing unit to be queried and the video memory resource occupancy data in response to the received query request for the operating status data.
[0102] Among them, the above-mentioned specified query period can be determined according to the operating status data of the specified chip, that is, if it is determined according to the operating status data of the specified chip that the number of processing units in the idle state in the specified chip is larger, the determined specified query period is longer.
[0103] It should be noted that since data transmission between the CPU and the specified chip needs to be carried out through the Peripheral Component Interconnect Express (PCIe), and the total amount of the PCIe transmission bandwidth (that is, the amount of data that the PCIe bus can transmit per unit time) between the CPU and the specified chip is limited. Therefore, when the monitoring and management module actively initiates a query request for the operating status data through a preset query interface, there may be a situation of PCIe transmission bandwidth competition because the PCIe transmission bandwidth between the CPU and the specified chip is being used by the computing tasks of each computing core of the specified chip, which may lead to a reduction in the execution efficiency of the target task.
[0104] As can be seen from the above, in order to avoid the impact of the queries initiated by the monitoring and management module on the execution efficiency of the target task, the monitoring and management module can combine two different methods to obtain the running status data of the specified chip. Among them, the operation management module monitors each computing core of the target processing unit through a control thread. When it is determined that there is a computing core that has completed running among the computing cores of the target processing unit, the computing core that has completed running is regarded as an idle computing core, and the status data of the idle computing core and the video memory resource occupancy data after the idle computing core has completed running are actively uploaded to the monitoring and management module. This method can reduce the frequency of queries initiated by the monitoring and management module actively, and thus can reduce the impact on the efficiency of executing the target task on the specified chip.
[0105] In addition, in order to avoid the PCIe resources of the specified chip being occupied, which may lead to a reduction in the efficiency of the specified chip in executing the target task, when the scheduling management module determines that there is a computing kernel in the specified chip that has interrupted its operation in the target processing unit, it can store it in the interrupted task waiting queue, or it can use the device-to-device (D2D) method to directly migrate the computing kernel that has interrupted its operation in the target processing unit to other processing units inside the specified chip to run, so as to avoid the occupation of the PCIe resources of the specified chip during the process of storing it in the interrupted task waiting queue. Specifically, as Figure 4 shown.
[0106] Figure 4 It is a schematic diagram of the computing kernel operation control method provided in this specification.
[0107] Combined Figure 4 As can be seen, when the scheduling management module determines that there is a computing kernel in the specified chip that has interrupted its operation and determines that the target processing unit is a computing unit, it can re-obtain the running status data of the specified chip. When it is determined according to the running status data that there is a computing core in an idle state among the processing units, the computing core in the idle state is used as the interrupted task processing unit, and the computing kernel that has interrupted its operation is used as the interrupted computing kernel. The interrupted computing kernel and the interrupted task processing unit are sent to the operation management module, so that the operation management module can schedule among the processing units of the execution chip after receiving the interrupted computing kernel and the interrupted task processing unit, so as to dispatch the interrupted computing kernel to the interrupted task processing unit to run.
[0108] As can be seen from the above, through the request management module of the above-mentioned computing kernel dynamic scheduling system, each computing kernel to be run belonging to different target tasks can be stored in different queues for management, and the scheduling management module determines the post coefficient of each computing kernel according to the degree of delay tolerance of each computing kernel in the queue. Furthermore, each computing kernel can be scheduled to run in the corresponding target computing kernel according to the post coefficient of each computing kernel, thereby improving the efficiency of running each target task through the specified chip.
[0109] For the sake of easy understanding, the following details the process of computing kernel dynamic scheduling through the above-mentioned computing kernel dynamic scheduling system, specifically as Figure 5 shown.
[0110] Figure 5 It is a schematic flowchart of a computing kernel dynamic scheduling method provided in this specification, including the following steps:
[0111] S501: Receive a task request, and obtain at least one computing kernel to be run through the request management module according to the task request. Among them, the computing kernel to be run is the computing service logic code segment for executing the target task corresponding to the task request.
[0112] S502: Receive, from the scheduling management module, for each computing kernel to be run, determine the post coefficient of the computing kernel to be run according to the execution characteristic data of the computing kernel to be run, and according to the post coefficient of each computing kernel to be run, and the running state data of the preset specified chip collected by the monitoring management module, determine the target computing kernel from each computing kernel to be run, and determine the target processing unit from each processing unit of the specified chip. Among them, the execution characteristic data is used to reflect the size of the delay allowed by the computing kernel to be run, including at least one of the task type of the target task to which the computing kernel to be run belongs, the waiting time of the computing kernel to be run, the estimated completion time of the computing kernel to be run, and the service performance index parameters corresponding to the computing kernel to be run. The running state data includes state data for reflecting the state of each processing unit in the specified chip and video memory resource occupancy data.
[0113] S503: Send the target computing kernel and the target processing unit to the operation management module, so that after receiving the target computing kernel and the target processing unit, the operation management module dispatches the target computing kernel to the target processing unit for running.
[0114] In this specification, the execution entity for implementing the dynamic scheduling method of computing kernels can refer to specified devices such as servers and embedded devices, or terminal devices such as desktop computers and laptop computers. For the convenience of description, hereinafter, only the case where the terminal device is the execution entity is taken as an example to illustrate the dynamic scheduling method of computing kernels provided in this specification.
[0115] In this specification, a terminal device can receive a task request, and the request management module of the computing kernel dynamic scheduling system running in the CPU deployed in the terminal device can, according to the received task request, obtain at least one computing kernel required to run when the target task corresponding to the received task request is executed, as the computing kernel to be run, where the computing kernel to be run is the computing service logic code segment for executing the target task corresponding to the task request.
[0116] Further, the terminal device can, through the scheduling management module of the computing kernel dynamic scheduling system, for each computing kernel to be run, determine the post coefficient of the computing kernel to be run according to the execution characteristic data of the computing kernel to be run, and according to the post coefficient of each computing kernel to be run, and the operation status data of the specified chip collected by the monitoring management module, determine the target computing kernel from the computing kernels to be run, and determine the target processing unit from each processing unit of the specified chip.
[0117] Among them, the above-mentioned execution characteristic data is used to reflect the allowable delay size of the computing kernel to be run, including: at least one of the task type of the target task to which the computing kernel to be run belongs, the waiting time of the computing kernel to be run, the estimated completion time of the computing kernel to be run, and the service performance index parameters corresponding to the computing kernel to be run.
[0118] The above-mentioned operation status data includes: status data for reflecting the status of each processing unit in the specified chip and video memory resource occupancy data.
[0119] When the terminal device obtains the target computing kernel and the target processing unit determined by the scheduling management module, it can send the target computing kernel and the target processing unit to the operation management module, so that after receiving the target computing kernel and the target processing unit, the operation management module dispatches the target computing kernel to the target processing unit for running.
[0120] As can be seen from the above, through the above-mentioned computing kernel dynamic scheduling method, each computing kernel to be run belonging to different target tasks can be stored in different queues for management. For each computing kernel in the queue, according to the degree of delay tolerance of each computing kernel, the postposition coefficient of each computing kernel is determined. Furthermore, each computing kernel can be scheduled to the corresponding target computing kernel according to the postposition coefficient of each computing kernel, thereby improving the efficiency of running each target task through the specified chip.
[0121] The above is the computing kernel dynamic scheduling system and method provided by one or more embodiments of this specification. Based on the same idea, this specification also provides a corresponding computing kernel dynamic scheduling device, as Figure 6 shown.
[0122] Figure 6 The schematic diagram of a computing kernel dynamic scheduling device provided by this specification specifically includes:
[0123] A first receiving module, configured to receive a task request, and obtain at least one computing kernel to be run through the request management module according to the task request, where the computing kernel to be run is a computing service logic code segment for executing the target task corresponding to the task request;
[0124] A second receiving module, configured to receive, from the scheduling management module, for each computing kernel to be run, according to the execution characteristic data of the computing kernel to be run, determine the postposition coefficient of the computing kernel to be run, and according to the postposition coefficient of each computing kernel to be run, and the running state data of the preset specified chip collected by the monitoring management module, determine the target computing kernel from each computing kernel to be run, and determine the target processing unit from each processing unit of the specified chip, where the execution characteristic data is used to reflect the size of the delay allowed by the computing kernel to be run, including: at least one of the task type of the target task to which the computing kernel to be run belongs, the waiting time of the computing kernel to be run, the estimated completion time of the computing kernel to be run, and the service performance index parameters corresponding to the computing kernel to be run, and the running state data includes: state data for reflecting the state of each processing unit in the specified chip and video memory resource occupancy data;
[0125] A scheduling module, configured to send the target computing kernel and the target processing unit to the running management module, so that after receiving the target computing kernel and the target processing unit, the running management module dispatches the target computing kernel to the target processing unit for running.
[0126] This specification also provides a computer-readable storage medium, which stores a computer program, and the computer program can be used to execute the aboveFigure 5 The computing kernel dynamic scheduling method shown
[0127] This specification also provides Figure 7 The schematic structural diagram of the electronic device shown. As Figure 7 described, at the hardware level, the electronic device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory. Of course, it may also include other hardware required for other services. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to implement the above Figure 3 The computing kernel dynamic scheduling method shown. Of course, in addition to the software implementation method, this specification does not exclude other implementation methods, such as logic devices or a combination of software and hardware, etc. That is to say, the execution subject of the following processing flow is not limited to each logic unit, and can also be hardware or logic devices.
[0128] In the 1990s, improvements to a technology could be clearly distinguished as either hardware improvements (e.g., improvements to circuit structures such as diodes, transistors, switches, etc.) or software improvements (improvements to method flows). However, with the development of technology, many method flow improvements today can be regarded as direct improvements to hardware circuit structures. Designers almost always obtain the corresponding hardware circuit structure by programming the improved method flow into the hardware circuit. Therefore, it cannot be said that an improvement to a method flow cannot be implemented using a hardware entity module. For example, a programmable logic device (PLD) (e.g., a field programmable gate array (FPGA)) is such an integrated circuit whose logical function is determined by a user's programming of the device. Designers can program themselves to "integrate" a digital system onto a single PLD, without having to ask a chip manufacturer to design and fabricate a dedicated integrated circuit chip. Moreover, nowadays, instead of manually fabricating integrated circuit chips, this programming is mostly implemented using "logic compiler" software, which is similar to the software compilers used in program development and writing. The original code before compilation also has to be written in a specific programming language, which is called a hardware description language (HDL), and there is not just one type of HDL, but many types, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used ones currently are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also be aware that by simply performing a little logical programming on the method flow using the above-mentioned several hardware description languages and programming it into an integrated circuit, it is easy to obtain the hardware circuit that implements the logical method flow.
[0129] The controller can be implemented in any suitable manner. For example, the controller can take the form of, for example, a microprocessor or a processor and a computer-readable medium storing computer-readable program code (such as software or firmware) executable by the (micro)processor, logic gates, switches, an application specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller. Examples of the controller include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art also know that in addition to implementing the controller in the form of pure computer-readable program code, it is entirely possible to make the controller implement the same function in the form of logic gates, switches, application specific integrated circuits, programmable logic controllers, and embedded microcontrollers by logically programming the method steps. Therefore, such a controller can be considered a hardware component, and the devices included therein for implementing various functions can also be regarded as structures within the hardware component. Or even, the devices for implementing various functions can be regarded as either software modules for implementing the method or structures within the hardware component.
[0130] The systems, devices, modules, or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer can be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or any combination of these devices.
[0131] For the convenience of description, when describing the above devices, they are described separately as various units according to their functions. Of course, when implementing this specification, the functions of each unit can be implemented in the same or multiple software and / or hardware.
[0132] Those skilled in the art should understand that the embodiments of this specification can be provided as a method, a system, or a computer program product. Therefore, this specification can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, this specification can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program code.
[0133] This specification is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the specification. It should be understood that each flow and / or block in the flowchart and / or block diagram, and combinations of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to produce a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices generate means for implementing the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or means for implementing the functions specified in one or more of the blocks.
[0134] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to operate in a specific manner, such that the instructions stored in the computer-readable memory produce a manufacture including instruction means that implement the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or means for implementing the functions specified in one or more of the blocks.
[0135] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or means for implementing the functions specified in one or more of the blocks.
[0136] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.
[0137] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory, such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of computer-readable media.
[0138] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.
[0139] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0140] It should be understood by those skilled in the art that the embodiments of this specification may be provided as methods, systems or computer program products. Therefore, this specification may take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware. Moreover, this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0141] This specification may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.
[0142] The various embodiments in this specification are described in a progressive manner. For the same or similar parts among the various embodiments, reference can be made to each other, and the key points of each embodiment are the differences from other embodiments. In particular, for the system embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and for the relevant parts, reference can be made to the partial description of the method embodiments.
[0143] The above description is only for the embodiments of this specification and is not intended to limit this specification. For those skilled in the art, various modifications and changes can be made to this specification. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this specification shall be included within the scope of the claims of this specification.
Claims
1. A computing kernel dynamic scheduling system, characterized in that: The computing kernel dynamic scheduling system includes: a request management module, a scheduling management module, a monitoring management module, and an operation management module; The request management module is used to obtain at least one computing kernel to be run, wherein the computing kernel to be run is a computing business logic code segment of a target task to be executed; The scheduling management module is used to determine the post-coefficient of each computing kernel to be run according to the execution characteristic data of the computing kernel to be run, and determine the target computing kernel from each computing kernel to be run according to the post-coefficient of each computing kernel to be run, and determine the target processing unit from each processing unit of the designated chip according to the operation status data of the preset designated chip collected by the monitoring management module, and send the target computing kernel and the target processing unit to the operation management module, wherein the execution characteristic data is used to reflect the delay size allowed by the computing kernel to be run, including: the task type of the target task to which the computing kernel to be run belongs, the waiting time of the computing kernel to be run, the estimated completion time of the computing kernel to be run, and the service performance index corresponding to the computing kernel to be run. At least one of the target parameters, the running status data includes: status data for reflecting the status of each processing unit in the specified chip and video memory resource occupancy data; the lower the post-coefficient of the computing kernel to be run, the higher the priority of the computing kernel to be run; the task type of the target task includes: real-time task and non-real-time task; the scheduling management module is used to determine the post-coefficient of the computing kernel to be run for each computing kernel to be run according to the service performance indicator parameters and estimated completion time corresponding to the computing kernel to be run when determining that the task type of the target task is a real-time task; the scheduling management module is used to determine the post-coefficient of the computing kernel to be run for each computing kernel to be run according to the waiting time and estimated completion time corresponding to the computing kernel to be run when determining that the task type of the target task is a non-real-time task; The operation management module is used to dispatch the target computing kernel to the target processing unit for execution after receiving the target computing kernel and the target processing unit.
2. The computing kernel dynamic scheduling system according to claim 1, characterized in that: The request management module is used to obtain a task request, determine a target task to be executed according to the task request, and determine at least one to-be-executed computing kernel that needs to be executed by a designated chip when executing the target task. For each to-be-executed computing kernel, according to the task type of the target task to which the to-be-executed computing kernel belongs, the to-be-executed computing kernel is stored in a waiting queue that matches the to-be-executed computing kernel, wherein the waiting queue includes: a real-time task waiting queue, a non-real-time task waiting queue, and an interrupt task waiting queue; The scheduling management module is used to determine the post-coefficient of each computing kernel to be run in each waiting queue.
3. The computing kernel dynamic scheduling system according to claim 1, characterized in that: The scheduling management module is used to determine that the computing kernel to be run is a target computing kernel according to the post-position coefficient of the computing kernel to be run, and when it is determined according to the running status data that there is no processing unit in an idle state among the processing units, determine the target processing unit from the processing units according to the service performance indicator parameters of the computing kernel being run by each processing unit, and send the target computing kernel and the target processing unit to the running management module; The operation management module is used to, upon receiving the target computing kernel and the target processing unit and determining that there is a computing kernel running in the target processing unit, stop the computing kernel running in the target processing unit and dispatch the target computing kernel to the target processing unit for execution.
4. The computing kernel dynamic scheduling system according to claim 3, characterized in that: The processing unit includes: a computing unit and a computing core in the computing unit; The scheduling management module is used for, when it is determined that the designated chip contains a computing core that interrupts the operation of the target processing unit and the target processing unit is a computing unit, reacquiring the operation status data of the designated chip, and when it is determined according to the operation status data that there is an idle computing core in each processing unit, using the idle computing core as an interrupt task processing unit, using the interrupted computing core as an interrupt computing core, and sending the interrupted computing core and the interrupted task processing unit to the operation management module; The operation management module is used for dispatching the interruption calculation kernel to the interruption task processing unit for execution after receiving the interruption calculation kernel and the interruption task processing unit.
5. The computing kernel dynamic scheduling system according to claim 1, characterized in that: The scheduling management module is used to determine that the computing kernel to be run is a target computing kernel according to the post-coefficient of the computing kernel to be run, and when it is determined according to the running status data that there is a processing unit in an idle state among the processing units, use the processing unit in the idle state as the target processing unit, and when it is determined that the number of floating-point operations required to run the computing kernel to be run is less than a preset threshold, determine at least part of the computing kernels to be run from the computing kernels to be run and the target computing kernel to form a target computing kernel set according to a target combination strategy, and send the target computing kernel set and the target processing unit to the running management module; The target combination strategy is determined from preset combination strategies, and the combination strategies include: at least one of a first combination strategy and a second combination strategy, wherein the first combination strategy is used to determine the target computing kernel set according to the video memory resource occupancy data of the specified chip and the video memory resource data required for the operation of each computing kernel to be run, and the second combination strategy is used to determine the target computing kernel set according to the number of floating-point operations required for the operation of each computing kernel to be run.
6. The computing kernel dynamic scheduling system according to claim 5, characterized in that: The processing unit includes: at least one of a computing unit and a computing core in the computing unit; The operation management module is used to determine whether the target processing unit is a computing unit after receiving the target computing kernel set. If so, for each computing kernel included in the target computing kernel set, determine a selected computing core from the computing cores of the target processing unit, and dispatch the computing kernel to the selected computing core for execution.
7. The computing kernel dynamic scheduling system according to claim 6, characterized in that: The operation management module is used to create a control thread in the designated computing core of the target processing unit, and through the control thread, determine a selected computing core from the computing cores of the target processing unit for each computing core included in the target computing core set, and dispatch the computing core to the selected computing core for execution.
8. The computing kernel dynamic scheduling system according to claim 7, characterized in that: The operation management module is used to monitor the computing cores of the target processing unit through the control thread, so that when it is determined that there is a computing core that has completed operation among the computing cores of the target processing unit, the computing core that has completed operation is used as an idle computing core, and the status data of the idle computing core and the video memory resource occupancy data after the idle computing core completes the operation are uploaded to the monitoring management module.
9. The computing kernel dynamic scheduling system according to claim 8, characterized in that: The monitoring management module is used to call a preset query interface to send an operation status data query request for the processing unit to be queried in the specified chip to the operation management module, and receive the status data of the processing unit to be queried and the video memory resource occupancy data returned by the operation management module according to the operation status data query request; as well as, Used to receive the status data of the idle computing core uploaded by the operation management module and the video memory resource occupancy data after the idle computing core completes the operation.
10. The computing kernel dynamic scheduling system according to claim 5, characterized in that: The scheduling management module is used to determine a target combination strategy from preset combination strategies based on the service performance indicator parameters corresponding to each computing kernel to be run, when it is determined that the number of floating-point operations required to run the computing kernel to be run is less than a preset threshold value, and according to the target combination strategy, determine at least some of the computing kernels to be run from the computing kernels to be run to form a target computing kernel set with the target computing kernel, and send the target computing kernel set and the target processing unit to the operation management module.
11. A computing kernel dynamic scheduling method, characterized in that: The method is applied to a computing kernel dynamic scheduling system, the computing kernel dynamic scheduling system includes: a request management module, a scheduling management module, a monitoring management module, and an operation management module, and the method includes: Receive a task request, and obtain at least one to-be-run computing kernel according to the task request through the request management module, wherein the to-be-run computing kernel is a computing business logic code segment that executes a target task corresponding to the task request; The scheduling management module receives, for each computing kernel to be run, the post-post coefficient of the computing kernel to be run according to the execution characteristic data of the computing kernel to be run, and determines the target computing kernel from each computing kernel to be run according to the post-post coefficient of each computing kernel to be run, and determines the target processing unit from each processing unit of the designated chip according to the operation status data of the preset designated chip collected by the monitoring management module, wherein the execution characteristic data is used to reflect the delay size allowed by the computing kernel to be run, including: the task type of the target task to which the computing kernel to be run belongs, the waiting time of the computing kernel to be run, the estimated completion time of the computing kernel to be run, and at least one of the service performance indicator parameters corresponding to the computing kernel to be run, and the operation status The state data includes: state data and video memory resource occupancy data for reflecting the state of each processing unit in the specified chip; the lower the post-coefficient of the computing kernel to be run, the higher the priority of the computing kernel to be run; the task type of the target task includes: real-time task and non-real-time task; the scheduling management module is used to determine the post-coefficient of the computing kernel to be run for each computing kernel to be run according to the service performance indicator parameters and estimated completion time corresponding to the computing kernel to be run when determining that the task type of the target task is a real-time task; the scheduling management module is used to determine the post-coefficient of the computing kernel to be run for each computing kernel to be run according to the waiting time and estimated completion time corresponding to the computing kernel to be run when determining that the task type of the target task is a non-real-time task; The target computing kernel and the target processing unit are sent to the operation management module, so that after receiving the target computing kernel and the target processing unit, the operation management module dispatches the target computing kernel to the target processing unit for execution.
Citation Information
Patent Citations
Task scheduling method and system and computer readable storage medium
CN115168000A
Data processing method and device based on core particles, medium and equipment
CN115829017A