Scheduling method and device of artificial intelligence processor, terminal and storage medium
By calculating the individual and joint processing latency of the AI processing modules on the terminal side, and selecting the scheme with the lowest power consumption, the problem of coordinated scheduling of AI processing modules is solved, and a balance between latency and energy consumption is achieved.
Patent Information
- Application Number
- CN202511059261.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-30
- Publication Date
- 2025-11-07
AI Technical Summary
In existing technologies, it is difficult for AI processing modules on the terminal side to coordinate and schedule, making it difficult to minimize task processing latency and energy consumption.
By calculating the individual processing latency and joint processing latency of the AI processing module, the lowest power consumption scheme is selected based on the timing requirements and priority order of the target task, and an execution scheme is generated to meet the latency requirements.
Minimize terminal power consumption while meeting latency requirements, significantly reducing energy consumption during AI task execution.
Smart Images

Figure CN120909726A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of artificial intelligence technology, specifically relating to a scheduling method, device, terminal, and storage medium for an artificial intelligence processor. Background Technology
[0002] Edge AI refers to the deployment of artificial intelligence algorithms for training or inference directly on terminal devices, rather than relying on cloud servers for data processing. By bringing computing power down to the device, it enables data collection, analysis, and decision-making to be completed locally, marking the evolution of AI technology from "cloud-centralized" to "terminal-distributed".
[0003] Currently, AI deployment solutions on the terminal side require at least one dedicated AI processing module, such as an NPU or GPU. Additionally, to apply AI in wireless connectivity, an AI processing module, known as Modem AI, also needs to be deployed in the modem. The processing speed and power consumption of AI processing modules deployed in different locations vary. Existing technologies typically schedule AI processing models based on task type; for example, an NPU / GPU might handle computationally intensive tasks, while a Modem AI module handles communication-related tasks. This makes coordinated scheduling between different AI processing modules difficult, hindering the minimization of task processing latency and energy consumption. Summary of the Invention
[0004] The purpose of this application is to provide a scheduling method, apparatus, terminal, and storage medium for an artificial intelligence processor.
[0005] To achieve the above objectives, the first aspect of this application provides a scheduling method for an artificial intelligence processor, applied to a terminal, wherein the terminal is equipped with a plurality of AI processing modules and a memory that are interconnected by signals, wherein the plurality of AI processing modules include at least one first AI processing module and at least one second AI processing module deployed in a modem.
[0006] The scheduling method includes:
[0007] Obtain target task information, including the type, size, and time delay requirements of the target task;
[0008] Based on the target task information, task data volume information is obtained, which includes AI model data volume, task input data volume, task output data volume, and task computation and inference volume.
[0009] obtain a single processing latency and / or a joint processing latency of the AI processing module based on the task data amount information, the first transmission rate and the second transmission rate, and determine whether the latency requirement is met, and if so, generate an execution scheme of the target task based on the single processing latency and / or the joint processing latency that meets the latency requirement, wherein the first transmission rate is a data transmission rate between the first AI processing module and the second AI processing module, and the second transmission rate is a data transmission rate between the AI processing module and the memory.
[0010] In one or more embodiments, the step of obtaining a single processing latency and / or a joint processing latency of the AI processing module based on the task data amount information, the first transmission rate and the second transmission rate, and determining whether the latency requirement is met, and if so, generating an execution scheme of the target task based on the single processing latency and / or the joint processing latency that meets the latency requirement, comprises:
[0011] calculating a single processing latency of the second AI processing module based on the task data amount information, the first transmission rate and the second transmission rate, and determining whether the latency requirement is met;
[0012] if not, calculating a single processing latency of the first AI processing module based on the task data amount information, the first transmission rate and the second transmission rate, and determining whether the latency requirement is met;
[0013] if not, calculating the joint processing latency under different task partitioning ratios based on the task data amount information, the first transmission rate and the second transmission rate, and determining whether the latency requirement is met;
[0014] if so, the first AI processing module and the second AI processing module jointly process the target task.
[0015] In one or more embodiments, the method further comprises:
[0016] if the second AI processing module meets the latency requirement, the second AI processing module processes the target task alone; or
[0017] if the first AI processing module meets the latency requirement, the first AI processing module processes the target task alone.
[0018] In one or more embodiments, the step of calculating the joint processing latency under different task partitioning ratios and determining whether the latency requirement is met comprises:
[0019] sequentially traversing different task segmentation ratios of the target task, and the traversal direction is to gradually increase the task processing ratio of the first AI processing module by a fixed step size, and when each task segmentation ratio is traversed, the joint processing delay under the task segmentation ratio is calculated, and it is determined whether the delay requirement is met;
[0020] If not, the next task segmentation ratio is continuously traversed until the task segmentation ratio meeting the delay requirement is obtained or the traversal is completed.
[0021] In one or more embodiments, based on the task data amount information, the first transmission rate and the second transmission rate, the individual processing delay and / or the joint processing delay of the AI processing module are obtained, and it is determined whether the delay requirement is met. If yes, the step of generating the execution scheme of the target task based on the individual processing delay and / or the joint processing delay meeting the delay requirement comprises:
[0022] Based on the task data amount information, the first transmission rate and the second transmission rate, the individual processing delay of each AI processing module and the joint processing delay under different task segmentation ratios are calculated and collected into a queue;
[0023] The processing delay meeting the delay requirement is filtered from the queue, and is sorted according to a preset priority order, the preset priority order comprises that the priority of the individual processing delay of the second AI processing module is higher than that of the individual processing delay of the first AI processing module, and the priority of the joint processing delay under different task segmentation ratios is higher, and the higher the task processing ratio of the second AI processing module is, the higher the priority is;
[0024] Based on the processing delay with the highest priority, the execution scheme of the target task is generated.
[0025] In one or more embodiments, the calculation method of the individual processing delay of the second AI processing module comprises:
[0026] Based on the AI model data amount and the second transmission rate of the second AI processing module, the model loading delay of the second AI processing module is obtained;
[0027] Based on the task calculation inference amount and the calculation rate of the second AI processing module, the inference delay of the second AI processing module is obtained;
[0028] Based on the model loading delay and the inference delay of the second AI processing module, the individual processing delay of the second AI processing module is obtained.
[0029] In one or more embodiments, the method for calculating the single processing latency of the first AI processing module comprises:
[0030] obtaining a model loading latency of the first AI processing module based on the AI model data volume and the second transmission rate of the first AI processing module;
[0031] obtaining an inference latency of the first AI processing module based on the task computation inference volume and the computation rate of the first AI processing module;
[0032] obtaining an input latency of the first AI processing module based on the task input data volume and the first transmission rate;
[0033] obtaining an output latency of the first AI processing module based on the task output data volume and the first transmission rate;
[0034] obtaining the single processing latency of the first AI processing module based on the model loading latency, the inference latency, the input latency and the output latency of the first AI processing module.
[0035] In one or more embodiments, the method for calculating the joint processing latency comprises:
[0036] dividing the target task into a first task and a second task according to a task division ratio, wherein the first task is executed by the first AI processing module and the second task is executed by the second AI processing module;
[0037] obtaining inference latencies of the first AI processing module and the second AI processing module based on the task computation inference volumes of the first task and the second task and the computation rates of the first AI processing module and the second AI processing module;
[0038] obtaining model loading latencies of the first AI processing module and the second AI processing module based on the AI model data volumes corresponding to the first task and the second task and the second transmission rate;
[0039] calculating first interaction latency and second interaction latency for different processing sequences of the first task and the second task, wherein the first interaction latency is the interaction latency when the first task is processed before the second task, and the second interaction latency is the interaction latency when the second task is processed before the first task;
[0040] obtaining first joint processing latency and second joint processing latency under the task division ratio based on the inference latencies, the model loading latencies of the first AI processing module and the second AI processing module, and the first interaction latency and the second interaction latency.
[0041] In one or more embodiments, the method for calculating the first interaction delay comprises:
[0042] obtaining the first interaction delay based on the task input data amount of the first task and the first transmission rate; or
[0043] The method for calculating the second interaction delay comprises:
[0044] obtaining the second interaction delay based on the task output data amount of the first task and the first transmission rate.
[0045] In one or more embodiments, if there is no single processing delay and / or joint processing delay that meets the delay requirement, further comprising:
[0046] sorting the single processing delay and / or joint processing delay according to size, and generating an execution scheme of the target task based on the smallest processing delay, and feeding back the execution result to the network side; or
[0047] abandoning to execute the target task, and feeding back execution failure information to the network side; or
[0048] sorting the single processing delay and / or joint processing delay according to size, and feeding back the smallest processing delay to the network side for the network side to determine whether to execute the target task.
[0049] To achieve the above object, the second aspect of the present application provides a scheduling device of an artificial intelligence processor, applied to a terminal, wherein the terminal is internally deployed with a plurality of AI processing modules and a memory which are signal connected with each other, wherein the plurality of AI processing modules comprise at least one first AI processing module and at least one second AI processing module deployed in a modem;
[0050] The scheduling device comprises:
[0051] a task acquisition module, configured to acquire target task information, wherein the target task information comprises the type, size and delay requirement of the target task;
[0052] a task amount acquisition module, configured to acquire task data amount information based on the target task information, wherein the task data amount information comprises AI model data amount, task input data amount, task output data amount and task calculation and inference amount;
[0053] The computing determining module is configured to obtain a single processing delay and / or a joint processing delay of the AI processing module based on the task data amount information, the first transmission rate and the second transmission rate, and determine whether the delay requirement is met, and if so, generate an execution scheme of the target task based on the single processing delay and / or the joint processing delay that meets the delay requirement, wherein the first transmission rate is a data transmission rate between the first AI processing module and the second AI processing module, and the second transmission rate is a data transmission rate between the AI processing module and the memory.
[0054] To achieve the above object, the third aspect of the present application provides a terminal, comprising:
[0055] at least one processor; and
[0056] a memory storing instructions that, when executed by the at least one processor, cause the at least one processor to perform the scheduling method of the artificial intelligence processor as described in any of the above embodiments.
[0057] To achieve the above object, the fourth aspect of the present application provides a machine-readable storage medium storing executable instructions that, when executed, cause the machine to perform the scheduling method of the artificial intelligence processor as described in any of the above embodiments.
[0058] Compared with the prior art, the present application has the following beneficial effects:
[0059] The scheduling method of the present application calculates the single processing delay and the joint processing delay of the AI processing module, selects the scheme with the lowest power consumption that can meet the delay requirement based on the timing requirement and the priority order of the target task, and generates the execution scheme of the target task, which can minimize the power consumption on the premise of meeting the delay requirement, and significantly reduce the power consumption of the terminal when executing the AI task. BRIEF DESCRIPTION OF DRAWINGS
[0060] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed in the embodiment or prior art description will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments described in the present application, and those skilled in the art can also obtain other drawings according to these drawings without creative labor.
[0061] Figure 1 is a schematic diagram of an embodiment of a terminal AI deployment scheme;
[0062] Figure 2 is a flowchart of an embodiment of the scheduling method of the artificial intelligence processor of the present application;
[0063] Figure 3 is a flowchart of an embodiment of a method for calculating the single processing latency of the second AI processing module of the present application;
[0064] Figure 4 is a flowchart of an embodiment of a method for calculating the single processing latency of the first AI processing module of the present application;
[0065] Figure 5 is a flowchart of an embodiment of a method for calculating the joint processing latency of the present application;
[0066] Figure 6 is Figure 2 is a flowchart of an embodiment corresponding to S300 in the method;
[0067] Figure 7 is Figure 6 is a flowchart of an embodiment corresponding to S303a in the method;
[0068] Figure 8 is Figure 2 is a flowchart of another embodiment corresponding to S300 in the method;
[0069] Figure 9 is a structural diagram of an embodiment of the scheduling device of the artificial intelligence processor of the present application;
[0070] Figure 10 is a structural diagram of an embodiment of the terminal of the present application. DETAILED DESCRIPTION
[0071] In order to enable those skilled in the art to better understand the technical solutions in the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by those skilled in the art without creative labor should fall within the scope of protection of the present disclosure.
[0072] Please refer to Figure 1 , Figure 1 is a schematic diagram of an embodiment of the terminal AI deployment scheme. As shown in Figure 1As shown in the existing terminal-side artificial intelligence (AI) deployment scheme, there are generally at least two AI processing modules, wherein the first AI processing module 11, as a dedicated AI processing module, can adopt NPU, GPU, etc., and is used for end-side application processing, such as picture beautification, voice assistant, etc., and can also be used for processing communication connection AI tasks; in addition, in order to cope with the application of AI in wireless connection, a second AI processing module 13 is generally also deployed in the modem 12 for processing communication connection AI tasks.
[0073] The processing speeds of AI processing modules deployed in different positions are different, and the power consumptions are also different; in the prior art, the scheduling of corresponding AI processing models is generally based on the type of task, for example, NPU / GPU is adopted to process computationally intensive tasks, and the Modem AI module is adopted to process communication-related tasks, and it is difficult to cooperatively schedule different AI processing modules, and it is difficult to minimize the task processing delay and energy consumption.
[0074] In order to solve the above problems, the applicant has developed a new artificial intelligence processor scheduling method, which can cooperatively schedule different AI processing modules, minimize the energy consumption of the terminal under the premise of meeting the delay requirement, and balance the delay and energy consumption.
[0075] Specifically, please refer to Figure 2 , Figure 2 is a flowchart of an embodiment of the artificial intelligence processor scheduling method of the applicant.
[0076] In this embodiment, the scheduling scheme is applied to a terminal as shown in Figure 1 , specifically, the terminal is internally deployed with a first AI processing module, a second AI processing module and a memory connected to each other, wherein the second AI processing module is deployed in the modem.
[0077] As shown in Figure 2 , the scheduling scheme includes:
[0078] S100, obtaining target task information.
[0079] The target task information includes the type, size and delay requirement of the target task.
[0080] It can be understood that the above target task information can be obtained based on the target task issued by the network side.
[0081] S200, obtaining task data volume information based on the target task information.
[0082] The task data volume information includes AI model data volume, task input data volume, task output data volume and task calculation and inference volume.
[0083] Since the type of the target task is known, the AI model for executing the target task is known, and the AI model data volume can be obtained; accordingly, the input data volume, the output data volume, and the computation and inference volume of the target task can be obtained based on the target task information.
[0084] S300, based on the task data volume information, the first transmission rate, and the second transmission rate, obtaining the individual processing delay and / or the joint processing delay of the AI processing module, and determining whether the delay requirement is met. If yes, generating an execution scheme of the target task based on the individual processing delay and / or the joint processing delay that meets the delay requirement.
[0085] The first transmission rate is the data transmission rate between the first AI processing module and the second AI processing module, and the second transmission rate is the data transmission rate between the AI processing module and the memory.
[0086] The scheduling method of the present application calculates the individual processing delay of the AI processing module in processing the target task, and the joint processing delay of the first AI processing module and the second AI processing module in jointly processing the target task, and determines whether the delay requirement is met. Based on the processing delay that meets the delay requirement, the corresponding AI processing module is selected to execute the target task, thereby ensuring that the execution scheme of the target task can meet the delay requirement.
[0087] Since the power consumption and the computation rate of the second AI processing module deployed in the scheduling modem are lower than those of the first AI processing module, in order to minimize the power consumption, the scheduling scheme of the present application prioritizes the use of the second AI processing module to process the target task individually under the premise of meeting the delay requirement, or maximizes the task processing proportion of the second AI processing module when only the joint processing delay meets the delay requirement.
[0088] The calculation method of the individual processing delay and the joint processing delay of the present application will be described in detail below.
[0089] Please refer to Figure 3 , Figure 3 is a flowchart of an embodiment of the calculation method of the individual processing delay of the second AI processing module of the present application.
[0090] As Figure 3 shown, the calculation method comprises:
[0091] S10a, based on the AI model data volume and the second transmission rate of the second AI processing module, obtaining the model loading delay of the second AI processing module.
[0092] Since the AI model data amount is known, and the transmission rate between the second AI processing module and the memory is known, the time for the second AI processing module to load the AI model from the memory, i.e., the model loading latency t21, can be calculated.
[0093] S20a, based on the task computation inference amount and the computation rate of the second AI processing module, the inference latency of the second AI processing module is obtained.
[0094] Since the task computation inference amount of the target task is known, and the computation rate of the second AI processing module is known, the time for the second AI processing module to specifically execute the target task, i.e., the inference latency t22, can be calculated.
[0095] S30a, based on the model loading latency and the inference latency of the second AI processing module, the single processing latency of the second AI processing module is obtained.
[0096] Since the second AI processing module is deployed in the modem, there is no additional input time of task data and output time of computation result, therefore, the single processing latency of the second AI processing module, i.e., the sum of its model loading latency t21 and inference latency t22, is equal to t21+t22.
[0097] Please refer to Figure 4 , Figure 4 is a flowchart of an embodiment of the calculation method of the single processing latency of the first AI processing module.
[0098] As shown in Figure 4 , the calculation method comprises:
[0099] S10b, based on the AI model data amount and the second transmission rate of the first AI processing module, the model loading latency of the first AI processing module is obtained.
[0100] The same as S10a, the time for the first AI processing module to load the AI model from the memory, i.e., the model loading latency t11, can be calculated first.
[0101] S20b, based on the task computation inference amount and the computation rate of the first AI processing module, the inference latency of the first AI processing module is obtained.
[0102] The same as S20b, the time for the first AI processing module to specifically execute the target task, i.e., the inference latency t12, can be calculated.
[0103] S30b, based on the task input data amount and the first transmission rate, the input latency of the first AI processing module is obtained.
[0104] Since the first AI processing module also needs to obtain the input data of the target task from the modem, an input delay will be generated, which is related to the task input data amount and the first transmission rate. Therefore, based on the task input data amount and the first transmission rate, the input delay t13 can be obtained.
[0105] S40b, based on the task output data amount and the first transmission rate, obtaining the output delay of the first AI processing module.
[0106] Since the first AI processing module needs to output the calculation result to the modem after the target task is executed, and the network side is fed back by the modem, an output delay will be generated, which is related to the task output data amount and the first transmission rate. Therefore, based on the task output data amount and the first transmission rate, the output delay t14 can be obtained.
[0107] S50b, based on the model loading delay, the inference delay, the input delay and the output delay of the first AI processing module, obtaining the single processing delay of the first AI processing module.
[0108] Specifically, the single processing delay is the sum of the above delays, which is equal to t11+t12+t13+t14+t15.
[0109] Please refer to Figure 5 , Figure 5 is a flowchart of an embodiment of the joint processing delay calculation method of the present application.
[0110] As Figure 5 shown, the calculation method comprises:
[0111] S10c, the target task is divided into a first task and a second task according to the task division ratio.
[0112] It can be understood that the joint processing delay is related to the task amount processed by the first AI processing module and the second AI processing module, i.e. the task division ratio. Therefore, first, the target task can be divided based on the task division ratio, wherein the first task is executed by the first AI processing module, and the second task is executed by the second AI processing module.
[0113] In one embodiment, the above S300 can only calculate the joint processing delay under the preset task division ratio, and in other embodiments, the joint processing delay under multiple different task division ratios can also be calculated, which can all achieve the effect of the present embodiment.
[0114] S20c, based on the task calculation inference amount of the first task and the second task and the calculation rate of the first AI processing module and the second AI processing module, obtaining the inference delay of the first AI processing module and the second AI processing module.
[0115] The task computation inference amount of the two tasks after the target task segmentation is known, and therefore the time for the first AI processing module and the second AI processing module to respectively process the respective tasks, i.e., the inference time delay t12a and t22a, can be calculated.
[0116] S30c, based on the AI model data amount corresponding to the first task and the second task and the second transmission rate, obtaining the model loading time delay of the first AI processing module and the second AI processing module.
[0117] Since the first task and the second task are known, the AI model data amount corresponding thereto can be obtained, and further, based on the second transmission rate, the time for the first AI processing module and the second AI processing module to respectively load the AI model of the respective tasks, i.e., the model loading time delay t11a and t21a, can be calculated.
[0118] S40c, for different processing sequences of the first task and the second task, the first interaction time delay and the second interaction time delay are calculated.
[0119] The first interaction time delay is the interaction time delay when the first task is processed before the second task, and the second interaction time delay is the interaction time delay when the second task is processed before the first task.
[0120] The exchange time delay between the first AI processing module and the second AI processing module is related to the task processing sequence. When the first task is a front part task, the task input data of the first task needs to be transmitted to the first AI processing module by the modem first, and at this time, the first interaction time delay t31 is generated.
[0121] The calculation method of the first interaction time delay includes: based on the task input data amount of the first task and the first transmission rate, obtaining the first interaction time delay.
[0122] When the first task is a rear part task, the task output data of the first task needs to be fed back to the modem, and at this time, the second interaction time delay t32 is generated.
[0123] The calculation method of the second interaction time delay includes: based on the task output data amount of the first task and the first transmission rate, obtaining the second interaction time delay.
[0124] S50c, based on the inference time delay, the model loading time delay, and the first interaction time delay and the second interaction time delay of the first AI processing module and the second AI processing module, obtaining the first joint processing time delay and the second joint processing time delay under the task segmentation ratio.
[0125] For the same task split ratio, two joint processing latencies can be obtained due to different processing sequences of the first task and the second task, wherein the first joint processing latency, i.e., the sum of the inference latency, the model loading latency and the first interaction latency of the first AI processing module and the second AI processing module, is equal to t12a+t22a+t11a+t21a+t31;
[0126] The second joint processing latency, i.e., the sum of the inference latency, the model loading latency and the second interaction latency of the first AI processing module and the second AI processing module, is equal to t12a+t22a+t11a+t21a+t32.
[0127] The generation method of the execution scheme of the target task will be described in detail below.
[0128] Please refer to Figure 6 , Figure 6 is Figure 2 the flowchart of an embodiment corresponding to S300 in
[0129] As Figure 6 shown, the generation method of the execution scheme of the target task includes:
[0130] S301a, based on the task data amount information, the first transmission rate and the second transmission rate, the single processing latency of the second AI processing module is calculated, and it is judged whether the latency requirement is met.
[0131] If not, then:
[0132] S302a, based on the task data amount information, the first transmission rate and the second transmission rate, the single processing latency of the first AI processing module is calculated, and it is judged whether the latency requirement is met.
[0133] If not, then:
[0134] S303a, based on the task data amount information, the first transmission rate and the second transmission rate, the joint processing latency under different task split ratios is calculated, and it is judged whether the latency requirement is met.
[0135] If yes, then:
[0136] S304a, the first AI processing module and the second AI processing module jointly process the target task.
[0137] If the second AI processing module meets the latency requirement, then:
[0138] S302b, the second AI processing module processes the target task alone.
[0139] If the first AI processing module meets the latency requirement, then:
[0140] S303b, the first AI processing module processes the target task alone.
[0141] Since the power consumption of the second AI processing module is significantly lower than that of the first AI processing module, and the model partitioning involved in the joint processing process will cause additional complexity, in the embodiment, the scheme of sequentially calculating the single processing delay of the second AI processing module, the single processing delay of the first AI processing module, and the joint processing delay is adopted, and whether the delay requirement is met is judged after each calculation, so as to realize the purpose of minimizing the power consumption under the premise of guaranteeing the delay requirement.
[0142] The specific steps of S303a described above will be described in detail below. Please refer to Figure 7 , Figure 7 is Figure 6 the flowchart of an embodiment corresponding to S303a in
[0143] As shown in Figure 7 , the steps of calculating the joint processing delay under different task partitioning ratios and judging whether the delay requirement is met include:
[0144] S3031, different task partitioning ratios of the target task are sequentially traversed, and the traversal direction is to gradually increase the task processing ratio of the first AI processing module by a fixed step size. When a task partitioning ratio is traversed, the joint processing delay under the task partitioning ratio is calculated, and whether the delay requirement is met is judged.
[0145] If not, then:
[0146] S3032, the next task partitioning ratio is continuously traversed until the task partitioning ratio meeting the delay requirement is obtained or the traversal is completed.
[0147] Since the power consumption of the second AI processing module is significantly lower than that of the first AI processing module, in the embodiment, different task partitioning ratios are sequentially traversed along the direction of gradually increasing the task processing ratio of the first AI processing module, and the joint processing delay under each task partitioning ratio is calculated, which helps to increase the task processing ratio of the second AI processing module and minimize the power consumption under the premise of meeting the delay requirement.
[0148] It should be noted that for different task orders under each task partitioning ratio, there are a first joint processing delay and a second joint processing delay. When judging whether the delay requirement is met, one of the joint processing delays meets the requirement. At this time, the execution scheme of the target task can be generated based on the task order and the task partitioning ratio corresponding to the joint processing delay.
[0149] And if both joint processing delays under the task partitioning ratio meet the delay requirement, the scheme with the lowest delay can be selected to generate the execution scheme of the target task.
[0150] Based on the scheme of the above embodiment, the power consumption is minimized under the premise of meeting the delay requirement by sequentially calculating the individual processing delay and the joint processing delay. In another embodiment, the processing delays of all schemes can also be calculated synchronously, and the optimal one is selected.
[0151] Specifically, referring to Figure 8 , Figure 8 is Figure 2 the flowchart of another embodiment corresponding to S300.
[0152] As Figure 8 shown, the method for generating the execution scheme of the target task includes:
[0153] S301c, based on the task data amount information, the first transmission rate and the second transmission rate, the individual processing delay of each AI processing module and the joint processing delay under different task segmentation ratios are calculated and collected into a queue.
[0154] S302c, the processing delay that can meet the delay requirement is screened from the queue, and is sorted according to a preset priority order.
[0155] The preset priority order includes that the priority of the individual processing delay of the second AI processing module is higher than that of the first AI processing module, and the priority of the joint processing delay is higher than that of the individual processing delay. Among the joint processing delays under different task segmentation ratios, the higher the task processing ratio of the second AI processing module is, the higher the priority is.
[0156] S303c, based on the processing delay with the highest priority, the execution scheme of the target task is generated.
[0157] In this embodiment, the individual processing delays and the joint processing delays of all schemes are calculated synchronously, and the scheme with the lowest power consumption that can meet the delay requirement is selected based on the preset priority order, and the execution scheme of the target task is generated, which can also achieve the purpose of minimizing the power consumption under the premise of meeting the delay requirement.
[0158] In S300 described above Figure 2 , if there is no individual processing delay and joint processing delay that meets the delay requirement, the terminal can choose to continue executing the target task or directly give up executing the target task.
[0159] Specifically, the individual processing delay and / or the joint processing delay can be sorted according to the size, and based on the smallest processing delay, the execution scheme of the target task is generated, and the execution result is fed back to the network side, and whether to adopt the execution result is determined by the network side.
[0160] Or, the target task can also be abandoned, and execution failure information is fed back to the network side, and the execution failure information can include the reason for the execution failure, i.e., the delay requirement cannot be met;
[0161] Or, the unit processing delay and / or joint processing delay can also be sorted according to size, and the smallest processing delay is fed back to the network side for the network side to judge whether to execute the target task.
[0162] It should be noted that the scheduling method of each embodiment described above is applied in a terminal in which a first AI processing module and a second AI processing module are deployed, and in other embodiments, the scheduling method of the present application can also be applied in a terminal in which multiple first AI processing modules and / or multiple second AI processing modules are deployed. For example, two first AI processing modules and one second AI processing module deployed in a modem can be deployed in the terminal, and the effects of the present embodiment can also be achieved.
[0163] The present application also provides a scheduling device of an artificial intelligence processor. Please refer to Figure 9 , Figure 9 is a structural schematic diagram of an embodiment of the scheduling device of the artificial intelligence processor of the present application.
[0164] The scheduling device is applied in a terminal, and the terminal is internally deployed with a plurality of AI processing modules and a memory which are signal-connected with each other, wherein the plurality of AI processing modules include at least one first AI processing module and at least one second AI processing module deployed in a modem.
[0165] As shown in Figure 9 , the scheduling device includes a task acquisition module 21, a task amount acquisition module 22, and a calculation and judgment module 23.
[0166] The task acquisition module 21 is configured to acquire target task information, and the target task information includes the type, size, and delay requirement of the target task.
[0167] The task amount acquisition module 22 is configured to acquire task data amount information based on the target task information, and the task data amount information includes AI model data amount, task input data amount, task output data amount, and task calculation and inference amount.
[0168] The calculation and judgment module 23 is configured to obtain the individual processing delay and / or joint processing delay of the AI processing module based on the task data amount information, a first transmission rate, and a second transmission rate, and judge whether the delay requirement is met. If yes, an execution scheme of the target task is generated based on the individual processing delay and / or joint processing delay that meets the delay requirement, wherein the first transmission rate is the data transmission rate between the first AI processing module and the second AI processing module, and the second transmission rate is the data transmission rate between the AI processing module and the memory.
[0169] As described above with reference to Figures 1 to 8 , the scheduling method of the artificial intelligence processor according to the embodiments of the present specification is described. The details mentioned in the above description of the method embodiments are also applicable to the scheduling apparatus of the artificial intelligence processor of the embodiments of the present specification. The above scheduling apparatus of the artificial intelligence processor can be implemented in hardware, or in software, or in a combination of hardware and software.
[0170] The present application also provides a terminal, please refer to Figure 10 , Figure 10 is a structural schematic diagram of an embodiment of the terminal of the present application. As Figure 10 indicated, the terminal 30 can include at least one processor 31, a memory 32 (for example, a non-volatile memory), a memory 33, a communication interface 34, at least one first AI processing module 36, at least one second AI processing module 38 deployed in a modem 37, and the at least one processor 31, the at least one first AI processing module 35, the at least one second AI processing module 37, the memory 32, the memory 33 and the communication interface 34 are connected together via an internal bus 35. The at least one processor 31 executes at least one computer readable instruction stored or encoded in the memory 32.
[0171] It should be understood that the computer executable instructions stored in the memory 32, when executed, cause the at least one processor 31 to perform various operations and functions described above in conjunction with Figures 1-8 the various embodiments of the present specification.
[0172] In the embodiments of the present specification, the terminal 30 can include, but is not limited to, a personal computer, a server computer, a workstation, a desktop computer, a laptop computer, a notebook computer, a mobile electronic device, a smart phone, a tablet computer, a cellular phone, a personal digital assistant (PDA), a handheld device, a messaging device, a wearable electronic device, a consumer electronic device, and the like.
[0173] According to one embodiment, a program product such as a machine readable medium is provided. The machine readable medium can have instructions (i.e., the above-mentioned elements implemented in software) that, when executed by a machine, cause the machine to perform various operations and functions described above in conjunction with Figures 1-8 the various embodiments of the present specification. Specifically, a system or apparatus equipped with a readable storage medium on which software program codes implementing the functions of any of the above-mentioned embodiments are stored, and causing the computer or processor of the system or apparatus to read out and execute the instructions stored in the readable storage medium can be provided.
[0174] In this case, the program code itself read from the readable medium can implement the functions of any of the above-described embodiments, and the machine-readable code and the readable storage medium storing the machine-readable code form part of the present specification.
[0175] Embodiments of the readable storage medium include floppy diskettes, hard disks, optical disks (e.g., CD-ROMs, CD-Rs, CD-RWs, DVD-ROMs, DVD-RAMs, DVD-RWs, DVD-RWs), magnetic tapes, non-volatile memory cards, and ROMs. Alternatively, the program code can be downloaded over a communications network from a server computer or from a cloud.
[0176] Those skilled in the art will understand that various embodiments disclosed above can be modified and altered without departing from the spirit of the invention. Therefore, the scope of protection of the present specification should be defined by the appended claims.
[0177] It should be noted that not all steps and units in the above-described flowcharts and system block diagrams are necessary, and some steps or units can be omitted according to actual needs. The execution order of each step is not fixed and can be determined as needed. The device structure described in each of the above embodiments can be a physical structure or a logical structure, i.e., some units can be implemented by the same physical client, or some units can be implemented by multiple physical clients, or some units can be implemented by some components in multiple independent devices.
[0178] In each of the above embodiments, a hardware unit or module can be implemented mechanically or with electronic components. For example, a hardware unit, module or processor can include dedicated circuitry or logic (e.g., an application-specific integrated circuit or ASIC) to perform the corresponding operations. A hardware unit or processor can also include programmable logic or circuitry (e.g., a general-purpose processor or other programmable processor) that can be temporarily reconfigured via software to perform the corresponding operations. A specific implementation of a hardware unit or processor (mechanical or dedicated circuitry, or programmable logic or circuitry) can be determined by a cost and time consideration, among other factors.
[0179] The specific implementations described above with reference to the attached drawings illustrate example embodiments in which the techniques can be carried out, but are not meant to be limiting in scope. The term "exemplary," as used in this specification, means "serving as an example, instance, or illustration," and not "preferred" over other embodiments. The detailed description includes specific details for the purpose of providing a thorough understanding of the techniques described herein. However, it will be apparent to those skilled in the art that these techniques can be practiced without these specific details. In some instances, well-known structures and devices are shown in block diagram form in order to avoid obscuring the concepts of the described embodiments.
[0180] The foregoing description of the present disclosure has been presented for the purposes of conduction and enabling those of ordinary skill in the art to make and use the present disclosure. Modifications to embodiments of the present disclosure implementing the principles of the present disclosure can occur to those skilled in the art with the benefit of the present disclosure. Therefore, what has been described above is merely illustrative of the principles of the present disclosure, and the present disclosure is to be given a broad interpretation.
Claims
1. A method of scheduling an artificial intelligence processor, the method comprising: The application is applied to a terminal, wherein a plurality of AI processing modules and a memory are arranged in the terminal and are connected with each other, wherein the plurality of AI processing modules comprise at least one first AI processing module and at least one second AI processing module arranged in a modem; The scheduling method comprises: obtaining target task information, wherein the target task information comprises the type, size and time delay requirement of a target task; obtaining task data volume information based on the target task information, wherein the task data volume information comprises AI model data volume, task input data volume, task output data volume and task calculation and inference volume; obtaining the individual processing time delay and / or joint processing time delay of the AI processing module based on the task data volume information, first transmission rate and second transmission rate, and determining whether the time delay requirement is met, and if yes, generating an execution scheme of the target task based on the individual processing time delay and / or joint processing time delay meeting the time delay requirement, wherein the first transmission rate is the data transmission rate between the first AI processing module and the second AI processing module, and the second transmission rate is the data transmission rate between the AI processing module and the memory.
2. The scheduling method of claim 1, wherein, The step of obtaining the individual processing time delay and / or joint processing time delay of the AI processing module based on the task data volume information, first transmission rate and second transmission rate, and determining whether the time delay requirement is met, and if yes, generating an execution scheme of the target task based on the individual processing time delay and / or joint processing time delay meeting the time delay requirement comprises: calculating the individual processing time delay of the second AI processing module based on the task data volume information, first transmission rate and second transmission rate, and determining whether the time delay requirement is met; if not, calculating the individual processing time delay of the first AI processing module based on the task data volume information, first transmission rate and second transmission rate, and determining whether the time delay requirement is met; if not, calculating the joint processing time delay under different task segmentation ratios based on the task data volume information, first transmission rate and second transmission rate, and determining whether the time delay requirement is met; if yes, the first AI processing module and the second AI processing module jointly process the target task.
3. The scheduling method of claim 2, wherein, Further comprising: if the second AI processing module meets the time delay requirement, the second AI processing module processes the target task individually; or, if the first AI processing module meets the time delay requirement, the first AI processing module processes the target task individually.
4. The scheduling method of claim 2, wherein, The step of calculating the joint processing time delay under different task segmentation ratios and determining whether the time delay requirement is met comprises: sequentially traversing different task segmentation ratios of the target task, and the traversal direction is to gradually increase the task processing ratio of the first AI processing module by a fixed step, and when each task segmentation ratio is traversed, the joint processing time delay under the task segmentation ratio is calculated, and it is determined whether the time delay requirement is met. If not, the next task segmentation ratio is continuously traversed until a task segmentation ratio satisfying the latency requirement is obtained or the traversal is completed.
5. The scheduling method of claim 1, wherein, Based on the task data amount information, the first transmission rate and the second transmission rate, the individual processing latency and / or the joint processing latency of the AI processing module are obtained, and it is determined whether the latency requirement is satisfied. If yes, the step of generating the execution scheme of the target task based on the individual processing latency and / or the joint processing latency satisfying the latency requirement comprises: Based on the task data amount information, the first transmission rate and the second transmission rate, the individual processing latency of each AI processing module and the joint processing latency under different task segmentation ratios are calculated and collected into a queue; The processing latency capable of satisfying the latency requirement is filtered from the queue and sorted according to a preset priority order. The preset priority order comprises that the priority of the individual processing latency of the second AI processing module is higher than that of the individual processing latency of the first AI processing module, and the priority of the joint processing latency under different task segmentation ratios is higher when the task processing ratio of the second AI processing module is higher. Based on the processing latency with the highest priority, the execution scheme of the target task is generated.
6. The scheduling method of claim 1, wherein, The calculation method of the individual processing latency of the second AI processing module comprises: Based on the AI model data amount and the second transmission rate of the second AI processing module, the model loading latency of the second AI processing module is obtained; Based on the task calculation inference amount and the calculation rate of the second AI processing module, the inference latency of the second AI processing module is obtained; Based on the model loading latency and the inference latency of the second AI processing module, the individual processing latency of the second AI processing module is obtained.
7. The scheduling method of claim 1, wherein, The calculation method of the individual processing latency of the first AI processing module comprises: Based on the AI model data amount and the second transmission rate of the first AI processing module, the model loading latency of the first AI processing module is obtained; Based on the task calculation inference amount and the calculation rate of the first AI processing module, the inference latency of the first AI processing module is obtained; Based on the task input data amount and the first transmission rate, the input latency of the first AI processing module is obtained; Based on the task output data amount and the first transmission rate, the output latency of the first AI processing module is obtained; Based on the model loading latency, the inference latency, the input latency and the output latency of the first AI processing module, the individual processing latency of the first AI processing module is obtained.
8. The scheduling method of claim 1, wherein, The calculation method of the joint processing latency comprises: The target task is segmented into a first task and a second task according to a task segmentation ratio, wherein the first task is executed by the first AI processing module, and the second task is executed by the second AI processing module; obtaining, based on the task computation inferences of the first task and the second task and the computation rates of the first AI processing module and the second AI processing module, inference time delays of the first AI processing module and the second AI processing module; obtaining, based on the AI model data amounts corresponding to the first task and the second task and the second transmission rate, model loading time delays of the first AI processing module and the second AI processing module; obtaining, for different processing sequences of the first task and the second task, first interaction time delays and second interaction time delays, the first interaction time delay being an interaction time delay when the first task is processed before the second task, and the second interaction time delay being an interaction time delay when the second task is processed before the first task; obtaining, based on the inference time delays, the model loading time delays, and the first interaction time delays and the second interaction time delays of the first AI processing module and the second AI processing module, first joint processing time delays and second joint processing time delays under the task segmentation ratio.
9. The scheduling method of claim 8, wherein, The method for calculating the first interaction time delay comprises: obtaining the first interaction time delay based on the task input data amount of the first task and the first transmission rate; or The method for calculating the second interaction time delay comprises: obtaining the second interaction time delay based on the task output data amount of the first task and the first transmission rate.
10. The scheduling method of claim 1, wherein, If there is no single processing time delay and / or joint processing time delay meeting the time delay requirement, the method further comprises: sorting the single processing time delays and / or joint processing time delays according to sizes, generating an execution scheme of the target task based on the smallest processing time delay, and feeding back an execution result to the network side; or giving up executing the target task and feeding back execution failure information to the network side; or sorting the single processing time delays and / or joint processing time delays according to sizes, and feeding back the smallest processing time delay to the network side for the network side to determine whether to execute the target task.
11. A dispatch apparatus of an artificial intelligence processor, characterized by comprising: The method is applied to a terminal, and the terminal internally deploys a plurality of AI processing modules and memories which are signal-connected to each other, wherein the plurality of AI processing modules comprise at least one first AI processing module and at least one second AI processing module deployed in a modem; The scheduling device comprises: a task acquisition module configured to acquire target task information, the target task information comprising a type, a size, and a time delay requirement of a target task; a task amount acquisition module configured to acquire task data amount information based on the target task information, the task data amount information comprising AI model data amounts, task input data amounts, task output data amounts, and task computation inferences; and a processing time delay calculation module configured to calculate single processing time delays and / or joint processing time delays of the target task based on the task data amount information. The computing determining module is configured to obtain a single processing delay and / or a joint processing delay of the AI processing module based on the task data amount information, the first transmission rate and the second transmission rate, and determine whether the delay requirement is met. If yes, an execution scheme of the target task is generated based on the single processing delay and / or the joint processing delay that meets the delay requirement. The first transmission rate is a data transmission rate between the first AI processing module and the second AI processing module, and the second transmission rate is a data transmission rate between the AI processing module and the memory.
12. A terminal comprising: at least one processor; at least one first AI processing module; at least one second AI processing module deployed in a modem; and a memory storing instructions that, when executed by the at least one processor, cause the at least one processor to perform the scheduling method of the artificial intelligence processor according to any one of claims 1 to 10.
13. A machine-readable storage medium storing executable instructions that, when executed, cause the machine to perform the scheduling method of the artificial intelligence processor according to any one of claims 1 to 10.