Maneuvering application scene-oriented AI model scheduling method and system

By building an AI model library and combining the execution timing and priority of the task set, energy consumption evaluation and model replacement are carried out, the problem of low resource utilization in maneuvering edge devices is solved, and resource balance and task execution efficiency are improved.

CN120256071AActive Publication Date: 2025-07-04ZHONGKE EDGE SMART INFORMATION TECH (SUZHOU) CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510743414.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-05
Publication Date
2025-07-04
Estimated Expiration
2045-06-05

AI Technical Summary

Technical Problem

In the prior art, the AI ​​model selection method of maneuvering edge devices cannot be dynamically adjusted according to the real-time requirements and priority of tasks, resulting in low utilization of computing power resources, low task execution efficiency, and unable to effectively control resource use when computing power resources consume more than power capacity, resulting in waste of resources or degradation of task execution quality.

Method used

By building an AI model library, combining the execution timing and priority of the task set, energy consumption evaluation and model replacement are carried out to ensure that the resource usage is balanced between 70% and 80%, and the accuracy and size of the model are dynamically adjusted to optimize resource utilization.

Benefits of technology

While ensuring the operation of tasks, maintain resource balance, ensure that the system operates stably under the power capacity limit, improve resource utilization and task execution efficiency, and effectively control the use of computing resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120256071A_ABST
    Figure CN120256071A_ABST
Patent Text Reader

Abstract

The invention provides a maneuvering application scene-oriented AI model scheduling method and system, and the method comprises the steps: carrying out the calculation and evaluation of a resource overall use mean value of each task in a task set, and carrying out the priority degradation and model replacement work in combination with the task priority and the overall usage amount of maneuvering environment resources. The model used by the task is adjusted in time according to the model precision, the model size and the task priority, so that resource balance is kept at 70%-80% while task running is guaranteed, it is guaranteed that the system can still run stably under the limitation of the power capacity, and resource utilization and task execution quality are better balanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of AI scheduling, and particularly relates to an AI model scheduling method and system for mobile application scenarios. Background Art

[0002] As a cutting-edge technology, mobile edge computing mainly executes computing tasks on mobile devices or temporarily deployed edge nodes. These devices, such as drones, vehicle-mounted systems, mobile base stations, etc., mainly face the challenge of resource limitations. With the rapid development of the Internet of Things and 5G technology, the role of mobile edge computing in key fields such as civilian life and emergency has become increasingly prominent. The application of artificial intelligence technology in mobile edge computing is particularly crucial, covering multiple aspects such as target recognition, environmental perception, and decision support. However, the complexity of the corresponding AI models and the high demand for computing resources significantly conflict with the resource limitations of mobile edge devices. Existing resource scheduling strategies, whether static resource allocation or dynamic resource adjustment, have certain limitations and are difficult to achieve efficient matching of resources and AI task requirements. With the continuous improvement of the complexity and accuracy of AI models, the demand for computing power resources is also continuously increasing. How to efficiently run multiple AI tasks with limited computing power resources has become an urgent problem to be solved.

[0003] The disadvantages existing in the prior art are that the model selection method is often relatively fixed and cannot dynamically adjust the size and accuracy of the model according to the real-time requirements and priorities of the tasks, resulting in low utilization rate of computing power resources and low task execution efficiency. When the power consumption of computing power resources exceeds the power capacity, the model operation state cannot be adjusted in time, and the efficient use of computing power resources cannot be effectively controlled while ensuring the accuracy of the model, resulting in resource waste or a decline in task execution quality. Summary of the Invention

[0004] The main problem solved by the present invention is how to maximize the utilization of computing power resources while ensuring that the accuracy of the model is not affected when the power consumption of computing power resources exceeds the preset value, and provides an AI model scheduling method and system for mobile application scenarios.

[0005] To solve the above technical problems, the technical solutions adopted are as follows: An AI model scheduling method for mobile application scenarios, comprising the following steps: Step 1: Construct an AI model library, and the attributes of each model in the model library include the resource configuration information of the model, the accuracy of the model, and the size of the model; Step 2: Obtain each task to be executed and sort each task according to the execution time sequence to form a task set, and each task in the task set has execution time sequence and priority attributes; Step 3: Match the corresponding models to each task in the task set according to the pre-set task and AI model selection rules, and then evaluate the energy consumption of each task in the task set; Step 4: If the energy consumption evaluation is within the range of resource balance, use the AI model matched by the current task to execute according to the execution time sequence of each task and end the scheduling. The resource balance means that the overall resource usage average is between 70% and 80%; If the resources are insufficient or sufficient after the energy consumption evaluation, replace the matching model of the task being executed in the task set and return to Step 3 for energy consumption evaluation until the energy consumption evaluation is resource balanced. The resource insufficiency means that the overall resource usage average exceeds 80%, and the resource sufficiency means that the overall resource usage average is lower than 70%.

[0006] Further, the method for energy consumption evaluation is: Step 3.1: Obtain the overall power consumption of the current mobile edge environment, the maximum capacity limits of GPU and memory MEM, and the actual resource utilization rates of power, GPU, and memory MEM in the mobile edge environment at the current moment; Step 3.2: Calculate the energy consumption value required to execute the tasks up to the current moment; Step 3.3: Conduct a lowest energy consumption game for the tasks in the task set: Select low-precision and small-size models for each task in the task set, and calculate the lowest energy consumption value from the moment of executing the first task in the task set to the current moment; Conduct an optimal energy consumption game for the tasks in the task set: Select models for each task in the task set according to the priority. High-priority tasks prefer to select models on the premise of model accuracy, and low-priority tasks prefer to select models on the premise of model size. Calculate the optimal energy consumption value from the moment of executing the first task in the task set to the current moment; Step 3.4: Calculate the overall average resource usage of the mobile edge environment corresponding to the lowest energy consumption game and the optimal energy consumption game respectively; Step 3.5: If the overall average resource usage is between 70% and 80%, use the current model configuration; If the overall average resource usage is lower than 70% or higher than 80%, replace the matching model of the task being executed in the task set.

[0007] Further, the method for replacing the matching model of the task being executed in the task set is: If the overall average resource utilization exceeds 80%, starting from the highest-priority task being executed, replace it with a model that has a precision one level lower than the model matched by the current task. Then, determine whether the overall average resource utilization is between 70% and 80%. If not, reduce the volume of the current model by one level and then determine whether the overall average resource utilization is between 70% and 80%. If not, continue to select the model of the next higher-priority task being executed for replacement until the overall average resource utilization is between 70% and 80%. If the overall average resource utilization is below 70%, starting from the lowest-priority task being executed, replace it with a model that has a precision one level higher than the model matched by the current task. Determine whether the overall average resource utilization is between 70% and 80%. If not, select a model with a volume one level larger than the model matched by the current task. Then, determine again whether the overall average resource utilization is between 70% and 80%. If not, continue to select the model of the next lower-priority task being executed for replacement until the overall average resource utilization is between 70% and 80%.

[0008] Furthermore, the pre-set task and AI model selection rule is: Select models for high-priority tasks with high precision as the priority, and select models for low-priority tasks with model minimization as the priority.

[0009] Furthermore, the method for calculating the energy consumption value required for the tasks executed up to the current moment is: ; wherein, represents the energy consumption value required for executing the task within the fixed time duration ; is the current moment , and the overall average utilization of the mobile edge environment resources; ; , and respectively represent the actual resource utilization rates of power consumption, GPU, and memory MEM in the mobile edge environment at the current moment ; ; wherein, represents a binary decision variable: if the task selects the model , then holds; otherwise, holds; , and respectively represent the power consumption, GPU, and memory resource usage of the current edge computing environment at the th moment; N represents the number of models in the model set, and M represents the number of tasks; , and respectively represent the maximum capacity limits of the overall power consumption, GPU, and memory MEM of the current mobile edge environment; ; represents the precision i th model in the model set for the j th task in the task set, and according to the priority j th task, select the th model, and the corresponding resource requirement values of the size i ; and size , including , , , , respectively represent the overall power consumption, GPU, and memory MEM resource requirement values of the current mobile edge environment for the i-th model, represents the task priority of the j th task, respectively represent the precision and size of the i th model, represents the precision and size of the selected i-th model.

[0010] The present invention also provides an AI model scheduling system for mobile application scenarios, using the steps of an AI model scheduling method for mobile application scenarios.

[0011] Adopting the above technical solutions, the present invention has the following beneficial effects: An AI model scheduling method and system for mobile application scenarios provided by the present invention calculate and evaluate the overall resource usage mean of each task in the task set, and then combine the task priority and the overall resource usage of the mobile environment to perform priority degradation and model replacement work. The model used by the task is adjusted in a timely manner according to the model precision, model size, and task priority, so as to maintain resource balance while ensuring task operation, and maintain it at 70% to 80%, ensuring that the system can still operate stably under the power capacity limit, and better balancing resource utilization and task execution quality. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Figure 1 is the system flow chart provided by the present invention; Figure 2 is the Gantt chart of AI model loading based on task timing. Detailed implementation manners

[0013] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are some, but not all, embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0014] Figure 1 A specific embodiment of an AI model scheduling method for a mobile application scenario of the present invention is shown, including the following steps: Step 1: Construct an AI model library, and the attributes of each model in the model library include the resource configuration information of the model, the accuracy of the model, and the size of the model.

[0015] In this embodiment, the constructed AI model library contains model libraries with various precisions and sizes, covering convolutional neural networks (CNNs), recurrent neural networks (RNNs), long short-term memory networks (LSTMs), etc. When performing model scheduling, different-sized and different-precision models are selected from the model library according to the known computing power resources, task timings, and the priorities of the tasks to be scheduled.

[0016] The set of models in the model library can be expressed as , where represents the rd model in the model library, and N represents the number of models in the model library. The set of resource configuration information of each model in the model library can be expressed as , represents the resource information of the rd model in the model library, , where , respectively represent the accuracy and size of the rd model.

[0017] Step 2: Obtain each task to be executed and sort each task according to the execution timing sequence to form a task set, and each task in the task set has execution timing and priority attributes.

[0018] In this embodiment, several factors affecting model scheduling are comprehensively considered, namely the front and back timings of task scheduling, the scheduling priority of the task itself, the accuracy of the model, the size of the model, and the power consumption of computing power resources. That is, when executing a task, both task conditions and model conditions need to be considered, and the most suitable model is configured for the task through comprehensive consideration.

[0019] The task set is marked as , where M represents the number of tasks. The tasks in the task set have been sorted in the order of their execution time sequences, is the first task to be executed. These tasks need to select corresponding types of models from the model library for execution. Each task is represented by a binary tuple , where represents the priority of the th task. The priorities are divided into high priority and low priority. For high-priority tasks, models are selected for high-precision purposes, and for low-priority tasks, models are selected for minimum model size purposes; represents the

[0020] Step 3: Match the corresponding models for each task in the task set according to the pre-set task and AI model selection rules, and then evaluate the energy consumption of each task in the task set.

[0021] In this embodiment, the pre-set task and AI model selection rules are: For high-priority tasks, models are preferentially selected for high-precision purposes, and for low-priority tasks, models are preferentially selected for minimum model size purposes. For each task , select the corresponding model from the model library .

[0022] In this embodiment, the method for energy consumption evaluation is: Step 3.1: Obtain the overall power consumption of the current mobile edge environment, the maximum capacity limits of the GPU and the memory MEM, and the actual resource utilization rates of the power consumption, GPU, and memory MEM in the mobile edge environment at the current time t; and respectively represent the overall power consumption of the current mobile edge environment, the maximum capacity limits of the GPU and the memory MEM.

[0023] , and respectively represent the actual resource utilization rates of the power consumption, GPU, and memory MEM in the mobile edge environment at the current time ; (1) Among them, represents a binary decision variable: if task selects model , then holds; otherwise, holds; , and respectively represent the current edge computing environment at the The power consumption at a moment, and the resource usage of the GPU and memory; N represents the number of models in the model set, and M represents the number of tasks.

[0024] Through the above formula, the model library and the task set are formalized as a cooperative game. In this cooperative game, the players are the selected models. According to the j priority of the ith task , the resource demand value corresponding to the ith model selected with precision priority or size priority is ; represents the resource demand value corresponding to the precision i and size j of the j ith task in the task set for the i ith model in the model set, selected according to the priority of the ith task , , , respectively representing the overall power consumption of the current mobile edge environment, the resource demand values of the GPU and the memory MEM of the ith model, represents the task priority of the j ith task, respectively represent the precision and size of the i ith model, and represents the precision and size of the selected ith model.

[0025] Step 3.2: Calculate the energy consumption value required to execute the tasks up to the current moment. The energy consumption value is used to measure the specific energy consumption value of the mobile edge device as the power consumption of the mobile edge environment, the GPU, and the memory resources change during a fixed-duration task time. The larger the energy consumption value, the greater the resource consumption; the smaller the energy consumption value, the smaller the resource consumption. When the energy consumption value is small, although the resource consumption decreases, it also brings a loss in the precision of the model.

[0026] In this embodiment, the method for calculating the energy consumption value required to execute the tasks from the moment when the first task is executed to the current moment is: (2) Among them, represents the energy consumption value required to execute the task within the fixed duration ; is the current moment , the overall average usage of mobile edge environmental resources; (3).

[0027] Step 3.3: Conduct minimum energy consumption game for the tasks in the task set: Select low-precision and small-volume models for each task in the task set, and calculate the minimum energy consumption value from the start time of the first task execution in the task set to the current time. After matching models for each task in the task set, use Formula 2 to calculate the minimum energy consumption value.

[0028] Conduct optimal energy consumption game for the tasks in the task set: Select models for each task in the task set according to the priority. High-priority tasks prefer to select models on the premise of model accuracy, and low-priority tasks prefer to select models on the premise of model size. Calculate the optimal energy consumption value from the start time of the first task execution in the task set to the current time. Use Formula 2 to calculate the optimal energy consumption value.

[0029] The game is carried out in two steps. First, conduct the minimum energy consumption game. Each task selects a low-precision and small-volume model, and finally calculates the minimum energy consumption value ; then conduct the optimal energy consumption game. Each task selects a model according to the priority. High-priority tasks prefer to ensure the scheduling and operation of the model on the premise of model accuracy, and low-priority tasks prefer to ensure the scheduling and operation of the model on the premise of model size. Substitute into the above formula, and monitor the overall average usage value of resources to see if it is in the range of 70% to 80%. If not, adjust the model. After several rounds of games, the players will achieve mutual win. At this time, the expected fairness among each model in the model library reaches the maximum, and finally calculate the optimal energy consumption value carried by this mobile edge device .

[0030] Step 3.4: Calculate the overall average usage of mobile edge environmental resources corresponding to the minimum energy consumption game and the optimal energy consumption game respectively; Step 3.5: If the overall average usage of resources is between 70% and 80%, use the current model configuration; If the overall average usage of resources is lower than 70% or higher than 80%, replace the matching model of the task being executed in the task set.

[0031] Step 4: If the energy consumption assessment is within the range of resource balance, use the AI model matched by the current task to execute according to the execution time sequence of each task and end the scheduling. The resource balance means that the overall average usage of resources is between 70% and 80%; If the resources are insufficient or sufficient after energy consumption assessment, replace the matching model of the task being executed in the task set and return to step 3 for energy consumption assessment until the energy consumption assessment shows balanced resources. The insufficient resources mean that the overall average resource usage exceeds 80%, and the sufficient resources mean that the overall average resource usage is below 70%.

[0032] In this embodiment, the method for replacing the matching model of the task being executed in the task set is as follows: If the overall average resource usage exceeds 80%, start from the highest-priority task being executed and replace it with a model whose accuracy is one level lower than that of the model matched by the current task. Determine whether the overall average resource usage is between 70% and 80%. If not, reduce the volume of the current model by one level, and then determine whether the overall average resource usage is between 70% and 80%. If not, continue to select the next higher-priority task being executed for model replacement until the overall average resource usage is between 70% and 80%. If the overall average resource usage is below 70%, start from the lowest-priority task being executed and replace it with a model whose accuracy is one level higher than that of the model matched by the current task. Determine whether the overall average resource usage is between 70% and 80%. If not, select a model whose volume is one level larger than that of the model matched by the current task, and then determine whether the overall average resource usage is between 70% and 80%. If not, continue to select the next lower-priority task being executed for model replacement until the overall average resource usage is between 70% and 80%.

[0033] In this embodiment, during the model scheduling and operation, by monitoring whether the overall average resource usage is in the state of balanced resources (maintained between 70% and 80%), the maximized utilization of computing power resources is ensured. If it is not balanced resources, the matching model of the current task needs to be replaced and other models of different sizes are reselected. Suppose when the GPU computing power resources are insufficient in the case of running multiple models simultaneously, on the premise of ensuring that the model accuracy does not affect the business, the priority of the high-priority tasks is automatically downgraded, and through the downgrading method, the model corresponding to the task is replaced, so that the overall business is effectively guaranteed and the existing computing power resources are optimally used.

[0034] The present invention can intelligently select AI models with different accuracies and different sizes for scheduling on the premise of ensuring the normal operation of AI recognition tasks, realize the optimal utilization of computing power resources, and calculate the minimum energy consumption requirement and the optimal energy consumption requirement for task execution within a fixed duration, ensuring that the energy equipment carried when starting to execute the task meets the basic operation of the task and the optimal guarantee of the task.

[0035] The following is a specific description through a real-time case: Background: Suppose there is a computing power resource pool with a power consumption limit of 400W, 8-core GPUs, and 32G of memory, and four AI recognition tasks need to be completed within a day. The tasks have different timing requirements and priorities, and there are models of different sizes and precisions available in the model library for selection. The resource balance degree does not exceed 80%.

[0036] As Figure 2 shown in the Gantt chart is described as follows: I. Details of AI model loading: (a) Task A: Starts at 8:00 am, has a high priority, and applies a high-precision model.

[0037] (b) Task B: Starts at 9:00 am, has a low priority, and applies a low-precision model.

[0038] (c) Task C: Starts at 12:00 pm, has a low priority, and applies a low-precision model.

[0039] (d) Task D: Starts at 2:00 pm, has a medium priority, and applies a high-precision model.

[0040] II. Gantt chart description: (a) Time axis: From 8:00 to 16:00, divided into hours.

[0041] (b) Task A: Starts at 8:00, is expected to run for 6 hours, and ends at 14:00.

[0042] (c) Task B: Starts at 9:00, is expected to run for 2 hours, and ends at 11:00.

[0043] (d) Task C: Starts at 12:00, is expected to run for 2 hours, and ends at 14:00.

[0044] (e) Task D: Starts at 14:00, is expected to run for 2 hours, and ends at 16:00.

[0045] III. Changes in resource balance degree: (a) 8:00: Task A starts. Resources are sufficient. According to the pre-set rules for task and AI model selection, a high-precision and large-volume model is matched, and the overall average resource utilization is 70%.

[0046] (b) 9:00: Task B starts. The resource balance state. According to the pre-set rules for task and AI model selection, a low-precision and medium-volume model is matched, and the overall average resource utilization is 80%.

[0047] (c) 11:00: Task B ends. The overall average resource utilization drops to 70%.

[0048] (d) 12:00: Task C starts. The remaining resources are balanced. According to the pre-set rules for task and AI model selection, a low-precision medium-volume model is matched. The overall average resource usage rises to 85%, and the power consumption exceeds 400W. Since the remaining resources are insufficient, the priority of Task A is then lowered from high to medium, and a low-precision small-volume model is matched for replacement scheduling. The resource balance drops to 75% equilibrium state, and the power consumption is reduced to 350W.

[0049] (e) 14:00: Task C ends and Task D starts. The remaining resources are balanced. According to the pre-set rules for task and AI model selection, a high-precision medium-volume model is matched, and the resource balance rises to 85%. Since the remaining resources are insufficient, at this time, the tasks being run are A and D, and the priorities of both tasks are medium. Starting from Task A in the same priority order, it is detected that it has been downgraded to use a low-precision small-volume model, so the downgrading operation continues with the next Task D, reducing the model size of Task D, replacing the running model, adopting a high-precision small-volume model and performing scheduling. The resource balance drops to 80%, and the power consumption is reduced to 320W.

[0050] (f) 16:00: Tasks A and D end simultaneously, and the resource balance drops to 0%.

[0051] It can be seen that through the selection of models with different precisions and volumes, the overall resource balance is maintained in the range of 70% to 80%, and the overall power consumption is controlled below 400W.

[0052] The present invention can dynamically adjust the size and precision of the AI model according to the timing and scheduling priority of the task, thereby significantly improving the utilization rate of GPU resources and the task execution efficiency. By means of model replacement, the use of GPU resources is effectively controlled to ensure that the system can still operate stably under the power capacity limit. During the model replacement process, on the premise of ensuring the model accuracy, the efficient utilization of GPU computing power resources can be achieved, always maintained in the range of 70% to 80%, better balancing resource utilization and task execution quality, and calculating the optimal energy consumption requirements for task execution.

[0053] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. An AI model scheduling method for mobile application scenarios, characterized in that, It includes the following steps: Step 1: Build an AI model library. The attributes of each model in the model library include the resource configuration information of the model, the accuracy of the model, and the size of the model; Step 2: Obtain each task to be executed and sort each task according to the execution time sequence of the tasks to form a task set. Each task in the task set has execution time sequence and priority attributes; Step 3: Match the corresponding models for each task in the task set according to the pre-set task and AI model selection rules, and then perform energy consumption evaluation on each task in the task set; Step 4: If the energy consumption evaluation is within the range of resource balance, use the AI model matched by the current task to execute according to the execution time sequence of each task and end the scheduling. The resource balance means that the overall resource usage average is between 70% and 80%; If the resources are insufficient or sufficient after the energy consumption evaluation, replace the matching model of the task being executed in the task set and return to Step 3 for energy consumption evaluation until the energy consumption evaluation is resource balanced. The resource insufficiency means that the overall resource usage average exceeds 80%, and the resource sufficiency means that the overall resource usage average is lower than 70%.

2. The AI model scheduling method according to claim 1, wherein The method of energy consumption evaluation is: Step 3.1: Obtain the overall power consumption of the current mobile edge environment, the maximum capacity limits of GPU and memory MEM, and the actual resource utilization rates of power, GPU, and memory MEM in the mobile edge environment at the current moment; Step 3.2: Calculate the energy consumption value required to execute the tasks up to the current moment; Step 3.3: Perform the lowest energy consumption game on the tasks in the task set: Select low-precision and small-size models for each task in the task set, and calculate the lowest energy consumption value from the moment when the first task in the task set starts to the current moment; Perform the optimal energy consumption game on the tasks in the task set: Select models for each task in the task set according to the priority. High-priority tasks preferentially select models on the premise of model accuracy, and low-priority tasks preferentially select models on the premise of model size, and calculate the optimal energy consumption value from the moment when the first task in the task set starts to the current moment; Step 3.4: Calculate the overall average resource usage of the mobile edge environment corresponding to the lowest energy consumption game and the optimal energy consumption game respectively; Step 3.5: If the overall average resource usage is between 70% and 80%, use the current model configuration; If the overall average resource usage is lower than 70% or higher than 80%, replace the matching model of the task being executed in the task set.

3. The AI model scheduling method according to claim 2, wherein The method of replacing the matching model of the task being executed in the task set is: If the overall average resource usage exceeds 80%, start from the highest-priority task being executed, use a model with an accuracy one level lower than the model matched by the current task for replacement, judge whether the overall average resource usage is between 70% and 80%. If not, reduce the volume of the current model by one level, and then judge whether the overall average resource usage is between 70% and 80%. If not, continue to select the next highest-priority task being executed for model replacement until the overall average resource usage is between 70% and 80%; If the overall average resource utilization is lower than 70%, starting from the lowest-priority task being executed, replace it with a model that has a precision one level higher than the model matched by the current task. Determine whether the overall average resource utilization is between 70% and 80%. If not, select a model whose volume is one level larger than the volume of the model matched by the current task, and again determine whether the overall average resource utilization is between 70% and 80%. If not, continue to select the next lowest-priority task being executed to start model replacement until the overall average resource utilization is between 70% and 80%.

4. The AI model scheduling method according to claim 3, wherein The pre-set task and AI model selection rules are as follows: For high-priority tasks, select models with high precision as the priority, and for low-priority tasks, select models with the smallest size as the priority.

5. The AI model scheduling method according to claim 4, wherein The method for calculating the energy consumption value required for the tasks executed up to the current moment is as follows: ; Among them, represents the energy consumption value required to execute a task within a fixed time duration ; is the current moment , and is the overall usage average of mobile edge environmental resources; ; , and respectively represent the power consumption, GPU, and actual resource utilization of memory MEM in the maneuvering edge environment at the current moment ; ; represents a binary decision variable: if the task selects the model , then holds; otherwise, holds; , and respectively represent the power consumption, GPU, and memory resource usage of the current edge computing environment at the th moment; N represents the number of models in the model set, and M represents the number of tasks; , and respectively represent the maximum capacity limits of the overall power consumption, GPU, and memory MEM of the current mobile edge environment; ; It represents that for the i -th model in the model set and the j -th task in the task set, according to the priority of the j -th task, select the precision i and size of the -th model, and the corresponding resource requirement values, including , , , which respectively represent the overall power consumption of the current maneuvering edge environment, the resource requirement values of GPU and memory MEM of the i-th model, represents the task priority of the j -th task, respectively represent the precision and size of the i -th model, and represent the precision and size of the selected i-th model.

6. An AI model scheduling system for mobile application scenarios, characterized in that, Use each step of the AI model scheduling method for a mobile application scenario described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Cooperative scheduling system, method and device for computing power resources and storage medium

    CN115562824A

  • Large model scheduling method and device based on NPU computing power

    CN119336457A

  • Maneuvering edge application scene-oriented computing power resource evaluation method and system

    CN119376955A

  • Multi-task dynamic resource scheduling method

    WO2021233261A1