Task scheduling method and system based on energy consumption optimization
By combining multi-agent deep reinforcement learning with multi-bee colony optimization algorithm, the task scheduling strategy is dynamically adjusted, which solves the problems of uneven resource allocation and difficulty in optimizing energy consumption in heterogeneous computing systems. It achieves synergistic optimization of computing efficiency and energy consumption, and improves the overall performance and resource utilization of the system.
Patent Information
- Application Number
- CN202511110416.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-08
- Publication Date
- 2025-11-18
AI Technical Summary
Traditional task scheduling methods struggle to dynamically match tasks and resources in heterogeneous computing systems, resulting in uneven resource allocation, poor adaptability, and difficulty in achieving synergistic optimization between computing efficiency and energy consumption. In particular, when facing dynamically changing high-load task scenarios, they suffer from local optima and low resource utilization.
A scheduling method combining multi-agent deep reinforcement learning and multi-bee colony optimization algorithm is adopted. By dynamically adjusting the task scheduling strategy, generating the initial task-resource allocation strategy using deep reinforcement learning, and performing global optimization through multi-bee colony optimization, the computational efficiency is maximized and the energy consumption is minimized.
It improves the computing efficiency and energy consumption of heterogeneous computing systems, achieves adaptive optimization of task scheduling, avoids local optimum traps, balances exploration and development, and improves system energy efficiency and resource utilization.
Smart Images

Figure CN120973534A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the field of artificial intelligence and computing resource scheduling optimization, in particular to a task scheduling method and system based on energy consumption optimization. BACKGROUND
[0002] With the rapid development of artificial intelligence and big data technology, a large number of computing-intensive and delay-sensitive AI tasks have emerged in the medical industry, such as CT image recognition, genetic data processing, pathological image analysis, and natural language medical record understanding. These tasks require high concurrency, high performance, and low energy consumption scheduling and computing on computing platforms.
[0003] To meet the diversity requirements of different tasks in computing density, bandwidth demand, and energy efficiency targets, current medical AI platforms usually deploy multi-element heterogeneous AI chip computing resources composed of CPUs, GPUs, FPGAs, NPUs, etc. to support multi-type task collaborative processing. However, as the complexity and dynamics of tasks increase, traditional scheduling methods face the following key challenges: (1) Difficulty in matching heterogeneous resources: Different tasks have significant differences in resource preferences for chip types, and fixed rules or static strategies are difficult to dynamically match; (2) System load fluctuation is severe: Task arrival has burstiness and periodicity, and resource scheduling needs to have high real-time performance and adaptability; (3) Energy efficiency targets are difficult to balance: There are conflicts between performance, energy consumption, and bandwidth utilization in multi-objective optimization, which is difficult to satisfy simultaneously.
[0004] In actual deployment, traditional task scheduling methods are mostly based on heuristic algorithms or centralized reinforcement learning strategies, which have certain optimization capabilities, but in complex and variable heterogeneous system environments, they have the following problems: (1) Prone to local optimum, lack of global perspective; (2) Uneven resource allocation, some chips are overloaded and some are idle; (3) Poor adaptability, lack of rapid response capability to sudden tasks.
[0005] In addition, existing methods often use only a single optimization algorithm (such as genetic algorithm, Q-learning, etc.), which has problems such as imbalance between exploration and development, limited search capability, etc. Especially in the face of dynamic and high-load task scenarios, it is difficult to achieve the coordinated optimization of task completion efficiency and system energy consumption. SUMMARY
[0006] To solve the above problems, the present disclosure proposes a task scheduling method and system based on energy consumption optimization, which dynamically adjusts the task scheduling strategy through multi-agent learning and performs global optimization using multi-swarm optimization to maximize the computing efficiency and minimize the energy consumption of tasks.
[0007] According to some embodiments, the present disclosure adopts the technical solutions as follows: A task scheduling method based on energy consumption optimization, comprising: obtaining dynamic task scheduling requirements, including a task set to be scheduled, a heterogeneous computing resource set, a task load characteristic matrix and a computing resource capability matrix; based on the dynamic task scheduling requirements, using a scheduling model that integrates deep reinforcement learning and a multi-swarm optimization algorithm, performing multi-objective optimization on a task-resource allocation strategy in terms of computing efficiency, storage bandwidth utilization and energy consumption according to task load and resource capability, and obtaining an optimal solution; wherein the scheduling model takes computing efficiency, storage bandwidth utilization and energy consumption as multiple objectives, generates an initial number of task-resource allocation strategies through deep reinforcement learning, performs global optimization on the initial number of task-resource allocation strategies using the multi-swarm optimization algorithm, and obtains an optimal task-resource allocation strategy.
[0008] According to some embodiments, the present disclosure adopts the technical solutions as follows: A task scheduling system based on energy consumption optimization, comprising: a requirement obtaining module configured to obtain dynamic task scheduling requirements, including a task set to be scheduled, a heterogeneous computing resource set, a task load characteristic matrix and a computing resource capability matrix; a strategy optimization module configured to, based on the dynamic task scheduling requirements, use a scheduling model that integrates deep reinforcement learning and a multi-swarm optimization algorithm, perform multi-objective optimization on a task-resource allocation strategy in terms of computing efficiency, storage bandwidth utilization and energy consumption according to task load and resource capability, and obtain an optimal solution; wherein the scheduling model takes computing efficiency, storage bandwidth utilization and energy consumption as multiple objectives, generates an initial number of task-resource allocation strategies through deep reinforcement learning, performs global optimization on the initial number of task-resource allocation strategies using the multi-swarm optimization algorithm, and obtains an optimal task-resource allocation strategy.
[0009] According to some embodiments, the present disclosure adopts the technical solutions as follows: A computer program product comprising a computer program, which, when executed by a processor, implements the task scheduling method based on energy consumption optimization.
[0010] According to some embodiments, the present disclosure adopts the technical solutions as follows: A non-transitory computer-readable storage medium for storing computer instructions, which, when executed by a processor, implements the task scheduling method based on energy consumption optimization.
[0011] According to some embodiments, the present disclosure adopts the technical solutions as follows: An electronic device, comprising: a processor, a memory and a computer program; wherein the processor is connected with the memory, and the computer program is stored in the memory; when the electronic device is running, the processor executes the computer program stored in the memory, so that the electronic device executes the method for task scheduling based on energy consumption optimization.
[0012] Compared with the prior art, the present disclosure has the beneficial effects as follows: The present application realizes task scheduling adaptive optimization through multi-agent reinforcement learning, uses multi-swarm optimization algorithm for global optimization, and finally improves the computing efficiency and energy consumption performance of the heterogeneous computing system, specifically: (1) Through the combination of multi-agent deep reinforcement learning and bee colony optimization algorithm, dynamic adaptation and global optimization of task scheduling are realized; (2) A dynamic adjustment mechanism based on entropy is introduced to effectively balance exploration and development and avoid local optimal trap; (3) Multi-objective collaborative optimization strategy improves system energy efficiency and resource utilization. BRIEF DESCRIPTION OF DRAWINGS
[0013] The drawings accompanying the specification of the present disclosure serve to provide further understanding of the present disclosure, and the illustrative embodiments of the present disclosure and their descriptions serve to explain the present disclosure, and do not constitute an improper limitation on the present disclosure.
[0014] Figure 1 Method flowchart of example 1. DETAILED DESCRIPTION The present disclosure will be further described below in conjunction with the drawings and examples.
[0015] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the present disclosure. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the present disclosure belongs.
[0016] It should be noted that the terms used herein are only for the purpose of describing specific embodiments, and are not intended to limit the exemplary embodiments according to the present disclosure. As used herein, unless the context clearly indicates otherwise, the singular form is intended to include the plural form, and in addition, it should be understood that when the terms "comprise" and / or "comprise" are used in the specification, they indicate the presence of the features, steps, operations, devices, components and / or their combinations.
[0017] Example 1 In an embodiment of the present disclosure, a task scheduling method based on energy consumption optimization is provided. The task scheduling strategy is dynamically adjusted through multi-agent learning, and global optimization is performed using multi-swarm optimization to maximize the calculation efficiency and minimize the energy consumption of the task, as shown in Figure 1 Specifically, the embodiment includes the following steps: Step S1: Obtain the dynamic task scheduling requirement, including a task set to be scheduled, a heterogeneous computing resource set, a task load characteristic matrix, and a computing resource capability matrix. Step S2: Based on the dynamic task scheduling requirement, use a scheduling model that combines deep reinforcement learning and multi-swarm optimization algorithm, and according to the task load and resource capability, perform multi-objective optimization on the task-resource allocation strategy with the calculation efficiency, storage bandwidth utilization, and energy consumption as the target to obtain the optimal solution. The scheduling model takes the calculation efficiency, storage bandwidth utilization, and energy consumption as the multi-objective, generates an initial number of task-resource allocation strategies through deep reinforcement learning, and uses the multi-swarm optimization algorithm to perform global optimization on the initial number of task-resource allocation strategies to obtain the optimal task-resource allocation strategy.
[0018] As an embodiment, the task scheduling method based on energy consumption optimization of the present disclosure is used in a medical AI platform with multiple heterogeneous AI chip types such as CPU, GPU, FPGA, and NPU for an intelligent auxiliary diagnosis system. The CPU is mainly responsible for task scheduling and system management, and the GPU is used for high-throughput medical image processing such as MRI or CT image segmentation and recognition. At the same time, the FPGA can be used to perform real-time analysis of time-sensitive ECG signals, and the NPU can efficiently run deep neural network models for disease prediction or classification. This multi-chip collaborative heterogeneous architecture not only improves the overall system calculation efficiency, but also takes into account the differentiated needs of different tasks for computing resources, helping to achieve a balance between high performance and low power consumption. The specific implementation process is as follows: 1. Task scheduling modeling This embodiment models the task scheduling problem in the medical AI platform as a resource allocation optimization problem on a multi-heterogeneous AI chip platform, which includes the following elements: (1) Medical task set denoted as Each task represents a different type of intelligent medical task, such as medical image segmentation, medical image classification, ECG signal analysis, EEG abnormality detection, real-time monitoring of vital signs, extraction of electronic medical record text information, disease risk prediction, drug recommendation, clinical event recognition, medical question and answer system, etc.
[0019] (2) Heterogeneous computing resource set denoted as , representing the heterogeneous AI chip types in the medical AI platform, including CPU, GPU, FPGA, NPU, etc.
[0020] (3) Medical task load characteristic matrix is defined as , where represents the task load vector on the resource .
[0021] (1) (4) Computing resource capability matrix represents the computing capability, bandwidth, memory upper limit, power consumption upper limit, etc. of each type of heterogeneous computing resource: (2) (5) Comprehensive optimization objective function To achieve the goal of maximizing computing efficiency and energy efficiency optimization, while considering resource constraints and dynamic changes in task load, the following comprehensive optimization objective function is constructed: 1) Computing efficiency : measures the effective amount of computation that can be completed per unit of time, defined as the total amount of computation actually completed by all tasks divided by the total running time: (3) where represents the computing demand of task on resource ; represents the latest completion time of all tasks; represents whether task is assigned to resource ; The higher the value, the higher the computing efficiency, and the goal is to maximize it.
[0022] 2) Storage bandwidth utilization : measures the effective use of storage resource bandwidth: (4) where represents the bandwidth of task on resource ; represents the maximum bandwidth capacity of resource ; The higher the value, the more fully utilized the bandwidth, and the goal is to maximize it.
[0023] 3) Energy consumption : represents the total energy consumed to complete all tasks: (5) where, denotes the task whether it is assigned to a resource ; the energy consumption of the task on the resource ; The lower the value, the lower the energy consumption, and the goal is to minimize.
[0024] To combine the objectives of different dimensions, normalize each objective function: , , (6) where, , the minimum and maximum values of the computational efficiency; , the minimum and maximum values of the storage bandwidth utilization; , the minimum and maximum values of the energy consumption.
[0025] Based on the above, the comprehensive optimization objective function is obtained: (7) where, are the weights of the computational efficiency, storage bandwidth utilization, and energy consumption objective functions, respectively, satisfying Adjust according to actual needs.
[0026] (6) Constraints There are two major categories of constraints: 1) The task scheduling must satisfy the constraint of "each task is assigned only once", which can be expressed in the formula as: (8) 2) Resource capacity constraints, which can be expressed in the formula as: (9) where, denotes the task the computational demand of the task on the resource, denotes the task bandwidth on the resource, the energy consumption of the task on the resource .
[0027] (7) Dynamic load adaptation mechanism To adapt to real-time changes in medical tasks (such as emergency data inflow or system energy consumption limit fluctuations), a dynamic scheduling mechanism is introduced to update the medical task load characteristics in real time , a scheduling model that combines deep reinforcement learning and multi-swarm algorithm is used to optimize the task-resource allocation strategy, and the optimal solution is obtained.
[0028] 2、Design multi-agent deep reinforcement learning (Multi-Agent Deep Reinforcement Learning, MADRL) to learn medical task scheduling strategy (1) First, introduce the basic knowledge of multi-agent Markov game model: The resource allocation optimization problem of a multi-element heterogeneous AI chip platform is modeled as a multi-agent collaborative decision problem. The scheduling control unit is regarded as an agent, and strategy learning and dynamic adjustment are performed based on local observation (task load, resource state) and global target (computing efficiency and energy efficiency optimization) to achieve efficient scheduling in complex environments. Multi-agent deep learning method is used to model its interactive behavior; each agent dynamically adjusts the strategy according to the local environment (medical task demand, resource state) and global target (performance and energy efficiency optimization).
[0029] Multi-agent deep reinforcement learning (Multi-Agent Deep Reinforcement Learning, MADRL) modeling process: The multi-agent scheduling control unit can be abstracted as a multi-agent Markov game model (Markov Game), defined as a five-tuple: (S,A, π,R,γ)(10) State set: represented by S, its elements represent the local resource state and task characteristics perceived by the i-th agent; Action set: represented by A, its elements is the action space of the i-th agent, representing the allocation decision of the task to the heterogeneous resource unit; Execution strategy: represented by π, used to describe the selection of a certain action from the action set A at time t and then replace the state of the agent, which is a mapping from the state set S to the action set A ( ).
[0030] Reward function: described by R, used to describe the instantaneous reward of the action by the environment after executing the action a, the reward function of the present embodiment adopts the comprehensive optimization objective function of formula (7).
[0031] Discount factor: described by γ, used to control the importance of future rewards.
[0032] (2) Based on the above multi-agent Markov game model, design a multi-agent deep reinforcement learning framework: Initialization: Construct the state space of the agent, design the action space, define the reward function, define the reward mechanism, ensure the maximization of global resource utilization efficiency, dynamically adjust the allocation of multiple heterogeneous resources, and achieve self-evolution and self-optimization.
[0033] Agent training: Multi-agent training is conducted using the Centralized Training and Decentralized Execution (CTDE) method.
[0034] Environment simulation: Simulate the dynamic changes in medical task load in the smart medical simulation platform and train deep learning models to adapt to different scenarios.
[0035] Model Deployment: The trained multi-agent system is deployed to a real medical AI platform to perform dynamic medical task scheduling and generate several initial task-resource allocation strategies.
[0036] 3. Multi-colony optimization algorithm The Multiple Bee Colony Optimization (MBCO) algorithm is introduced. Based on the deep learning scheduling strategy, the initial task-resource allocation strategy obtained by multi-agent deep reinforcement learning (MADRL) is globally optimized, that is, the optimal solution for medical resource allocation is globally searched. The specific steps are as follows: (1) Multi-swarm initialization: Based on the task load characteristic matrix and computing resource capability matrix, initialize the task scheduling strategy representation of each individual bee colony.
[0037] Each individual bee colony corresponds to one of several initial task-resource allocation strategies. Its behavior is guided by heuristic rules during the search process. The global search and local optimization are balanced by the exploration and exploitation mechanisms. The quality of the individual solution is measured by the comprehensive optimization objective function of formula (7), thereby gradually approaching the global optimal scheduling strategy.
[0038] (2) Local search and global adjustment: In the resource scheduling optimization of the medical AI platform, the multi-bee colony algorithm is introduced to search for the optimal solution of task-resource allocation through the collaborative mechanism of exploration and development.
[0039] While maintaining the diversity of local scheduling strategies, the algorithm dynamically adjusts the global resource allocation structure, optimizes the utilization of computing resources and the efficiency of storage access, thereby improving the overall system's computing performance and reducing energy consumption.
[0040] (3) Exploring and developing a balance mechanism: Exploration: Encourage individuals to randomly explore within the search space, avoiding local optima.
[0041] Exploitation: Strengthen the focused search on high-quality solutions, optimizing medical task allocation through reinforcement learning.
[0042] (4) By real-time sensing of computing resource status (such as CPU / GPU load, storage bandwidth, power fluctuation, etc.), feedback to the swarm algorithm optimization process, used for dynamic adjustment of individual fitness and search direction, realizing online optimization of medical task scheduling strategy.
[0043] To realize the dynamic adjustment of exploration and exploitation weights in MABC, a dynamic adjustment mechanism based on entropy is introduced. In the early exploration stage, the entropy value is increased to increase randomness, and in the later optimization stage, the entropy value is reduced to focus optimization. Specifically: Dynamic adjustment formula of exploration and exploitation weights: (11) (12) (13) where, are exploration weight and exploitation weight, is the entropy value at the tth generation, used to represent population diversity; is the initial entropy value (the maximum entropy value in the initial stage of the population); is the minimum entropy value (the minimum entropy value in the later optimization stage of the population); T is the maximum number of iterations; t is the current iteration number; is the exploration weight of the tth iteration; is the exploitation weight of the tth iteration.
[0044] Entropy value decreases gradually with the increase of iteration number t, in the initial stage , the population diversity is larger, and the randomness is stronger, which is helpful for exploration, in the later stage , the population diversity is smaller, focusing on the local optimal solution, which is helpful for exploitation.
[0045] Exploration weight is positively related to the current entropy value, which is larger in the early stage and gradually decreases in the later stage. Exploitation weight is complementary to exploration weight, which is smaller in the early stage and gradually increases in the later stage.
[0046] The algorithm is initially exploratory, increasing the global search ability of the population by increasing randomness; in the later stage, it is mainly development, improving the accuracy of local search through centralized optimization. This dynamic adjustment mechanism can effectively balance exploration and development, avoid the algorithm falling into local optimal solution, and improve the global optimization performance.
[0047] In combination with the multi-artificial bee colony algorithm, multiple objectives are optimized simultaneously in the population. A reasonable balance and optimization are achieved among computational efficiency, storage bandwidth utilization, and energy consumption.
[0048] 5. Joint optimization mechanism In the scheduling strategy generation stage, the deep learning model quickly generates an initial number of task-resource allocation schemes according to the system state, serving as the starting point for the optimization process. Subsequently, a multi-bee colony algorithm is introduced to search and optimize the scheduling strategy in the global range, overcoming the problem of local optimal solution that the deep model may fall into. This two-level optimization framework composed of deep reinforcement learning and multi-bee colony algorithm combines the efficient representation ability of deep learning and the global search ability of swarm intelligence algorithm, effectively improving the overall efficiency and global optimization level of the scheduling strategy.
[0049] To improve the intelligence and convergence performance of the multi-bee colony algorithm in the global search process, the Deep Q Learning (DQL) mechanism is embedded in the behavior decision-making process of the bee colony individuals. DQL provides strategy guidance for bee colony individuals based on the current task load state and resource usage information, optimizes their search direction and jumping actions, thereby improving the efficiency and stability of scheduling optimization, and realizing the dynamic scheduling driven by the cooperation of swarm intelligence optimization and reinforcement learning strategy. The initialization of the multi-bee colony algorithm embedded with DQL is as follows: 1) Q-value function : represents the cumulative reward expectation value that can be obtained when selecting action a in state s; 2) State s: the current environmental information of the bee colony individual, including individual position, position distribution of neighbor individuals, quality of current solution, etc. 3) Action a: the action that the bee colony individual can take, including: randomly searching for a new solution (exploration action); approaching high-quality solutions (development action); exchanging information with neighbors; 4) Reward r: the immediate feedback obtained by the individual after taking a certain action, such as the degree of quality improvement of the new solution; 5) Action selection strategy: ε-greedy strategy is adopted, with probability ε selecting a random action (exploration) and probability 1-ε selecting the current optimal action (development).
[0050] In the above initialization parameters, the deep Q learning is combined with the artificial bee colony, which is reflected in the following points: The update rule of deep Q learning is used, and the Q value update formula is: (14) wherein: is the current state; is the current selected action; is the immediate reward of the current action; is the new state after performing the action; is the learning rate, controlling the update step size; is the discount factor, balancing the current reward and future reward.
[0051] The action is selected using an epsilon-greedy policy, wherein the current optimal action is selected according to the cumulative reward expectation represented by the Q value function: (15) reward value is defined according to the action effect of the individual in the bee swarm (16) wherein, is the improvement value of the objective function between the new solution and the old solution; is a penalty term, preventing the individual from wasting computational resources in invalid regions.
[0052] In deep Q learning (DQL), a deep neural network (DNN) is used to represent the Q function, and the Q(s, a) function is approximated by a deep neural network, with the input being the state s and the output being the Q value of each action a.
[0053] Deep Q learning dynamically adjusts the action selection strategy of the individual through state-action mapping, achieving an effective balance between exploration and development; deep learning introduces environmental perception capabilities, enabling the bee swarm algorithm to better avoid local optimal traps; the individual dynamically selects the optimal action based on the environment and historical experience, improving the efficiency of the algorithm and achieving a balance between exploration and development.
[0054] Embodiment 2 In an embodiment of the present disclosure, a task scheduling system based on energy consumption optimization is provided, comprising: a demand acquisition module configured to acquire dynamic task scheduling demand, including a task set to be scheduled, a heterogeneous computing resource set, a task load characteristic matrix, and a computing resource capability matrix; a strategy optimization module configured to, based on the dynamic task scheduling demand, utilize a scheduling model that combines deep reinforcement learning and a multi-bee swarm algorithm, and according to the task load and resource capability, perform multi-objective optimization on the task-resource allocation strategy with the objectives of computing efficiency, storage bandwidth utilization, and energy consumption, to obtain an optimal solution. The scheduling model is a multi-objective model with calculation efficiency, storage bandwidth utilization and energy consumption, generates initial task-resource allocation strategies through deep reinforcement learning, and performs global optimization on the initial task-resource allocation strategies through a multi-swarm optimization algorithm to obtain an optimal task-resource allocation strategy.
[0055] Embodiment 3 In an embodiment of the present disclosure, a computer program product is provided, comprising a computer program which, when executed by a processor, implements the task scheduling method based on energy consumption optimization.
[0056] Embodiment 4 In an embodiment of the present disclosure, a non-transitory computer-readable storage medium is provided for storing computer instructions, which, when executed by a processor, implement the task scheduling method based on energy consumption optimization.
[0057] Embodiment 5 In an embodiment of the present disclosure, an electronic device is provided, comprising a processor, a memory, and a computer program; wherein the processor is connected with the memory, and the computer program is stored in the memory; when the electronic device is running, the processor executes the computer program stored in the memory, so that the electronic device implements the task scheduling method based on energy consumption optimization.
[0058] The present disclosure is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams can be implemented by computer program instructions. These computer program instructions can be provided to a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to produce a machine, so that the instructions executed by the computer or other programmable data processing devices produce a device that implements the functions specified in the flowcharts and / or block diagrams. Figure 1 The functions specified in one or more flows and / or blocks Figure 1 The functions specified in one or more flows and / or blocks
[0059] These computer program instructions can also be loaded onto a computer or other programmable data processing device to cause a series of operational steps to be performed on the computer or other programmable data processing device to produce a computer-implemented process, so that the instructions executed by the computer or other programmable data processing device provide a process for implementing the functions specified in the flowcharts and / or block diagrams. Figure 1 The functions specified in one or more flows and / or blocks Figure 1 The functions specified in one or more flows and / or blocks
[0060] The specific embodiments of the present disclosure are described above with reference to the accompanying drawings, but are not intended to limit the protection scope of the present disclosure, and those skilled in the art should understand that various modifications or changes made on the basis of the technical solutions of the present disclosure without creative labor are still within the protection scope of the present disclosure.
Claims
1. A task scheduling method based on energy consumption optimization, characterized in that, include: Obtain dynamic task scheduling requirements, including the set of tasks to be scheduled, the set of heterogeneous computing resources, the task load characteristic matrix, and the computing resource capacity matrix; Based on the dynamic task scheduling requirements, a scheduling model integrating deep reinforcement learning and multi-bee colony algorithm is used. Based on task load and resource capacity, the task-resource allocation strategy is optimized in multiple objectives, with computational efficiency, storage bandwidth utilization and energy consumption as objectives, to obtain the optimal solution. The scheduling model generates several initial task-resource allocation strategies through deep reinforcement learning, and then uses a multi-bee colony optimization algorithm to globally optimize these initial strategies to obtain the optimal task-resource allocation strategy.
2. The task scheduling method based on energy consumption optimization as described in claim 1, characterized in that, The task set includes several tasks; The heterogeneous computing resource set includes several heterogeneous AI chip types; The task load characteristic matrix represents the load vector of different tasks on different heterogeneous computing resources; The computing resource capability matrix represents the computing power, bandwidth, memory limit, and power consumption limit of different heterogeneous computing resources.
3. The task scheduling method based on energy consumption optimization as described in claim 1, characterized in that, The objective function, which considers computational efficiency, storage bandwidth utilization, and energy consumption as multiple objectives, is expressed as follows: in, For task-resource allocation strategy, Let these represent the objective functions for computational efficiency, storage bandwidth utilization, and energy consumption, respectively. These are the weights of the objective functions for computational efficiency, storage bandwidth utilization, and energy consumption, respectively. , , These represent the computational efficiency and the minimum and maximum values of the computational efficiency, respectively. , , These represent the storage bandwidth utilization rate and the minimum and maximum values of the storage bandwidth utilization rate, respectively. , , These represent energy consumption and the minimum and maximum energy consumption, respectively.
4. The task scheduling method based on energy consumption optimization as described in claim 3, characterized in that, The process of generating initial task-resource allocation strategies through deep reinforcement learning involves using resource status and the characteristics of tasks to be scheduled as states, task-to-resource allocation decisions as actions, and a comprehensive optimization objective function as a reward function. The process employs a multi-agent deep reinforcement learning method to learn the task-resource allocation strategies.
5. The task scheduling method based on energy consumption optimization as described in claim 1, characterized in that, The multi-bee colony optimization algorithm introduces an entropy-based dynamic adjustment mechanism to dynamically adjust the weights of exploration and development in the multi-bee colony optimization algorithm.
6. The task scheduling method based on energy consumption optimization as described in claim 3, characterized in that, The process of globally optimizing the initial task-resource allocation strategies specifically involves: Each individual bee in the colony corresponds to a task-resource allocation scheme. Its actions are guided by heuristic rules during the search process. By combining exploration and development mechanisms, it balances global search and local optimization. The quality of individual solutions is measured by a comprehensive optimization objective function, and the overall optimal scheduling strategy is gradually approached.
7. A task scheduling system based on energy consumption optimization, characterized in that, include: The requirement acquisition module is configured to acquire dynamic task scheduling requirements, including the set of tasks to be scheduled, the set of heterogeneous computing resources, the task load characteristic matrix, and the computing resource capability matrix. The strategy optimization module is configured to: based on dynamic task scheduling requirements, utilize a scheduling model that integrates deep reinforcement learning and multi-bee colony algorithm, and optimize the task-resource allocation strategy based on task load and resource capabilities, with computational efficiency, storage bandwidth utilization and energy consumption as objectives, to obtain the optimal solution; The scheduling model takes computational efficiency, storage bandwidth utilization and energy consumption as multiple objectives. It generates several initial task-resource allocation strategies through deep reinforcement learning, and uses an improved multi-bee colony optimization algorithm to globally optimize the initial task-resource allocation strategies to obtain the optimal task-resource allocation strategy.
8. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the energy-efficient task scheduling method according to any one of claims 1-6.
9. A non-transitory computer-readable storage medium, characterized in that, The non-transitory computer-readable storage medium is used to store computer instructions, which, when executed by a processor, implement a task scheduling method based on energy consumption optimization as described in any one of claims 1-6.
10. An electronic device, characterized in that, include: The device includes a processor, a memory, and a computer program; wherein the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device is running, the processor executes the computer program stored in the memory to enable the electronic device to perform a task scheduling method based on energy consumption optimization as described in any one of claims 1-6.