Intelligent computing resource allocation method based on reinforcement learning

Through the intelligent computing resource allocation method based on reinforcement learning, the problems of inflexible resource allocation and unoptimized task scheduling in the existing technology are solved, dynamic scheduling of resources and multi-objective optimization are realized, and the efficiency and resource utilization of the system are improved.

CN120104323AInactive Publication Date: 2025-06-06天津华信惠悦科技有限公司
View PDF 0 Cites 7 Cited by

Patent Information

Application Number
CN202510174567.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-18
Publication Date
2025-06-06
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing intelligent computing resource allocation methods are insufficient in handling complex tasks and multi-task environments, resulting in low resource utilization, inflexible task scheduling and poor system performance. Especially when load changes and task requirements change, the resource allocation strategy cannot be quickly adjusted.

Method used

Using intelligent computing resource allocation method based on reinforcement learning, a resource scheduling decision-making mechanism with multi-strategy collaborative optimization is realized through intelligent resource state modeling and adaptive state space design. This method includes deep autoencoder for data dimensionality reduction, multi-task learning predicts resource requirements, generating adversarial networks to generate multiple resource scheduling strategies, distributed Q-learning scheduling decisions, deep reinforcement learning dynamic evaluation of task priorities, and optimizing resource allocation and scheduling strategies through Bayesian optimization and genetic algorithms.

Benefits of technology

Dynamic scheduling of intelligent computing resources is realized, and resource allocation strategies are automatically adjusted according to real-time task requirements and system load, improving the efficient operation of the system under different loads and task types, reducing resource waste, and improving response speed and flexibility. At the same time, through multi-objective optimization and Pareto cutting-edge optimization, we balance task completion time, resource utilization rate and energy consumption, and improve resource utilization rate and overall system performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120104323A_ABST
    Figure CN120104323A_ABST
Patent Text Reader

Abstract

The invention relates to an intelligent computing resource allocation method based on reinforcement learning. According to the technical scheme, the method comprises the steps that monitoring data from different sources including hardware monitoring, software logs and network bandwidths are fused, and a multi-dimensional time sequence state space is formed; carrying out dimensionality reduction and de-noising processing on the multi-modal data by using a depth auto-encoder; multi-task learning MTL is introduced into time sequence modeling, and resource requirements and state evolution of multiple tasks are predicted at the same time; generating a plurality of predictive resource scheduling strategies by using a GAN (Generative Adversarial Network), and dynamically selecting a strategy scheme when a load changes; each strategy is realized by an independent sub-network, and part of core knowledge is shared; through a strategy evolution mechanism, according to a historical feedback optimization strategy combination including task completion time and resource consumption, in a heterogeneous resource environment including a plurality of cloud computing platforms and edge computing nodes, based on difference of resource types, fine-grained scheduling of strategies is carried out; and scheduling the decision by using a distributed Q-learning mechanism in the reinforcement learning model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an intelligent computing resource configuration method, and more specifically to an intelligent computing resource configuration method based on reinforcement learning. Background Art

[0002] At present, in the intelligent computing resource configuration methods, many traditional methods rely on static configuration, regular scheduling and optimization strategies based on simple rules. In particular, some resource configuration methods such as China Invention Patent No. 2020111342760 have obvious deficiencies and drawbacks in dealing with complex tasks and multi-task environments. These problems are particularly prominent in modern intelligent computing resource configuration, limiting the maximization of resource utilization, the flexibility of task scheduling and the overall performance of the system.

[0003] First, most existing methods rely on fixed operator configurations and static rules based on task reservations for resource allocation. This method usually performs static resource pre-allocation before the task arrives, ignoring the dynamic changes of tasks and the real-time nature of system load. Therefore, when the load changes, task requirements change, or the system state fluctuates, the existing methods often cannot quickly adjust the resource allocation strategy. For example, the resource allocation method in the patent configures resources for each operator in the task and determines the resource requirements based on the current task request and server status. However, when the task complexity increases or the task type changes, the system still cannot dynamically adjust the resource allocation. This static resource allocation strategy may lead to inefficient use of resources, or even over-configuration or resource shortages, especially in high-load or multi-task concurrent environments, the system cannot flexibly respond to task requirements, resulting in low task execution efficiency and resource waste. Secondly, the current method lacks an effective task scheduling and priority management mechanism when facing multi-tasks and complex data processing models. When faced with multiple task requests, the patented method uses a sub-method of collaborative resource allocation to handle the resource allocation problem between multiple tasks. However, this type of method often allocates tasks based on fixed strategies and lacks real-time perception and adaptive capabilities for dynamic factors such as different task priorities, delay sensitivity, and resource consumption. When multiple tasks are concurrent, the resource requirements of tasks may change dramatically. How to dynamically adjust resource allocation according to the actual needs of tasks, optimize the scheduling order of tasks, and avoid excessive competition for resources are problems that have not been effectively solved in current methods. This static or semi-static priority configuration can easily lead to high-priority tasks not getting enough resources, while low-priority tasks occupy too many resources, ultimately affecting the overall operating efficiency of the system and the timely completion of tasks.

[0004] In addition, the resource allocation and scheduling strategies of existing methods lack intelligent decision-making support, rely more on static decisions of rules and algorithms, and lack self-learning and adaptive mechanisms based on real-time tasks and system feedback. For example, the step of "obtaining the second task and performing collaborative resource allocation" mentioned in the patent mainly relies on traditional resource allocation methods, and is unable to perform fine-grained analysis of the load of different tasks, and lacks the ability to provide feedback on dynamic changes during task execution (such as delays, resource usage, task completion time, etc.). With the increase in task complexity, single rules and static scheduling methods have gradually become less flexible and are not adapted to the needs of modern large-scale intelligent computing tasks, especially in heterogeneous systems such as edge computing and multi-cloud environments. The changes in task requirements are very complex, and it is difficult for traditional methods to respond quickly and accurately.

[0005] In addition, the existing methods are also relatively limited in their ability to optimize the overall resource utilization efficiency of the system. In traditional resource allocation methods, only the execution requirements of the task are often considered, while the global performance optimization of the system is ignored. When dealing with multiple tasks executing concurrently, how to avoid excessive resource consumption and energy waste of the system while ensuring the efficient completion of the tasks remains a difficult problem to solve. The resource allocation mechanism in the patented method often relies on the resource usage of each operator or task for allocation, ignoring the dynamic collaboration and optimization of different resource types (such as computing resources, storage resources, bandwidth, etc.). For example, when certain computing tasks have a large demand for bandwidth or storage, the system should be able to dynamically adjust according to the current distribution of resources, but the existing methods fail to take into account the fine-grained dynamic allocation of resources, which can easily lead to excessive use of some resources while other resources are not fully utilized. Summary of the invention

[0006] The purpose of the present invention is to provide an intelligent computing resource configuration method based on reinforcement learning, so as to solve some of the drawbacks and shortcomings pointed out in the background technology.

[0007] The present invention solves the above-mentioned technical problems by adopting the following technical solution, which includes the following steps:

[0008] S1. Intelligent resource state modeling and adaptive state space design:

[0009] S1.1. Fusion of monitoring data from different sources including hardware monitoring, software logs, and network bandwidth to form a multi-dimensional time series state space; and use of deep autoencoders to reduce the dimensionality and denoise the multimodal data;

[0010] S1.2, introduce multi-task learning (MTL) in time series modeling to simultaneously predict the resource requirements and state evolution of multiple tasks;

[0011] S2. Resource scheduling decision-making mechanism for multi-strategy collaborative optimization:

[0012] S2.1. Generate multiple predicted resource scheduling strategies using generative adversarial networks (GANs) and dynamically select strategy solutions when the load changes.

[0013] S2.2. Each strategy is implemented by an independent sub-network, sharing some core knowledge. Through the strategy evolution mechanism, the strategy combination is optimized according to historical feedback including task completion time and resource consumption.

[0014] S2.3. In a heterogeneous resource environment including multiple cloud computing platforms and edge computing nodes, fine-grained scheduling of strategies based on differences in resource types; scheduling decisions using the distributed Q-learning mechanism in the reinforcement learning model;

[0015] S3. Resource allocation and dynamic priority adjustment based on multi-task learning:

[0016] S3.1. Dynamically evaluate the priority of each task through deep reinforcement learning. The reward function includes the task completion time, as well as the delay sensitivity and resource consumption history of the task execution to calculate the priority of the predicted task under the current system load.

[0017] S3.2. Adjust the priorities between tasks through Bayesian optimization, introduce the load perception mechanism in reinforcement learning, intelligently identify the load characteristics of different types of tasks, and dynamically adjust resource allocation;

[0018] S4. Evolution-based reward adjustment and realization of long-term optimization goals:

[0019] S4.1. Collect the execution data of various tasks during system operation, including execution time and resource utilization, to build a multi-objective optimization model;

[0020] S4.2, use genetic algorithm GA to adjust the weight of each goal and dynamically optimize the goal combination according to real-time task requirements and historical performance;

[0021] S4.3. Make trade-offs between different objectives through Pareto frontier optimization.

[0022] Furthermore, the resource scheduling decision-making mechanism method of multi-strategy collaborative optimization is:

[0023] Generator, used to generate multiple candidate resource scheduling strategies based on the current system status information including CPU load, memory usage, task type, and network bandwidth, where each candidate strategy is represented as a scheduling solution θ i , and the generator is trained through a reinforcement learning RL network to generate multiple optional strategy combinations;

[0024] The discriminator is used to evaluate the effectiveness of each candidate strategy based on the performance evaluation function F(θ i ) measures the performance of each candidate strategy, where F(θ i ) is defined as a weighted function of multiple optimization objectives including task completion time, resource consumption, and system stability:

[0025] F(θ i )=αT(θ i )+βC(θ i )+γS(θ i )

[0026] Among them, T(θ i ) is the task completion time, representing the strategy θ i The time required to complete the task; C(θ i ) is the resource consumption, indicating the strategy θ i The computing resources consumed when executing tasks; S(θ i ) is the system stability, indicating the strategy θ i The stability of the system operation, including failure rate and response delay; α, β, γ are the weighting coefficients required for scheduling task priorities, which determine the degree of each indicator in the overall evaluation.

[0027] Furthermore, the resource scheduling decision-making mechanism method of multi-strategy collaborative optimization is:

[0028] By monitoring the system load changes in real time, the generator generates multiple scheduling strategy candidate solutions θ i , the discriminator selects the optimal strategy θ in real time based on the current load situation opt ; Adjust strategy according to load perception θ i , through the strategy selection function Select a resource scheduling strategy, where G(θ i ) is the performance evaluation function of the resource scheduling strategy, defined as:

[0029]

[0030] Among them, G(θ i ) represents the strategy θ i The total benefit is the performance indicator of the strategy selection under load changes; R(t) is the change of resource utilization over time t, indicating that the strategy θ i The proportion of resources used at time t; P(t) is the change of power consumption over time t, which represents the strategy θ i The power or energy consumed at time t; η 1 ,η 2 is the weighted coefficient between resource utilization and power consumption, reflecting the different requirements of the system for resource utilization and energy consumption.1 >η 2 , to prioritize the effective use of computing resources; T is the evaluation time period, which is the time window for scheduling decisions.

[0031] Furthermore, the resource scheduling decision-making mechanism method of multi-strategy collaborative optimization is:

[0032] Each generated strategy is implemented by an independent sub-network, and different sub-networks share core knowledge. Each candidate strategy is optimized through an evolutionary algorithm, and the strategy is updated through historical task execution data and feedback information. The specific calculation is performed through the evolution function:

[0033]

[0034] Among them, E(θ i ) is the performance evaluation of the evolved strategy, indicating that strategy θ i The effect after historical feedback optimization; θ j is the parameter value in the historical strategy, indicating the parameters of the strategy executed by each task in the history; θ max is the upper limit prediction strategy value, indicating the upper limit of the performance indicator in the strategy; δ j is the weight factor for strategy adjustment, reflecting the influence of historical tasks on the current strategy optimization during the evolution process; N is the total amount of feedback historical data, the size of the data set used for strategy evolution, and indicates the number of tasks executed in history.

[0035] Furthermore, the method for achieving the evolution-based reward adjustment and long-term optimization goal is:

[0036] First, a multi-objective optimization model is constructed; different objective functions are combined by weighted sum; the balance point is found by adjusting the weights between the objective functions so that the conflicts between the various objectives are resolved; and the optimization problem is expressed by the following formula:

[0037]

[0038] in:

[0039] Φ(W,X) is the overall optimization objective function, which represents the comprehensive performance of task scheduling. The performance of the system is measured by the weighted sum of each objective function. i is the weight of the i-th objective, which determines the importance of the objective in the optimization; f i (X) is the i-th objective function, which represents the actual performance of the i-th performance indicator in task scheduling; X is the system state vector, which contains information related to resource usage and load status of the current task, reflecting the real-time status of the system; n is the number of objective functions, including task completion time, resource utilization rate, and energy consumption.

[0040] Furthermore, the method for achieving the evolution-based reward adjustment and long-term optimization goal is:

[0041] Genetic algorithm GA is introduced to dynamically adjust the target weight; through continuous iterative optimization, the target weight configuration of the current task requirements and system status is found; fitness evaluation is a key part of the genetic algorithm to measure the pros and cons of each target weight combination; the fitness evaluation formula is as follows:

[0042]

[0043] in:

[0044] A(W i ,X) is the fitness value, indicating the target combination W i Performance under the current system state X; the higher the fitness value, the better the target combination; α j is the importance coefficient of target j, reflecting the contribution of different targets to system optimization; f j (X) is the value of the jth objective function, which represents the actual performance of the objective under the given system state X; f j (X tre ) is the target value in an ideal state, which means the target that should be achieved without any resource limitation or delay.

[0045] Furthermore, the method for achieving the evolution-based reward adjustment and long-term optimization goal is:

[0046] Pareto frontier optimization is introduced to balance multiple objectives to find a solution that is optimal in one objective and not inferior in other objectives. In multi-objective optimization problems, the goal of Pareto frontier optimization is to calculate multiple solutions and find a solution set that cannot further improve the objective without sacrificing other objectives. The formula for this process is as follows:

[0047]

[0048] in:

[0049] P(W,X) is the Pareto frontier solution set, which represents the optimal trade-off between different objectives. Each solution in the solution set represents a weight configuration that is superior to other solutions in some objectives. i is the target weight combination corresponding to the i-th solution; through Pareto frontier optimization, the best weight combination will be selected; f j (X i ) represents the performance of the i-th solution under target j; f j (X j ) is the performance of other solutions under target j; when f j (X i)≤f j (X j ), then W i At least not inferior in terms of target j.

[0050] The intelligent computing resource allocation method of the present invention has significant beneficial effects, which are mainly reflected in the following aspects:

[0051] Dynamic scheduling of intelligent computing resources is achieved through reinforcement learning models (such as Q-learning and multi-strategy collaborative optimization mechanisms). This method can automatically adjust the resource allocation strategy according to real-time task requirements and system load conditions to ensure that the system can maintain efficient operation under different loads and task types. Compared with traditional static resource allocation methods, reinforcement learning can adapt to environmental changes in real time, reduce resource waste, and improve the response speed and flexibility of the system.

[0052] The multi-objective optimization method can find the optimal balance point among multiple conflicting objectives such as task completion time, resource utilization, and energy consumption. By introducing Pareto frontier optimization and weighted sum strategy, the system can effectively balance multiple objectives and avoid performance bottlenecks caused by ignoring a certain objective in traditional methods. This not only improves the efficiency of task execution, but also optimizes resource utilization and energy consumption, further reducing the operating costs of the system.

[0053] The weights of optimization objectives are dynamically adjusted through evolutionary methods such as genetic algorithms, and the best resource scheduling strategy that adapts to the current task requirements and system status is gradually found. The introduction of genetic algorithms enables the system to continuously improve the strategy configuration through historical data feedback and improve stability and performance in long-term operation. This feedback-based strategy optimization mechanism enables the system to cope with the scheduling needs of long-term tasks and respond quickly and accurately to different load changes. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] Figure 1 This is a flow chart of the intelligent computing resource configuration method based on reinforcement learning of the present invention.

[0055] Figure 2 This is a flow chart of the resource scheduling decision-making mechanism method for multi-strategy collaborative optimization of the present invention.

[0056] Figure 3 This is a flow chart of the method for implementing the evolution-based reward adjustment and long-term optimization goals of the present invention. DETAILED DESCRIPTION

[0057] The specific implementation modes of the present invention will be described in detail below in conjunction with the accompanying drawings.

[0058] Combined with Figure 1As shown in the process, the intelligent computing resource configuration method based on reinforcement learning has key technical innovations in intelligent resource state modeling and adaptive state space design. First, the method integrates monitoring data from different sources, including hardware status (such as CPU load, memory usage), software logs (such as task execution information, application behavior), network bandwidth usage, etc. These data provide the system with a multi-dimensional, time-series state space that can fully reflect the current state of the system and its evolution trend. However, the original monitoring data usually contains a lot of redundant information and noise, so it must be effectively processed.

[0059] To this end, a deep autoencoder is used to reduce the dimensionality and denoise these multimodal data. The deep autoencoder compresses the high-dimensional input data into a low-dimensional space through nonlinear mapping, captures the most important features, and removes unnecessary redundant information and noise, thereby enhancing the validity and usability of the data. This data processing method can reduce the computational burden of the system and improve the efficiency of subsequent model training. In addition, the output of the deep autoencoder not only retains the key information of the original data, but also provides a more refined state representation for the reinforcement learning model, which is crucial for accurately judging resource allocation decisions.

[0060] Next, in terms of time series modeling, in order to fully capture the dynamic changes in resource requirements and the long-term evolution of the state, the multi-task learning (MTL) method was introduced. Multi-task learning refers to learning multiple related tasks in the same model at the same time. It improves the generalization ability of the model by sharing knowledge between different tasks. In this method, multi-task learning not only helps to predict the resource requirements of different tasks, but also can jointly model the usage of various resources in the system. Specifically, MTL can handle the resource status prediction problem of multiple tasks, such as simultaneously predicting the CPU usage, memory consumption, network bandwidth requirements of different applications or services, avoiding the complexity of building a separate prediction model for each task. Under the framework of multi-task learning, the goal of each task is to maximize the efficiency of its resource use, and the relationships and resource requirements of different tasks are learned through shared neural network layers. In this way, the system can effectively capture the potential dependencies and correlations in resource configuration.

[0061] Generate multiple predicted resource scheduling strategies through generative adversarial networks (GAN). Generative adversarial networks consist of two main parts: generator and discriminator. In this process, the generator generates multiple candidate resource scheduling strategies based on the current system state (such as resource load, task priority, etc.), while the discriminator evaluates the effectiveness of these strategies and determines which strategies can optimize resource utilization, reduce task completion time, reduce energy consumption and other goals under specific conditions. In this way, the generator can not only generate a single strategy, but also provide multiple possible scheduling decisions, so that the system can choose the best solution under different load conditions and task requirements. As the system load changes, the scheduling strategy can also be dynamically adjusted based on the feedback of the discriminator. The system does not rely on preset fixed rules or a single strategy, but intelligently switches strategies according to the real-time load situation to achieve efficient resource utilization and flexibility in task scheduling. This multi-strategy generation method based on GAN has stronger adaptability than the traditional single static scheduling strategy method, and can cope with more complex and changing computing environments and load fluctuations.

[0062] Each scheduling strategy is implemented by an independent sub-network, and these sub-networks share some core knowledge. Specifically, the shared part between sub-networks is usually achieved by sharing the weights or parameters of certain network layers, so that each strategy can utilize the same system knowledge framework. This sharing mechanism ensures that different sub-strategies are not "out of touch" when solving problems and can complement each other. Although each sub-strategy has its own network and goals, the core knowledge they share can improve the learning efficiency of each sub-strategy and promote the collaborative optimization of the system when performing different scheduling tasks. In addition, the shared knowledge between strategies can accelerate the learning process and reduce the waste of computing resources.

[0063] The strategy evolution mechanism plays a vital role in this approach. Evolutionary algorithms, such as genetic algorithms (GA) or other heuristic search methods, are used to optimize the combination of different strategies. The historical performance of each strategy (such as task completion time, resource consumption, etc.) serves as feedback information to help evaluate the pros and cons of the strategy. Through the evolutionary mechanism, the system can continuously adjust and optimize the combination of strategies based on this feedback information. Over time, the combination of strategies can gradually tend to the optimal solution, so that it can better cope with changing load conditions and resource requirements. The evolutionary mechanism not only enhances the adaptability of the scheduling strategy, but also improves the stability and long-term efficiency of the overall system performance.

[0064] In heterogeneous resource environments, including multi-cloud computing platforms, edge computing nodes, etc., differences in resource types require fine-grained policy scheduling. Due to the obvious differences in computing power, storage capacity, bandwidth and other characteristics of different resources, a single policy may not be able to cope with the scheduling requirements of all resource types. In this environment, fine-grained policy scheduling based on differences in resource types becomes particularly important. The system will select the most appropriate policy for scheduling based on the characteristics of each resource, thereby maximizing resource utilization efficiency and reducing resource conflicts. For example, for a cloud computing platform with strong computing power, the system may choose a computing-intensive task scheduling policy, while for an edge computing node with limited bandwidth, it may choose a bandwidth-friendly task allocation policy. The distributed Q-learning mechanism in reinforcement learning is introduced for scheduling decisions. Q-learning is a model-free method in reinforcement learning that can continuously learn the optimal policy through interaction with the environment. In distributed Q-learning, multiple learning agents (such as different computing nodes) learn their own optimal scheduling policies respectively and coordinate decisions in a distributed manner. Each agent adjusts the scheduling policy according to the resource node and task load situation where it is located, and the global decision of the system is achieved through the collective behavior of different agents. Distributed Q-learning can efficiently solve resource scheduling problems in large-scale systems, avoid bottlenecks that may be encountered in centralized learning, and improve the scalability and flexibility of the system.

[0065] The key to reinforcement learning is to guide the agent to learn the optimal behavior strategy through the reward mechanism. Here, the reward function not only takes into account the completion time of the task, but also introduces the delay sensitivity and resource consumption history in the task execution to predict the priority of the task under the current system load. Specifically, the task completion time is an important indicator to measure whether the task can be completed efficiently, the delay sensitivity reflects the tolerance of the task to the execution delay, and the resource consumption history reveals the performance of the task in terms of resource usage. By integrating these factors into the reward function, the system can evaluate the priority of each task and dynamically adjust the resource allocation strategy according to the current load status. Deep reinforcement learning can effectively respond to load changes and optimize task scheduling by continuously adjusting the evaluation mechanism of task priority to reduce the total time of task execution, improve the overall utilization efficiency of resources, and reduce the energy consumption of the system. Next, Bayesian optimization is introduced into the process of task priority adjustment. Bayesian optimization is a global optimization method based on a probabilistic model, which is suitable for handling task scheduling problems with large resource consumption or slow feedback. In this method, Bayesian optimization iteratively updates the priority adjustment strategy to reduce the computational overhead in the exploration and optimization process. Bayesian optimization can continuously adjust the priority allocation between tasks based on historical data, thereby ensuring that the scheduling between different tasks can meet real-time resource needs and performance requirements, and avoid over-scheduling or resource waste. In the process of task priority adjustment, by introducing the load perception mechanism in reinforcement learning, the system can intelligently identify the load characteristics of different types of tasks. For example, for computationally intensive tasks, the system will recognize their high demand for computing resources and allocate them to resources with stronger computing power based on this feature; for data-intensive tasks, more bandwidth resources will be allocated first.

[0066] The objectives of a multi-objective optimization model usually involve multiple aspects, such as task completion time, resource utilization efficiency, energy efficiency, etc. These objectives are often conflicting, so effective methods need to be used to balance and optimize them. Task execution data not only contains the performance of the current task, but also records the historical execution of the task, thus providing an important basis for subsequent resource allocation and optimization decisions. By building a multi-objective optimization model, the system can make more accurate resource allocation decisions based on the priorities, resource requirements and other characteristics of different tasks, thereby improving resource utilization efficiency and task execution efficiency. In this process, a genetic algorithm (GA) is used to adjust the weight of each objective and dynamically optimize the objective combination based on real-time task requirements and historical performance. The genetic algorithm simulates the process of natural selection and finds the global optimal solution through the variation and crossover of the objective weights of individuals in the population. First, the genetic algorithm randomly generates a set of initial objective weight combinations, and then evaluates the fitness value of each objective combination (i.e., resource efficiency, task completion time and other indicators under different combinations) based on the execution data of the task. Objective combinations with higher fitness values ​​will be retained in the population, while those with lower fitness values ​​will be eliminated.

[0067] Through crossover operations, the genetic algorithm generates new combinations of target weights, while the mutation operation increases the diversity of the search space by randomly changing the weight values. After multiple generations of iteration and optimization, the genetic algorithm can find the target weight configuration that best suits the current system and task requirements. The introduction of the genetic algorithm enables the system to dynamically adapt to different environmental conditions such as load changes and task requirements, thereby optimizing the resource allocation strategy and ensuring the long-term and efficient operation of the system. At the same time, Pareto frontier optimization, as one of the core technologies in this method, can help the system make trade-offs between different objectives. Pareto frontier is an important concept in multi-objective optimization problems. It represents the optimal solution set that cannot further improve a certain objective without compromising other objectives under different objectives. In this method, the system will find the Pareto optimal solution based on the solution set under different objective combinations.

[0068] Embodiment 1:

[0069] Combined with Figure 2 As shown in the process, in a certain intelligent data processing platform, three types of tasks are set to be scheduled: data analysis tasks, deep learning training tasks, and video rendering tasks. Each task has different resource requirements. For example, data analysis tasks require a large amount of CPU and storage space, deep learning training tasks require a large amount of GPU resources and high-bandwidth network connections, and video rendering tasks have high requirements for CPU and memory, and also have high requirements for system stability, because long-term processing may cause system failures.

[0070] Step 1: The generator generates multiple candidate scheduling strategies

[0071] Based on the current system state (e.g., CPU load, memory usage, task type, network bandwidth), the generator (a reinforcement learning model) generates multiple candidate scheduling strategies based on this information. Specifically, the system generates strategies based on the following factors:

[0072] 1. CPU load: Set the CPU load of the current platform to 80%. The generator needs to consider how to allocate tasks to avoid CPU overload. 2. Memory usage: Set the current memory usage to 70%. The system needs to adjust the memory allocation according to the memory requirements of the task. 3. Task type: For example, data analysis tasks have higher requirements on the CPU, while deep learning training tasks have higher requirements on the GPU. 4. Network bandwidth: Set the network bandwidth of the current platform to 500Mbps. The generator needs to consider allocating network bandwidth to tasks that require large amounts of data transmission.

[0073] Based on this information, the generator may generate multiple scheduling strategies. For example, strategy 1 (θ 1 ) may allocate more CPU resources to data analysis tasks, strategy 2 (θ 2 ) may prioritize GPU resources for deep learning tasks, while strategy 3 (θ 3 ) may prioritize bandwidth allocation to video rendering tasks based on the current network bandwidth.

[0074] Step 2: Discriminator evaluates candidate strategies

[0075] After the generator generates multiple candidate strategies, the discriminator needs to evaluate the effectiveness of these strategies. The evaluation is done through the performance evaluation function F(θ i ), which takes into account task completion time, resource consumption and system stability.

[0076] Calculation formula:

[0077] F(θ i )=αT(θ i )+βC(θ i )+γS(θ i )

[0078] in:

[0079] T(θ i ): Task completion time, indicating that in strategy θ i The time required to complete the task.

[0080] C(θ i ): Resource consumption, indicating that in strategy θ i The computing resources consumed when executing the task.

[0081] S(θi ): System stability, which means that under the strategy θ i The stability of the system operation, including failure rate, response delay, etc.

[0082] α, β, γ: are weighted coefficients that control the importance of task completion time, resource consumption, and system stability in the overall evaluation.

[0083] Weighting coefficient value:

[0084] α∈[0.3,0.5]

[0085] β∈[0.2,0.4]

[0086] γ∈[0.1,0.3]

[0087] The range of these coefficients is determined by the characteristics of different tasks and the priority settings of the system. For example, some tasks may focus more on completion time, while some tasks may focus more on system stability.

[0088] Example:

[0089] When performing deep learning training tasks, the system evaluation results are as follows:

[0090] Task completion time T(θ 1 ): Using strategy 1, the task completion time is 12 hours. Resource consumption C(θ 1 ): Under strategy 1, the resource consumption is 300 computing units. System stability S(θ 1 ): Under strategy 1, the system stability score is 80 (0-100 points, 80 indicates higher stability).

[0091] If the weighting coefficients are set to: α = 0.4, β = 0.3, γ = 0.3, the evaluation results are:

[0092] F(θ 1 ) = 0.4 × 12 + 0.3 × 300 + 0.3 × 80 = 4.8 + 90 + 24 = 118.8

[0093] Similarly, other strategies can be evaluated, such as strategy 2θ 2 and Strategy 3 (θ 3 ), calculate the corresponding F(θ i ) value. Assuming the evaluation results of strategy 2 and strategy 3 are 130 and 110 respectively, it is obvious that strategy 1 (θ 1 ) has the lowest evaluation result, so the discriminator will select strategy 1 as the optimal strategy.

[0094] Step 3: Dynamically select the optimal strategy

[0095] Through the evaluation of the discriminator, the system will dynamically select the most suitable scheduling strategy based on the current task requirements and historical performance.

[0096] For example, set the current platform load as follows:

[0097] CPU Load: 85%

[0098] Memory usage: 75%

[0099] Network bandwidth: 600Mbps

[0100] The system selects strategy 1 (θ 1 ) to schedule tasks. The generator will continue to optimize the strategy combination through reinforcement learning training so that it can select the optimal strategy more quickly and intelligently when the load changes next time.

[0101] On the intelligent computing platform of this embodiment, the current system is executing multiple tasks simultaneously, including data analysis, deep learning training, and video rendering tasks. System resources include CPU, GPU, memory, storage, and network bandwidth. The load of the platform is constantly changing. For example, some tasks may take up more computing resources at high loads, while they may reduce resource consumption and improve system stability at low loads. In order to cope with these changes, it is necessary to monitor the system load in real time and select the optimal scheduling strategy based on this.

[0102] The generator generates multiple candidate scheduling strategies θ according to the current system load (such as CPU occupancy, memory usage, network bandwidth, etc.) and task requirements. i . Set the current system load as follows:

[0103] CPU Load: 80%

[0104] Memory usage: 70%

[0105] Network bandwidth: 500Mbps

[0106] Task type: Data analysis tasks require a lot of computing resources, deep learning tasks require GPU acceleration, and video rendering tasks have high requirements for CPU and memory.

[0107] Based on this information, the generator may generate three candidate strategies:

[0108] Strategy 1 (θ 1 : Prioritize the allocation of computing resources to data analysis tasks, while maintaining a certain amount of resource redundancy to meet the needs of other tasks.

[0109] Strategy 2 (θ 2: Prioritize GPU resources for deep learning tasks and reduce CPU usage.

[0110] Strategy 3 (θ 3 : Give priority to the resource requirements of video rendering tasks and reduce the load on the CPU and memory.

[0111] Next, the discriminator is based on the performance evaluation function G(θ i ) evaluates the effectiveness of each strategy. The performance evaluation function is defined as:

[0112]

[0113] in:

[0114] R(t): The change in resource utilization over time t, representing the strategy θ i The proportion of resources used at time t.

[0115] P(t): Power consumption changes over time t, representing strategy θ i The power or energy consumed at time t.

[0116] η 1 ,η 2 : Weighted coefficient, used to balance the impact of resource utilization and power consumption. Set η 1 =0.7,η 2 =0.3, indicating that the system pays more attention to resource utilization and gives priority to improving the effective use of resources.

[0117] T: The evaluation period is set as the time window for scheduling decisions, for example, 1 hour.

[0118] In order to evaluate each strategy, we first need to simulate the changes of resource utilization and power consumption over time under each strategy. For example, under strategy 1, the resource utilization and power consumption are set to change over time as follows:

[0119] During the first 30 minutes, resource utilization was 90% and power consumption was 250 watts.

[0120] Over the next 30 minutes, resource utilization dropped to 80 percent, with power consumption at 230 watts.

[0121] In the last 30 minutes, resource utilization was 85% and power consumption was 240 watts.

[0122] Based on these data, the performance evaluation function G(θ) of strategy 1 can be calculated 1 ). Assuming that the power consumption and resource utilization change steadily in each time period, the following results can be obtained:

[0123] Phase 1 (0-30 minutes):

[0124] G 1 =0.7×0.90.3×250=0.6375=-74.37

[0125] Phase 2 (30-60 minutes):

[0126] G 2 =0.7×0.80.3×230=0.5669=-68.44

[0127] Phase 3 (60-90 minutes):

[0128] G 3 =0.7×0.850.3×240=0.59572=-71.405

[0129] Therefore, the total benefit of strategy 1 is:

[0130] G(θ 1 )=-74.37+(-68.44)+(-71.405)=-214.215

[0131] Step 3: Choose the best strategy

[0132] Similarly, strategies 2 and 3 will be evaluated separately to calculate their G(θ 2 ) and G(θ 3 ) value. If the evaluation results of strategy 2 and strategy 3 are -220 and -210 respectively, then obviously strategy 3 will be selected as the optimal strategy θ opt , because it has the best overall benefit. When the load changes, the system adjusts the strategy based on real-time load data. For example, if the load suddenly increases, it may choose to allocate more resources to compute-intensive tasks, while if the load is low, it may prioritize strategies with lower resource consumption.

[0133] This embodiment introduces a resource scheduling decision-making mechanism for multi-strategy collaborative optimization. This mechanism generates multiple candidate strategies and optimizes them in combination with evolutionary algorithms, thereby continuously improving the intelligence level of resource scheduling. In this process, each generated strategy is implemented by an independent sub-network, and core knowledge is shared between different sub-networks to ensure the overall synergy of the system. In order to further optimize the scheduling strategy, historical task execution data and feedback information are introduced, and each candidate strategy is updated through an evolutionary algorithm. In this optimization process, the evolution function is the core, which effectively adjusts the strategy based on historical data and feedback from strategy execution. The form of the evolution function is:

[0134]

[0135] in:

[0136] E(θ i ) represents the performance evaluation of the evolved strategy, and represents the strategy θ i The effect after optimization through historical feedback;

[0137] θ j The parameter values ​​representing the strategy executed by each task in the history;

[0138] θ max is the upper limit of the performance indicator in the strategy, indicating the highest performance that can be achieved in the historical task;

[0139] δ j It is the weight factor of strategy adjustment, which reflects the influence of historical tasks on the current strategy optimization during the evolution process;

[0140] N is the total amount of feedback history data, indicating the number of tasks executed in history.

[0141] In order to better understand the application of this method, this multi-strategy collaborative optimization mechanism is deployed on an intelligent computing platform, where different types of tasks such as data analysis, image processing, and video rendering are running. The goal of the system is to dynamically schedule computing resources based on real-time load and task requirements to maximize performance and reduce energy consumption.

[0142] Example background:

[0143] The task load of a certain intelligent computing platform is set as follows:

[0144] Data analysis tasks require high CPU resources and high task completion time requirements;

[0145] Image processing tasks have strong demands on GPUs and require a large amount of parallel computing;

[0146] Video rendering tasks have greater demands on CPU and memory usage, and need to ensure stable running time.

[0147] The system first generates multiple candidate strategies based on the current load (such as CPU, GPU occupancy, memory and network bandwidth) and the requirements of different tasks. Each strategy is generated by an independent sub-network and optimized based on the shared core knowledge. Three candidate strategies are generated:

[0148] Strategy 1 (θ 1 : Give priority to ensuring the resource requirements of data analysis tasks and reduce the computing resource allocation of image processing tasks and video rendering tasks;

[0149] Strategy 2 (θ 2 : Prioritize GPU resources for image processing tasks and allocate less to data analysis tasks;

[0150] Strategy 3 (θ_3): Consider all tasks comprehensively and dynamically allocate resources to balance the load.

[0151] These strategies need to go through an evolution process to ensure that they optimize resource usage under different historical load conditions.

[0152] The evolutionary algorithm optimizes the strategy through feedback from historical tasks. Specifically, the system calculates the evolution function of the strategy based on the execution of each historical task. It is assumed that data from five historical tasks are collected, namely Task A, Task B, Task C, Task D, and Task E. Each task has executed different strategies, and the system records their resource usage and performance.

[0153] For strategy 1 (θ 1 , assuming that it performs well in historical task A, and the resource consumption and completion time of task A are ideal. The corresponding policy parameter θ j is 0.8 (representing resource utilization) and the performance of task A is θ max =1.0 (i.e., the maximum performance of task A under this strategy). Therefore, historical task A has a greater impact on the evolution of strategy 1.

[0154] For task B, strategy 1 performs relatively poorly, the strategy parameter is 0.6, and the execution performance index of task B is θ max =0.9. Strategy 2 performs relatively better, with a parameter value of 0.9. Based on these historical feedbacks, the evolution function optimizes the strategy by weighting the impact of each historical task.

[0155] The evolution function is calculated as follows:

[0156]

[0157] Among them, δ j Represents the weight of each historical task, and sets for task A, δ 1 =0.4, for task B, δ 2 =0.3, tasks C, D, and E are δ 3 =0.1,δ 4 =0.1,δ 5 =0.1, so the strategic benefit after evolution can be calculated.

[0158] Set the policy parameters θ for each task j and θ max as follows:

[0159] Task A: θ 1 =0.8,θ max =1.0

[0160] Task B: θ 2 =0.6,θ max =0.9

[0161] Task C: θ 3 =0.7,θ max =1.1

[0162] Task D: θ 4 =0.75,θ max =1.0

[0163] Task E: θ 5 =0.85,θ max =1.0

[0164] The evolution function is calculated as follows:

[0165]

[0166] E(θ 1 )=0.4·0.2+0.3·0.3333+0.1·0.3636+0.1·0.25+0.1·0.15

[0167] E(θ 1 )=0.08+0.1+0.0364+0.025+0.015=0.2564

[0168] After calculation by the evolutionary algorithm, the evolutionary benefit of strategy 1 is 0.2564. Similarly, the evolutionary benefits of other candidate strategies are calculated. For example, the evolutionary benefit of strategy 2 is 0.4, and the evolutionary benefit of strategy 3 is 0.3. Therefore, strategy 2 will be selected as the optimal strategy, and the system will give priority to strategy 2 in resource scheduling. By optimizing the historical task execution data through the evolutionary algorithm, the resource scheduling strategy can be adjusted dynamically. During each task execution, the system will continuously optimize resource allocation, improve resource utilization and system stability based on the feedback of real-time load and historical data, and ensure that each task can obtain the optimal resource scheduling strategy under different load conditions.

[0169] Embodiment 2:

[0170] Combined with Figure 3 As shown in the flow chart, this embodiment constructs a multi-objective optimization model and combines different objective functions in a weighted sum manner to optimize the overall task scheduling performance. The core of this method is to adjust the weights between the various objective functions to find a balance point and resolve the conflict between the various objectives. This optimization method is expressed by the following formula:

[0171]

[0172] in:

[0173] Φ(W,X) represents the overall optimization objective function, which reflects the comprehensive performance of task scheduling and combines multiple performance indicators such as task completion time, resource utilization, and energy consumption; W i is the weight coefficient of the i-th objective, which determines the importance of this objective in the overall optimization; f i (X) is the i-th objective function, which indicates the actual performance of the i-th performance indicator in task scheduling; X is the system state vector, which contains information such as resource usage and load of the current task, reflecting the real-time status of the system; n is the number of objective functions. In this example, the objective functions considered include task completion time, resource utilization rate, energy consumption, etc.

[0174] In order to further understand the practical application of this method, an example is used to show how to use this optimization method to schedule resources. Suppose there is an intelligent computing platform that is used to handle various types of tasks, including data processing, image analysis, video rendering, etc. The resource requirements of each task are different, and there may be differences in priority between tasks. The resource scheduling problem of this platform is regarded as a multi-objective optimization problem, with the following specific goals:

[0175] 1. Task Completion Time: Minimize the execution time of the task to ensure that the task is completed as soon as possible; 2. Resource Utilization: Maximize the utilization of computing resources (such as CPU, GPU, memory) to avoid resource waste; 3. Energy Consumption: Minimize energy consumption to reduce the operating cost and environmental burden of the platform.

[0176] Step 1: Construct the system state vector

[0177] Before resource scheduling, it is necessary to collect system status information in order to construct the state vector X. Assume that at a certain moment, the system status is as follows:

[0178] The current CPU load is 70%;

[0179] GPU load is 50%;

[0180] Memory usage is 60%;

[0181] The current tasks include: a data analysis task (CPU occupies 70%, memory occupies 30%), an image processing task (GPU occupies 50%, memory occupies 20%), and a video rendering task (CPU occupies 60%, memory occupies 50%).

[0182] The total load is 90%, which means the system is running close to saturation.

[0183] This information can be summarized into a system state vector X = [0.7, 0.5, 0.6, 0.9], which represents the current resource usage and load level of the system.

[0184] Step 2: Determine the objective function and weights

[0185] Next, define each objective function f i (X), and assign a weight W to each target i The system is set to attach more importance to task completion time, and to attach moderate importance to resource utilization and energy consumption. The weights assigned to the goals are as follows:

[0186] W 1 =0.5 (the weight of task completion time);

[0187] W 2 =0.3 (weight of resource utilization);

[0188] W 3 =0.2 (weight of energy consumption).

[0189] The objective function is defined as follows:

[0190] f 1 (X): Task completion time, which represents the sum of the execution time of all tasks. The average time of task execution is proportional to the CPU, GPU, and memory load, that is, the task completion time is the weighted average of the resource load. The execution time of the data analysis task is set to 2 hours, the execution time of the image processing task is set to 1.5 hours, and the execution time of the video rendering task is set to 3 hours. The average task execution time is:

[0191]

[0192] f 2 (X): Resource usage, indicating the resource consumption of the task. Set to simply calculate the average usage of the current resources:

[0193]

[0194] f 3 (X): Energy consumption, which represents the power consumed during the execution of all tasks. The power consumption of each task is assumed to be proportional to its resource usage. The power consumption of the data analysis task, image processing task, and video rendering task is assumed to be 50W, 40W, and 60W respectively, and is linearly related to resource usage. The energy consumption is:

[0195]

[0196] Step 3: Calculate the overall optimization objective function

[0197] According to the above calculations, the actual value of each objective can be obtained and substituted into the weighted sum optimization objective function. The overall optimization objective function is:

[0198] Φ(W,X)=W 1 ·f 1 (X)+W 2 ·f 2 (X)+W 3 ·f 3 (X)

[0199] Substituting in the known values:

[0200] Φ(W,X)=0.5·1.85+0.3·0.6+0.2·48.33=0.925+0.18+9.666=10.771

[0201] Step 4: Adjust weights to resolve conflicting goals

[0202] At this time, the overall optimization target is 10.771. If the system considers the task completion time to be the most important, it can increase W 1 The weight of W is reduced 2 and W 3 By adjusting the weights, a balance can be found between different goals. For example, if the system increases the importance of task completion time to 0.7, and reduces the importance of resource utilization and energy consumption to 0.2 and 0.1 respectively, the optimization objective function can be recalculated:

[0203] Φ(W,X)=0.7·1.85+0.2·0.6+0.1·48.33=1.295+0.12+4.833=6.248

[0204] At this point, the system significantly reduced the overall optimization target value through weight adjustment, indicating that the optimization of task completion time was given priority and other goals were appropriately compromised.

[0205] This embodiment constructs the weighted sum of the objective function as a multi-objective optimization problem. Under the framework of the genetic algorithm, the optimal objective combination is found by continuously iterating and optimizing the weights. The core of this process is fitness evaluation, which determines which weight combinations are more conducive to the overall optimization of the system by calculating the performance of each objective weight combination under the current system state.

[0206] The fitness evaluation formula is as follows:

[0207]

[0208] in:

[0209] A(Wi ,X) represents the fitness value, which measures the target combination W i The performance under the current system state X, the higher the fitness value, the better the target combination; α j is the importance coefficient of target j, which reflects the contribution of different targets to system optimization. It usually ranges from [0,1] and represents the impact of the target on the overall performance of the system. j (X) is the value of the jth objective function, which represents the actual performance of the objective under a given system state X, such as task completion time, resource utilization, energy consumption, etc.; f j (X tre ) is the target value under ideal conditions, which means the goal that should be achieved without any resource constraints or delays. It is usually taken as the theoretical optimal value or the best system performance.

[0210] On this basis, the genetic algorithm will select and evolve the optimal target weight configuration according to the fitness value, so as to achieve long-term optimization of task scheduling under the conditions of changes in system load, task type, etc.

[0211] Suppose there is an intelligent computing platform for processing different types of tasks (such as data analysis, image processing, video rendering, etc.). The platform resources are limited, and the task execution process needs to balance the task completion time, resource utilization and energy consumption. Tasks in the platform usually have different priorities and resource requirements. The goal of the system is to dynamically optimize task scheduling, maximize resource utilization efficiency, minimize energy consumption, and ensure that tasks are completed within a reasonable time.

[0212] Step 1: Initial system state and objective function definition

[0213] At a certain moment, the platform's resource status is as follows:

[0214] The current CPU load is 65%;

[0215] GPU load is 40%;

[0216] Memory usage is 55%;

[0217] The task types include: data analysis tasks (CPU occupies 60%, memory occupies 25%), image processing tasks (GPU occupies 40%, memory occupies 30%), and video rendering tasks (CPU occupies 70%, memory occupies 45%).

[0218] For these tasks, three objective functions are defined:

[0219] 1. Task completion time (f 1(X)): The task completion time is proportional to the current resource usage. Set the average completion time of the current task to 2 hours, and adjust the task completion time based on the current resource usage.

[0220] 2. Resource utilization (f 2 (X)): Average utilization of system resources. Set the current resource utilization to 60%.

[0221] 3. Energy consumption (f 3 (X)): Set the power consumption of each task to be proportional to its resource usage. The energy consumption of the current platform is 45W.

[0222] Step 2: Define the desired state target value

[0223] Ideally, if system resources are not limited and there is no delay, the objective function should achieve the following ideal value:

[0224] Task completion time 1 (X tre ) = 1 hour (theoretical optimal time);

[0225] Resource utilization f 2 (X tre ) = 1 (all resources are fully utilized);

[0226] Energy consumption 3 (X tre )=40W (minimum energy consumption).

[0227] Step 3: Calculate the fitness value

[0228] Next, the fitness formula is used to calculate the fitness value under the current target weight configuration. For example, the initial weight value of the objective function is set to W 1 =0.4, W 2 =0.3, W 3 =0.3, and the importance coefficient of each target is α j They are α 1 =0.5, α 2 =0.3, α 3 =0.2.

[0229] According to actual calculations, under the current state, the task completion time, resource utilization and energy consumption are set to 2 hours, 0.6 and 45W respectively, and the target values ​​under the ideal state are 1 hour, 1 and 40W respectively. Then the fitness value is calculated as follows:

[0230] A(W 1 ,X)=0.5·|21|+0.3·|0.61|+0.2·|4540|

[0231] A(W 1 ,X)=0.5·1+0.3·0.4+0.2·5

[0232] A(W 1 ,X)=0.5+0.12+1=1.62

[0233] The genetic algorithm will evaluate the pros and cons of the current target weight configuration based on the fitness value, select the weight combination with higher fitness for crossover and mutation, and thus generate a new weight configuration. Through continuous iterative optimization, the genetic algorithm will eventually find an optimal target weight configuration, achieving the best balance between multiple goals such as task completion time, resource utilization, and energy consumption.

[0234] After multiple iterations, the final weight configuration of the system is set to W 1 =0.5, W 2 =0.3, W 3 =0.2, the fitness value is calculated as follows:

[0235] A(W 1 ,X)=0.5·|1.51|+0.3·|0.751|+0.2·|4340|

[0236] A(W 1 ,X)=0.5·0.5+0.3·0.25+0.2·3

[0237] A(W 1 ,X)=0.25+0.075+0.6=0.925

[0238] After optimization, the fitness value of the system is reduced, indicating that the resource scheduling performance has been improved, and the balance between task completion time, resource utilization and energy consumption has been optimized. Through the introduction of genetic algorithms, the system can dynamically adjust the target weight according to the execution data and feedback information of historical tasks, so that the resource scheduling plan is continuously optimized.

[0239] This embodiment develops an intelligent computing resource scheduling platform for processing multiple computing tasks running in parallel, which may involve different types of computing tasks such as data analysis, image rendering, and video processing. The resource requirements of each task are different. For example, data analysis tasks require higher CPU occupancy and lower GPU occupancy, while image rendering tasks rely more on the computing power of the GPU. In this context, the task scheduling goal of the platform is to balance the completion time, resource utilization, and energy consumption of the task to ensure that the task can be completed efficiently and with low energy consumption.

[0240] To achieve this goal, the core problem that the platform needs to deal with is how to make reasonable trade-offs between multiple conflicting optimization objectives (such as task completion time, resource consumption, and system stability). To deal with these problems, the platform adopts Pareto frontier optimization.

[0241] Step 1: Define the multi-objective optimization function

[0242] The optimization goals of the platform can be summarized into the following three items:

[0243] 1. Task completion time (f 1 (X)): The time required for the task from start to finish. The completion time of the data analysis task is set to 2 hours, the image rendering task is 3 hours, and the video processing task is 4 hours.

[0244] 2. Resource utilization rate (f 2 (X)): The utilization of various resources in the system (such as CPU, GPU, and memory). Set the resource utilization of the current platform to 60%.

[0245] 3. Energy consumption (f 3 (X)): The energy consumed by the platform when executing tasks. The power consumption of each task is assumed to be proportional to its resource requirements. The current platform energy consumption is 45W.

[0246] Through weighted summation, these objectives constitute a multi-objective optimization problem. It is hoped that an optimal weight combination can be found to minimize task completion time, maximize resource utilization, and minimize energy consumption.

[0247] Step 2: Introducing Pareto frontier optimization

[0248] In multi-objective optimization problems, the goal of Pareto frontier optimization is to find solutions that are better than other solutions in one objective and at least not inferior in other objectives. In this optimization process, the solution set P(W,X) contains multiple weight combinations, each weight combination W i Corresponding to a specific scheduling strategy.

[0249] The mathematical formula for Pareto frontier optimization is:

[0250]

[0251] in:

[0252] P(W,X) is the Pareto frontier solution set, which represents the optimal solution set among different objectives;

[0253] W i is the target weight combination corresponding to the i-th solution;

[0254] f j (X i ) is the performance of the ith solution under objective j, indicating the actual value of the jth objective in task scheduling;

[0255] f j (X j ) is the performance of other solutions under target j.

[0256] Step 3: Select the Pareto optimal solution

[0257] In the platform, multiple candidate task scheduling strategies are generated through the system's reinforcement learning and genetic algorithm, corresponding to different target weight combinations. For each strategy combination, it is necessary to evaluate it based on the system's feedback, determine its performance under different goals, and find the optimal solution set through Pareto frontier optimization.

[0258] For example, suppose there are several candidate scheduling strategies:

[0259] W 1 =(0.5, 0.3, 0.2): has a higher weight on task completion time, and a lower weight on resource utilization and energy consumption; W 2 =(0.4,0.4,0.2): has a relatively balanced weight on resource utilization and task completion time; W 3 =(0.3,0.5,0.2): tends to prioritize resource utilization and ignore energy consumption. For each candidate solution, the specific performance of each goal is calculated by performing task scheduling simulation. The setting results are as follows:

[0260] For W 1 ,The task completion time is 2.1 hours, the resource utilization rate is 0.58, and the energy consumption is 50W;

[0261] For W 2 ,The task completion time is 2.5 hours, the resource utilization rate is 0.62, and the energy consumption is 48W;

[0262] For W 3 ,The task completion time is 2.3 hours, the resource utilization rate is 0.65, and the energy consumption is 46W.

[0263] According to the above calculation results, different solutions are compared. According to the definition of Pareto optimization, if a solution is better than other solutions in one objective and is not inferior in other objectives, then the solution belongs to the Pareto frontier solution set.

[0264] W 1 =(0.5,0.3,0.2) performs best in task completion time (2.1 hours), but is slightly inferior in resource utilization and energy consumption;

[0265] W 2 =(0.4,0.4,0.2) is more balanced among multiple goals, but is inferior to W in terms of task completion time. 1 ;

[0266] W 3 =(0.3, 0.5, 0.2) performs best in terms of resource utilization, but is poor in terms of task completion time and energy consumption.

[0267] Therefore, the solution contained in the Pareto frontier solution set P(W,X) is W 1 and W 2 , because they provide a better balance between different objectives. 3 Although it is superior in one objective, it performs poorly in other objectives and is therefore not in the frontier solution set. Ultimately, the system will select the optimal solution in the Pareto frontier solution set based on the current task requirements and system status. For example, if the current system has a higher demand for task completion time, it may choose W 1 As a scheduling strategy. If the current system is more concerned with resource utilization, W may be selected 3 .

Claims

1. The intelligent computing resource configuration method based on reinforcement learning is characterized by The following steps are involved: S1. Intelligent resource state modeling and adaptive state space design: S1.

1. Fusion of monitoring data from different sources including hardware monitoring, software logs, and network bandwidth to form a multi-dimensional time series state space; and use of deep autoencoders to reduce the dimensionality and denoise the multimodal data; S1.2, introduce multi-task learning (MTL) in time series modeling to simultaneously predict the resource requirements and state evolution of multiple tasks; S2. Resource scheduling decision-making mechanism for multi-strategy collaborative optimization: S2.

1. Generate multiple predicted resource scheduling strategies using generative adversarial networks (GANs) and dynamically select strategy solutions when the load changes. S2.2, each strategy is implemented by an independent sub-network, sharing some core knowledge; through the strategy evolution mechanism, the strategy combination is optimized according to historical feedback including task completion time and resource consumption; S2.

3. In a heterogeneous resource environment including multiple cloud computing platforms and edge computing nodes, fine-grained scheduling of strategies based on differences in resource types; scheduling decisions using the distributed Q-learning mechanism in the reinforcement learning model; S3. Resource allocation and dynamic priority adjustment based on multi-task learning: S3.

1. Dynamically evaluate the priority of each task through deep reinforcement learning. The reward function includes the task completion time, as well as the delay sensitivity and resource consumption history of the task execution to calculate the priority of the predicted task under the current system load. S3.

2. Adjust the priorities between tasks through Bayesian optimization, introduce the load perception mechanism in reinforcement learning, intelligently identify the load characteristics of different types of tasks, and dynamically adjust resource allocation; S4. Evolution-based reward adjustment and realization of long-term optimization goals: S4.

1. Collect the execution data of various tasks during system operation, including execution time and resource utilization, to build a multi-objective optimization model; S4.2, use genetic algorithm GA to adjust the weight of each goal and dynamically optimize the goal combination according to real-time task requirements and historical performance; S4.

3. Make trade-offs between different objectives through Pareto frontier optimization.

2. The intelligent computing resource configuration method based on reinforcement learning according to claim 1 is characterized in that The resource scheduling decision-making mechanism method of multi-strategy collaborative optimization: Generator, used to generate multiple candidate resource scheduling strategies based on the current system status information including CPU load, memory usage, task type, and network bandwidth, where each candidate strategy is represented as a scheduling solution θ i , and the generator is trained through a reinforcement learning RL network to generate multiple optional strategy combinations; The discriminator is used to evaluate the effectiveness of each candidate strategy based on the performance evaluation function F(θ i ) measures the performance of each candidate strategy, where F(θ i ) is defined as a weighted function of multiple optimization objectives including task completion time, resource consumption, and system stability: F(θ i )=αT(θ i )+βC(θ i )+γS(θ i ) Among them, T(θ i ) is the task completion time, representing the strategy θ i The time required to complete the task; C(θ i ) is the resource consumption, indicating the strategy θ i The computing resources consumed when executing tasks; S(θ i ) is the system stability, indicating the strategy θ i The stability of the system operation, including failure rate and response delay; α, β, γ are the weighting coefficients required for scheduling task priorities, which determine the degree of each indicator in the overall evaluation.

3. The intelligent computing resource configuration method based on reinforcement learning according to claim 2 is characterized in that The resource scheduling decision-making mechanism method of multi-strategy collaborative optimization: By monitoring the system load changes in real time, the generator generates multiple scheduling strategy candidate solutions θ i , the discriminator selects the optimal strategy θ in real time based on the current load situation opt ; Adjust strategy according to load perception θ i , through the strategy selection function Select a resource scheduling strategy, where G(θ i ) is the performance evaluation function of the resource scheduling strategy, defined as: Among them, G(θ i ) represents the strategy θ i The total benefit is the performance indicator of the strategy selection under load changes; R(t) is the change of resource utilization over time t, indicating that the strategy θ i The proportion of resources used at time t; P(t) is the change of power consumption over time t, which represents the strategy θ i The power or energy consumed at time t; η1, η2 are the weighted coefficients between resource utilization and power consumption, reflecting the different requirements of the system for resource utilization and energy consumption. When η1>η2, the effective use of computing resources is prioritized; T is the evaluation time period, which is the time window for scheduling decisions.

4. The intelligent computing resource configuration method based on reinforcement learning according to claim 3 is characterized in that The resource scheduling decision-making mechanism method of multi-strategy collaborative optimization is as follows: each generated strategy is implemented by an independent sub-network, and core knowledge is shared between different sub-networks. Each candidate strategy is optimized through an evolutionary algorithm, and the strategy is updated through historical task execution data and feedback information.

5. The intelligent computing resource configuration method based on reinforcement learning according to claim 1 is characterized in that The method for achieving the evolution-based reward adjustment and long-term optimization goal is: First, a multi-objective optimization model is constructed; different objective functions are combined by weighted sum; the balance point is found by adjusting the weights between the objective functions so that the conflicts between the various objectives are resolved; and the optimization problem is expressed by the following formula: in: Φ(W,X) is the overall optimization objective function, which represents the comprehensive performance of task scheduling. The performance of the system is measured by the weighted sum of each objective function. i is the weight of the i-th objective, which determines the importance of the objective in the optimization; f i (X) is the i-th objective function, which represents the actual performance of the i-th performance indicator in task scheduling; X is the system state vector, which contains information related to resource usage and load status of the current task, reflecting the real-time status of the system; n is the number of objective functions, including task completion time, resource utilization rate, and energy consumption.

6. The intelligent computing resource configuration method based on reinforcement learning according to claim 5 is characterized in that The method for achieving the evolution-based reward adjustment and long-term optimization goal is as follows: a genetic algorithm GA is introduced to dynamically adjust the target weight; through continuous iterative optimization, the target weight configuration of the current task requirements and system status is found; fitness evaluation is a key part of the genetic algorithm to measure the pros and cons of each target weight combination.

7. The intelligent computing resource configuration method based on reinforcement learning according to claim 6 is characterized in that The method for achieving the evolution-based reward adjustment and long-term optimization goal is: Pareto frontier optimization is introduced to balance multiple objectives to find a solution that is optimal in one objective and not inferior in other objectives. In multi-objective optimization problems, the goal of Pareto frontier optimization is to calculate multiple solutions and find a solution set that cannot further improve the objective without sacrificing other objectives. The formula for this process is as follows: in: P(W,X) is the Pareto frontier solution set, which represents the optimal trade-off between different objectives. Each solution in the solution set represents a weight configuration that is superior to other solutions in some objectives. i is the target weight combination corresponding to the i-th solution; through Pareto frontier optimization, the best weight combination will be selected; f j (X i ) represents the performance of the i-th solution under target j; f j (X j ) is the performance of other solutions under target j; when f j (X i )≤f j (X j ), then W i At least not inferior in terms of target i.

Citation Information

Cited By

  • Multi-application resource allocation method and device

    CN120429125A

  • AI-based cloud data intelligent analysis and service management system and method

    CN120658632A

  • Distributed stream processing system parameter tuning method based on multi-objective reinforcement learning

    CN120973549A

  • Distributed stream processing system parameter tuning method based on multi-objective reinforcement learning

    CN120973549B

  • Intelligent data center optimization method based on distributed GPU computing power scheduling

    CN121166383A