Heterogeneous cluster task scheduling fusion method based on Q learning and genetic algorithm

By combining Q-Learning and genetic algorithms to optimize heterogeneous cluster task scheduling, the problems of resource utilization efficiency and dynamic adaptability in heterogeneous clusters are solved, and task completion time is shortened and resource utilization is improved, which is suitable for efficient task scheduling of large-scale heterogeneous clusters.

CN120256066AActive Publication Date: 2025-07-04ARTIFICIAL INTELLIGENCE INNOVATION RES INST OF ZHEJIANG UNIV OF TECH BINJIANG DISTRICT HANGZHOU

Patent Information

Application Number
CN202510737337.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-04
Publication Date
2025-07-04
Estimated Expiration
2045-06-04

AI Technical Summary

Technical Problem

Existing task scheduling strategies are difficult to take into account resource utilization efficiency, task execution time and system dynamic adaptability in heterogeneous cluster environments, especially in terms of tenant fairness and computing cost in large-scale GPU clusters.

Method used

Combining Q-Learning and genetic algorithm, the state space and action space of heterogeneous clusters are defined, and task allocation is optimized through Q-value tables and genetic algorithm populations, and the -greedy strategy is used to select actions and calculate reward values. New populations are generated by combining the cross- and mutated operations of the genetic algorithm, and the scheduling strategy is dynamically adjusted to adapt to cluster state changes.

Benefits of technology

Significantly shorten the task completion time, improve the adaptability of resource utilization and scheduling strategies, avoid local optimal solutions, and improve task processing efficiency and flexibility of large-scale heterogeneous clusters.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120256066A_ABST
    Figure CN120256066A_ABST
Patent Text Reader

Abstract

The invention discloses a heterogeneous cluster task scheduling fusion method based on Q learning and a genetic algorithm, and belongs to the technical field of intelligent scheduling. Collaborative optimization is carried out through a dynamic feedback mechanism of Q-Learning and the global search capability of the genetic algorithm: in an initialization stage, a cluster state space and an action space are defined, and a multi-Q-value table and an initial population are generated; during task scheduling, nodes are selected based on a # imgabs0 #-greedy strategy, reward values are calculated, and a Q value table is updated; meanwhile, the scheduling scheme is coded into chromosomes, a new population is generated through roulette selection, crossover and variation, and a Q value table is optimized; in the iteration process, the convergence is judged according to the Q value change or the population fitness change, and the strategy is dynamically adjusted. According to the method, the advantages of double algorithms are fused, local optimum is avoided, the task processing efficiency and the resource utilization rate are remarkably improved, adaptive strategy updating during cluster state change is supported, and the method is suitable for an efficient task scheduling scene of a large-scale heterogeneous cluster.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer cluster task scheduling, and particularly relates to a heterogeneous cluster task scheduling fusion method based on Q-learning and genetic algorithm. Background Art

[0002] In a heterogeneous cluster environment, task scheduling faces many challenges. Traditional scheduling strategies often struggle to balance resource utilization efficiency, task execution time, and system dynamic adaptability simultaneously.

[0003] On the one hand, some existing rule-based scheduling algorithms, such as First-Come-First-Served (FCFS), Shortest Job First (SJF), etc., often cannot adapt to the complexity of heterogeneous clusters because they lack consideration of task and resource heterogeneity. To address this issue, researchers have proposed various intelligent algorithm-based scheduling strategies, such as genetic algorithm, Particle Swarm Optimization (PSO), and ant colony algorithm, etc. For example, in the design solution with patent application number CN202311605127.1, the nodes within the cluster are divided into regions and the main nodes are labeled. The ant colony algorithm is used to determine the optimal path for traversing the main nodes, and then the optimal region and optimal node are determined. Based on the SDN controller, the tasks to be scheduled are assigned to the optimal nodes. Such solutions optimize task scheduling through swarm intelligence behavior, but they usually require a large amount of computing resources and may be difficult to quickly adapt in a dynamically changing environment.

[0004] On the other hand, the concept of tenants is rarely considered in GPU cluster scheduling. However, cloud providers often involve multi-tenant resource allocation when providing GPUs. It has become a difficult problem to improve cluster efficiency while ensuring tenant fairness. For example, in the design solution with patent application number CN202310226362.1, tenants submit tasks to the waiting queue, obtain task information of cluster computing nodes, select tasks that meet the conditions, and schedule them to appropriate computing nodes according to affinity. Before tenants submit tasks, the GPUs are virtualized into fine-grained resources and global and local schedulers are started. In such solutions, when there are many tenants and complex tasks, the computational cost of determining task priorities and computing node affinity may be high, and the fine-grained partitioning and management of GPU resources increase system complexity, which to a certain extent limits the application effect in large-scale GPU cluster scheduling.

[0005] With the wide application of heterogeneous clusters in various fields, the demand for efficient and adaptive task scheduling strategies is becoming increasingly urgent. Existing task scheduling strategies have many deficiencies when facing the complexity and dynamics of heterogeneous clusters. Therefore, a heterogeneous cluster adaptive task scheduling fusion strategy based on Q-Learning and genetic algorithm is needed to solve these problems. Summary of the Invention

[0006] To solve the above technical problems, the present invention proposes a heterogeneous cluster task scheduling fusion method and system based on Q-learning and genetic algorithm to solve the problems existing in the above prior art.

[0007] In a first aspect, to achieve the above object, the present invention provides a heterogeneous cluster task scheduling fusion method based on Q-learning and genetic algorithm, including the following steps:

[0008] S1. Initialize the Q-Learning algorithm and genetic algorithm, define the state space and action space of the heterogeneous cluster, where the state space includes the resource attribute information of the nodes, the action space is the set of actions for task allocation to different nodes, and generate an initial Q-value table and a genetic algorithm population;

[0009] S2. Receive a scheduling task, select an action from the Q-value table according to the current cluster state, perform task allocation and calculate the reward value based on task execution time and resource utilization rate, and update the Q-value table;

[0010] S3. Encode the task scheduling scheme into a chromosome and add it to the genetic algorithm population, perform selection, crossover and mutation operations on the population based on the fitness function, generate a new population and update the Q-value table;

[0011] S4. Repeat the iterative process of S2 to S3, judge the algorithm convergence according to the change of the Q-value table or the change of the population fitness, and determine the final task scheduling strategy.

[0012] Optionally, in S1, the state space includes the CPU load, memory load and network bandwidth of the nodes, the initial Q-value table is initialized by uniformly distributed random numbers, and multiple Q-value tables are generated according to the maximum number of schedulable tasks.

[0013] Optionally, in S2, the process of selecting an action from the Q-value table according to the current cluster state, performing task allocation and calculating the reward value based on task execution time and resource utilization rate includes: the selection of actions adopts - greedy strategy, with probability randomly select an action, with probability select the action with the largest current Q-value; the reward value is obtained by weighted summation of the reciprocal of the task execution time and the reciprocal of the resource utilization rate.

[0014] Optionally, in S2, the update of the Q-value table includes the following operations: according to the Q-value of the current state and action, the obtained reward value, the expected maximum Q-value of the next state, combined with the learning rate and discount factor, update the Q-value corresponding to the current state and action; and encode the executed task scheduling scheme into a binary chromosome, where the gene position in the chromosome represents whether the task is allocated to the corresponding node, and add it to the genetic algorithm population for subsequent evolution operations.

[0015] Optionally, in S3, the fitness function is determined based on task execution time and resource utilization. Population evolution includes roulette wheel selection of parental individuals, crossover operation to generate offspring individuals, and mutation operation to introduce new genes. The updated Q-value table is used to optimize subsequent task scheduling strategies.

[0016] Optionally, in S4, during the judgment process of convergence, the conditions for judging convergence include: the change in Q-value is less than a preset change threshold or the change in the optimal fitness of the population is less than a preset fitness threshold within a continuous number of scheduling cycles.

[0017] In a second aspect, the present invention also provides a heterogeneous cluster task scheduling fusion system based on Q-learning and genetic algorithm for implementing a heterogeneous cluster task scheduling fusion method based on Q-learning and genetic algorithm. The system includes:

[0018] An initialization module for initializing the Q-learning algorithm and genetic algorithm, defining the state space and action space of the heterogeneous cluster, and generating an initial Q-value table and a genetic algorithm population;

[0019] A scheduling module for receiving scheduling tasks, selecting actions from the Q-value table according to the current cluster state, performing task allocation, calculating the reward value based on task execution time and resource utilization, and updating the Q-value table simultaneously;

[0020] A genetic algorithm module for encoding the task scheduling scheme into chromosomes and adding them to the population, and performing selection, crossover, and mutation operations on the population based on the fitness function to generate a new population, and updating the Q-value table simultaneously;

[0021] A convergence judgment module for detecting the change in the Q-value table or the optimal fitness of the population within continuous scheduling cycles, and judging the algorithm convergence accordingly to determine the final task scheduling strategy.

[0022] Optionally, in the initialization module, the defined state space includes the CPU load, memory load, and network bandwidth of each node, and the initial Q-value table is initialized with uniformly distributed random numbers, and multiple Q-value tables are generated according to the maximum number of schedulable tasks.

[0023] In a third aspect, the present invention also provides a computer terminal device, including:

[0024] One or more processors;

[0025] A memory coupled to the processor for storing one or more programs;

[0026] When the one or more programs are executed by the one or more processors, the one or more processors implement a heterogeneous cluster task scheduling fusion method based on Q-learning and genetic algorithm.

[0027] In a fourth aspect, the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements a heterogeneous cluster task scheduling fusion method based on Q-learning and genetic algorithms.

[0028] Compared with the prior art, the present invention has the following advantages and technical effects:

[0029] The heterogeneous cluster task scheduling fusion method and system based on Q-learning and genetic algorithms provided by the present invention deeply integrate the reinforcement learning feedback mechanism of Q-Learning and the global search ability of genetic algorithms, overcoming the problem of insufficient adaptability of traditional heterogeneous cluster scheduling strategies in dynamic environments. Q-Learning dynamically selects task allocation actions based on real-time cluster states (such as CPU load, memory load, network bandwidth), and optimizes the scheduling strategy through a reward function (integrating task execution time and resource utilization rate), significantly shortening the task completion time; genetic algorithms generate diverse scheduling schemes through crossover and mutation operations, expanding the search space and preventing Q-Learning from falling into local optimal solutions. The collaborative iterative optimization of the two (Q-value table update and population evolution) improves the global optimality of the scheduling strategy while reducing the computational cost of a single algorithm. In addition, the algorithm dynamically judges convergence based on changes in Q-values or population fitness, and restarts the learning process when the cluster state changes significantly, ensuring the adaptive ability of the scheduling strategy. The present invention effectively solves problems in the prior art such as high computational cost of intelligent algorithms, complex multi-tenant resource allocation, and slow response to dynamic environments, and is applicable to the efficient task scheduling requirements in large-scale heterogeneous cluster scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] The accompanying drawings that form a part of this application are used to provide a further understanding of this application. The schematic embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation to this application. In the drawings:

[0031] Figure 1 is a flowchart of the method of an alternative embodiment in Embodiment 1 of the present invention;

[0032] Figure 2 is a schematic diagram of the overall structure of the scheduler based on Q-Learning in an embodiment of the present invention;

[0033] Figure 3 is a flowchart of the method in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0034] It should be noted that, without conflict, the embodiments in this application and the features in the embodiments can be combined with each other. The following will refer to the accompanying drawings and combine with the embodiments to detail this application.

[0035] It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0036] The present invention proposes a fusion model that combines the global search ability of the genetic algorithm and the adaptive learning ability of Q-Learning to prevent the algorithm from falling into local optima, enabling the system to learn the overall optimal decision in the interaction with the environment to solve the complex heterogeneous cluster task scheduling problem.

[0037] The method proposed in the present invention can dynamically adjust the scheduling strategy according to the real-time state of the cluster and the task requirements to adapt to environmental changes.

[0038] Embodiment 1

[0039] As Figure 3 shown, this embodiment provides a heterogeneous cluster task scheduling fusion method based on Q-learning and genetic algorithm, including:

[0040] S1. Initialize the Q-Learning algorithm and the genetic algorithm, define the state space and action space of the heterogeneous cluster, where the state space includes the resource attribute information of the nodes, the action space is the set of actions for task allocation to different nodes, and generate an initial Q-value table and a genetic algorithm population;

[0041] S2. Receive the scheduling task, select an action from the Q-value table according to the current cluster state, execute the task allocation and calculate the reward value based on the task execution time and resource utilization rate, and update the Q-value table;

[0042] S3. Encode the task scheduling scheme into a chromosome and add it to the genetic algorithm population, perform selection, crossover, and mutation operations on the population based on the fitness function, generate a new population and update the Q-value table;

[0043] S4. Repeat the iterative process of S2 to S3, judge the convergence of the algorithm according to the change of the Q-value table or the change of the population fitness, and determine the final task scheduling strategy.

[0044] As Figure 1 shown, as an alternative implementation:

[0045] 1. Initialize the Q-leraing algorithm, establish a Q-value table for recording the expected return, and initialize the genetic evolution algorithm.

[0046] 2. Receive the scheduling task, select the corresponding Q-value table, obtain the current state of the cluster, use the Q-Learning algorithm to select the optimal action, and calculate the obtained reward.

[0047] 3. Analyze the statistical data of the reward function based on the reward obtained in step 2, and update the Q-value table of the Q-learning algorithm.

[0048] 4. Use the genetic evolution algorithm to perform genetic evolution operations on the scheduling schemes in the Q-value table, generate better scheduling schemes, and update them to the Q-value table used in step 2 to prevent the Q-learning algorithm from falling into local optimality.

[0049] 5. Repeat the iterative optimization process until the scheduling strategy is stable and meets the predetermined performance metrics.

[0050] As an implementation method in this embodiment, in S1, the state space includes the CPU load, memory load, and network bandwidth of the nodes. The initial Q-value table is initialized with uniformly distributed random numbers, and multiple Q-value tables are generated according to the maximum number of schedulable tasks.

[0051] Specifically, step S1 includes:

[0052] 1. Policy initialization:

[0053] Define the state space and policy space, initialize the parameters of the Q function and the genetic algorithm, and generate the initial population. The specific steps are as follows:

[0054] Step 1-1: Define the state space and action space:

[0055] Represent the state of the heterogeneous cluster as a vector containing various attribute information of the nodes, such as the CPU load of the nodes , memory load , network bandwidth , etc. That is, the state space :

[0056] ;

[0057] The action is to assign tasks to different nodes. Assume that there are nodes in the cluster, then the action space :

[0058] ;

[0059] where represents assigning the task to node , .

[0060] Step 1-2: Initialize the Q-value table:

[0061] Let be the Q-value for taking action in state . Initially, set to a random number uniformly distributed within the interval to initialize the Q-value table. According to the maximum number of tasks r that can be scheduled simultaneously, r Q-value tables are initialized to improve the pertinence of the scheduling algorithm for scheduling tasks.

[0062] Step 1-3: Initialize the genetic algorithm parameters:

[0063] Determine the number of individuals (i.e., task scheduling schemes) in the initial population of the genetic algorithm, which is the population size ; set the crossover probability and the mutation probability ; encode the task scheduling scheme into a chromosome using binary encoding, where each gene bit represents whether a task is assigned to the corresponding node.

[0064] As an implementation method in this embodiment, in S2, the process of selecting an action from the Q-value table according to the current cluster state, performing task allocation, and calculating the reward value based on task execution time and resource utilization includes: The selection of the action adopts the -greedy strategy, randomly selects an action with probability , and selects the action with the largest current Q-value with probability ; the reward value is obtained by the weighted sum of the reciprocal of the task execution time and the reciprocal of the resource utilization.

[0065] Specifically, step S2 includes:

[0066] 2. Iterative optimization of the scheduling strategy:

[0067] The overall structure of the scheduler based on Q-Learning is as Figure 2 shown. The specific steps in each task scheduling cycle are as follows:

[0068] Step 2-1: Observe the current state :

[0069] Receive a scheduling task, and select the corresponding Q-value table according to the number of tasks that need to be scheduled simultaneously. Obtain the current state information of the heterogeneous cluster and determine the state in the state space .

[0070] Step 2-2: Select an action according to Q-Learning :

[0071] Calculate each action The Q value in the current state , select an action using -greedy strategy, with probability randomly select an action, with probability select the action with the maximum Q value. That is: if , then , where is a randomly selected action from the action space , otherwise,[[]] , where represents the independent variable value corresponding to the maximum Q value .

[0072] Step 1-3: Execute task scheduling and obtain rewards :

[0073] Execute task scheduling according to the selected action, that is, allocate tasks to corresponding nodes. Define the reward function according to the task execution time and resource utilization rate , that is:

[0074] ;

[0075] where and are weight coefficients used to balance the importance of task execution time and resource utilization rate in the reward. Observe the result of task execution and calculate the obtained reward .

[0076] As an implementation method in this embodiment, in S2, the operation of updating the Q value table includes: updating the Q value corresponding to the current state and action according to the Q value of the current state and action, the obtained reward value, and the expected maximum Q value of the next state, in combination with the learning rate and the discount factor; and encoding the executed task scheduling scheme into a binary chromosome, where the gene bits in the chromosome represent whether the task is allocated to the corresponding node, and adding it to the genetic algorithm population for subsequent evolution operations.

[0077] Specifically, the process of updating the Q value table in step S2 includes:

[0078] 3. Update the Q value table:

[0079] Step 3-1: Update the Q function and add the task scheduling scheme to the genetic algorithm population;

[0080] Update the Q value according to the update formula of Q-Learning, that is:

[0081] ;

[0082] Among them, is the learning rate, is the discount factor, is the executed action and the next state transferred to after represents the next state reached after executing the action All possible actions in the action space under the next state is to take the maximum value among the Q-values corresponding to all possible actions in the state for calculating and updating the Q-value of the current state-action pair

[0083] Encode the task scheduling scheme after execution into the form of a chromosome and add it as an individual to the population of the genetic algorithm.

[0084] As an implementation method in this embodiment, in S3, the fitness function is determined based on the task execution time and resource utilization rate. The population evolution includes roulette wheel selection of parental individuals, crossover operation to generate offspring individuals, and mutation operation to introduce new genes. The updated Q-value table is used to optimize the subsequent task scheduling strategy.

[0085] Specifically, step S3 includes:

[0086] 4. Optimize the scheduling strategy using the genetic algorithm:

[0087] Step 4-1: Calculate the fitness of each individual in the population:

[0088] Based on the task execution time and the resource utilization rate Define the fitness function, that is:

[0089] ;

[0090] Among them and are the weight coefficients. Calculate the fitness of each individual (i.e., the task scheduling scheme) in the population according to the fitness function.

[0091] Step 4-2: Select, crossover, and mutate to generate a new population:

[0092] According to the individual fitness, select some individuals from the population as parents using the roulette wheel selection strategy. The higher the fitness of an individual, the greater the probability of being selected. The probability of an individual being selected is proportional to its fitness. That is:

[0093] ;

[0094] Among them is the probability of an individual being selected,​​ is the fitness of an individual, is the sum of the fitnesses of all individuals in the population.

[0095] For the selected parent individuals, according to the crossover probability perform crossover operations. Through the crossover operations, the excellent genes of the parent individuals can be combined to generate offspring individuals with new gene combinations, increasing the diversity of the population.

[0096] For the individuals after crossover, according to the mutation probability perform mutation operations. The mutation operations can introduce new genes, prevent the population from converging to the local optimal solution prematurely, and maintain the diversity and search ability of the population.

[0097] After selection, crossover, and mutation operations, a new population of task scheduling schemes is generated, the corresponding Q-value table is updated, and the scheduling strategy is optimized.

[0098] As an implementation manner in this embodiment, in S4, during the determination process of convergence, the conditions for determining convergence include: the change in the Q-value is less than a preset change threshold or the change in the optimal fitness of the population is less than a preset fitness threshold within a continuous number of scheduling cycles.

[0099] Specifically, step S4 includes:

[0100] 5. Continuous optimization and convergence:

[0101] Step 5-1: Repeated iteration:

[0102] Continuously repeat the operations of the above scheduling cycle. As the number of iterations increases, the Q-Learning algorithm continuously learns a better task scheduling strategy, and the genetic algorithm also continuously evolves the population, providing more potential excellent task scheduling schemes.

[0103] Step 5-2: Convergence judgment:

[0104] If the change in the Q-value is less than a threshold within a continuous number of scheduling cycles, or the change in the fitness of the optimal individual in the population is less than a threshold , it is considered that the algorithm has converged.

[0105] Step 5-3: Final decision determination:

[0106] When the algorithm converges, the final task scheduling strategy can be determined according to the converged Q-value or the optimal task allocation scheme in the population. In practical applications, the cluster status can continue to be monitored. When the status changes significantly, restart the learning and evolution process to maintain the adaptive ability of the task scheduling strategy.

[0107] In summary, in the existing technical solutions, there is no technical solution that combines reinforcement learning with other technologies to improve heterogeneous cluster task scheduling. The present invention uses a genetic algorithm to provide diverse initial solutions, expand the search space, and prevent reinforcement learning from falling into local optimal solutions; the feedback mechanism of reinforcement learning can help the genetic algorithm better adjust the search direction in subsequent iterations.

[0108] The present invention can improve scheduling flexibility and task processing efficiency. By intelligently selecting the optimal nodes and allocating tasks, it is expected to shorten the task completion time and improve the task processing efficiency of the entire cluster.

[0109] Moreover, the present invention combines the global search ability of the genetic algorithm and the adaptive learning ability of Q-Learning, and dynamically adjusts the key parameters in the genetic algorithm through the Q-Learning algorithm.

[0110] The present invention designs a reward mechanism to calculate rewards based on the results of task scheduling and use these rewards to update the Q-value table of Q-Learning. This reward mechanism encourages the scheduling strategy to develop towards the optimization goal.

[0111] Based on this, a heterogeneous cluster task scheduling fusion method based on Q-learning and genetic algorithm provided by an embodiment of the present invention effectively overcomes the limitations of traditional scheduling strategies in heterogeneous cluster environments by deeply integrating the global search ability of the genetic algorithm and the adaptive learning ability of Q-Learning. The genetic algorithm provides diverse initial scheduling solutions and expands the search space to prevent reinforcement learning from falling into local optimal solutions, while the real-time feedback mechanism of Q-Learning dynamically adjusts the search direction of the genetic algorithm to enhance the flexibility of the strategy. By designing a reward mechanism based on task execution time and resource utilization, it guides the continuous optimization of the scheduling strategy, significantly shortens the task completion time, and improves the overall efficiency of the cluster. In addition, after the algorithm converges, the learning and evolution process can be restarted according to the real-time cluster state to ensure the adaptability of the scheduling strategy in a dynamic environment, thereby achieving the coordinated improvement of resource utilization and task processing efficiency in complex heterogeneous cluster scenarios.

[0112] Embodiment 2

[0113] In this embodiment, a computer terminal device is provided, including:

[0114] One or more processors;

[0115] A memory, coupled to the processor, for storing one or more programs;

[0116] When the one or more programs are executed by the one or more processors, the one or more processors implement the method in the above embodiment.

[0117] In this embodiment, a computer-readable storage medium is further provided, on which a computer program is stored. When the computer program is executed by a processor, the methods in the above embodiments are implemented.

[0118] In this embodiment, an electronic device is further provided, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the methods in the above embodiments.

[0119] The above program can run in the processor, or can also be stored in the memory (or referred to as a computer-readable medium). The computer-readable medium includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computing device.

[0120] These computer programs can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate computer-implemented processing. Thus, the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in a process Figure 1 a process or multiple processes and / or blocks Figure 1 The steps corresponding to different steps can be implemented by different modules.

[0121] In this embodiment, such a device or system is provided. The system is called a heterogeneous cluster task scheduling fusion system based on Q-learning and genetic algorithms, including:

[0122] An initialization module, configured to initialize the Q-learning algorithm and the genetic algorithm, define the state space and action space of the heterogeneous cluster, and generate an initial Q-value table and a genetic algorithm population;

[0123] A scheduling module, configured to receive a scheduling task, select an action from the Q-value table according to the current cluster state, execute task allocation, calculate a reward value according to the task execution time and resource utilization rate, and update the Q-value table at the same time;

[0124] A genetic algorithm module, which is used to encode the task scheduling scheme into chromosomes and add them to the population, and perform selection, crossover, and mutation operations on the population based on the fitness function to generate a new population, and update the Q-value table at the same time;

[0125] A convergence judgment module, which is used to detect the changes in the Q-value table or the optimal fitness of the population within consecutive scheduling cycles, and judge the algorithm convergence accordingly to determine the final task scheduling strategy.

[0126] As an implementation manner in this embodiment, in the initialization module, the defined state space includes the CPU load, memory load, and network bandwidth of each node, and the initial Q-value table is initialized with uniformly distributed random numbers, and multiple Q-value tables are generated according to the maximum number of schedulable tasks.

[0127] As an implementation manner in this embodiment, in the scheduling module, actions are selected from the Q-value table according to the current cluster state to perform task allocation, and the reward value based on the task execution time and resource utilization rate is calculated. Among them, the action selection adopts the ε-greedy strategy, that is, an action is randomly selected with probability ε, and the action with the largest current Q-value is selected with probability 1-ε, and the reward value is obtained by weighted summation of the reciprocal of the task execution time and the reciprocal of the resource utilization rate.

[0128] As an implementation manner in this embodiment, the process of updating the Q-value table in the scheduling module includes: based on the Q-value of the current state and the selected action, the obtained reward value, and the expected maximum Q-value of the next state, and combined with the learning rate and the discount factor, update the corresponding Q-value; at the same time, encode the executed task scheduling scheme into a binary chromosome, where each gene bit in the chromosome represents whether the task is allocated to the corresponding node, and add this chromosome to the population of the genetic algorithm module for subsequent evolution operations.

[0129] As an implementation manner in this embodiment, in the genetic algorithm module, the fitness function is determined based on the task execution time and resource utilization rate. The population evolution operations include: selecting parent individuals by roulette wheel selection, generating offspring individuals through crossover operations, and introducing new genes through mutation operations, and the updated Q-value table is used to further optimize the task scheduling strategy.

[0130] As an implementation manner in this embodiment, in the convergence judgment module, the judgment conditions for system convergence include: within consecutive several scheduling cycles, the change of each Q-value in the Q-value table is less than the preset change threshold, or the change of the optimal fitness in the population is less than the preset fitness threshold.

[0131] This system or device is used to implement the functions of the method in the above embodiment. Each module in this system or device corresponds to each step in the method, and those that have been described in the method will not be repeated here.

[0132] Through the above embodiments, the problem of the fusion of heterogeneous cluster task scheduling based on Q-learning and genetic algorithm in the related art is solved. In the present invention, the global search ability of the genetic algorithm is deeply fused with the adaptive learning ability of Q-Learning, effectively overcoming the limitations of traditional scheduling strategies in heterogeneous cluster environments. The genetic algorithm provides diverse initial scheduling schemes and expands the search space to avoid the reinforcement learning from falling into local optimal solutions, while the real-time feedback mechanism of Q-Learning dynamically adjusts the search direction of the genetic algorithm and enhances the flexibility of the strategy. By designing a reward mechanism based on task execution time and resource utilization rate, the scheduling strategy is guided to continuously optimize, significantly shortening the task completion time and improving the overall efficiency of the cluster. In addition, after the algorithm converges, the learning and evolution process can be restarted according to the real-time cluster state to ensure the adaptability of the scheduling strategy in a dynamic environment, thereby achieving a coordinated improvement of resource utilization rate and task processing efficiency in complex heterogeneous cluster scenarios.

[0133] The above are only the preferred specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed in the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A heterogeneous cluster task scheduling fusion method based on Q-learning and genetic algorithm, characterized in that It includes the following steps: S1. Initialize the Q-Learning algorithm and the genetic algorithm, define the state space and action space of the heterogeneous cluster, where the state space includes the resource attribute information of the nodes, the action space is the set of actions for task allocation to different nodes, and generate an initial Q-value table and a genetic algorithm population; S2. Receive a scheduling task, select an action from the Q-value table according to the current cluster state, perform task allocation and calculate the reward value based on the task execution time and resource utilization rate, and update the Q-value table; S3. Encode the task scheduling scheme into a chromosome and add it to the genetic algorithm population, perform selection, crossover, and mutation operations on the population based on the fitness function, generate a new population and update the Q-value table; S4. Repeat the iterative process of S2 to S3, judge the convergence of the algorithm according to the change of the Q-value table or the change of the population fitness, and determine the final task scheduling strategy.

2. The method according to claim 1, wherein In S1, the state space includes the CPU load, memory load, and network bandwidth of the nodes, the initial Q-value table is initialized with uniformly distributed random numbers, and multiple Q-value tables are generated according to the maximum number of schedulable tasks.

3. The method according to claim 1, characterized in that, In S2, the process of selecting an action from the Q-value table according to the current cluster state, performing task allocation, and calculating the reward value based on task execution time and resource utilization includes: The selection of actions adopts - a greedy strategy, with a probability of randomly select an action, with a probability of select the action with the largest current Q value; the reward value is obtained by weighted summation of the reciprocal of the task execution time and the reciprocal of the resource utilization rate.

4. The method according to claim 1, characterized in that In S2, the update of the Q-value table includes the following operations: update the Q-value corresponding to the current state and action according to the Q-value of the current state and action, the obtained reward value, the expected maximum Q-value of the next state, in combination with the learning rate and the discount factor; And encode the executed task scheduling scheme into a binary chromosome, where the gene bits in the chromosome represent whether the task is allocated to the corresponding node, and add it to the genetic algorithm population for subsequent evolution operations.

5. The method according to claim 1, characterized in that In S3, the fitness function is determined based on the task execution time and resource utilization rate, the population evolution includes roulette wheel selection of parent individuals, crossover operation to generate offspring individuals, and mutation operation to introduce new genes, and the updated Q-value table is used to optimize the subsequent task scheduling strategy.

6. The method according to claim 1, wherein In S4, during the judgment process of the convergence, the conditions for judging convergence include: the change of the Q-value is less than the preset change threshold or the change of the population optimal fitness is less than the preset fitness threshold within a continuous number of scheduling cycles.

7. A heterogeneous cluster task scheduling fusion system based on Q - learning and genetic algorithm, characterized in that, The system includes: An initialization module, used to initialize the Q-learning algorithm and the genetic algorithm, define the state space and action space of the heterogeneous cluster, and generate an initial Q-value table and a genetic algorithm population; A scheduling module, used to receive a scheduling task, select an action from the Q-value table according to the current cluster state, perform task allocation, and calculate the reward value according to the task execution time and resource utilization rate, and at the same time update the Q-value table; A genetic algorithm module, used to encode the task scheduling scheme into a chromosome and add it to the population, and perform selection, crossover, and mutation operations on the population based on the fitness function to generate a new population, and at the same time update the Q-value table; A convergence judgment module, used to detect the change of the Q-value table or the population optimal fitness within continuous scheduling cycles, and judge the convergence of the algorithm accordingly to determine the final task scheduling strategy.

8. The system according to claim 7, characterized in that In the initialization module, the defined state space includes the CPU load, memory load, and network bandwidth of each node, and the initial Q-value table is initialized with uniformly distributed random numbers, and multiple Q-value tables are generated according to the maximum number of schedulable tasks.

9. A computer terminal device, characterized in that, It includes: One or more processors; A memory, coupled to the processor, for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the heterogeneous cluster task scheduling fusion method based on Q-learning and genetic algorithm as described in any one of claims 1-6.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the heterogeneous cluster task scheduling fusion method based on Q-learning and genetic algorithm as described in any one of claims 1-6.

Citation Information

Patent Citations

  • GPU cluster scheduling method and device

    CN116431329A

  • Cluster task scheduling method and device and storage medium

    CN117596245A

  • Heterogeneous platform task scheduling method and system based on Q learning

    CN112256422A

  • Associated task scheduling method based on evolutionary algorithm

    CN112346839A

  • Big data analysis-based plan scheduling optimization method

    CN117076077A

Cited By

  • Self-adaptive weight task scheduling method and system based on time sequence differential learning

    CN120973493A

  • Inland river port unmanned shore tackle automatic scheduling method and system based on artificial intelligence

    CN121169038A