Task scheduling optimization method and device based on reinforcement learning, equipment and medium

By using reinforcement learning algorithms to process task priorities and resource requirements, constructing state space and action space, generating the optimal task allocation plan, and combining historical data to update the model, it solves the adaptability problem of traditional scheduling methods in complex dynamic environments, and achieves accurate mapping of resource status and task requirements and efficient scheduling.

CN120780432APending Publication Date: 2025-10-14SHAOGUAN XINGCHENG NETWORK TECH CO LTD

Patent Information

Application Number
CN202510892776.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-30
Publication Date
2025-10-14

AI Technical Summary

Technical Problem

Traditional task scheduling methods have difficulty achieving accurate and efficient task scheduling in complex dynamic environments, especially in cloud computing and industrial automation production lines. Existing technologies such as deep neural networks lack long-term feedback and adaptive capabilities and are prone to falling into local optimal solutions.

Method used

By obtaining real-time system resource status data, using reinforcement learning algorithms to process task priorities and resource demand vectors, constructing state space and action space, generating the optimal task allocation plan through policy iteration, and updating the reinforcement learning model in combination with historical scheduling data, a closed-loop optimization strategy is formed.

Benefits of technology

It achieves accurate mapping of resource status and task requirements, overcomes the lack of adaptability of traditional static scheduling to complex scenarios, and improves the accuracy of task scheduling and resource utilization efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120780432A_ABST
    Figure CN120780432A_ABST
Patent Text Reader

Abstract

The invention relates to a task scheduling optimization method and device based on reinforcement learning, equipment and a medium. The method comprises the steps that firstly, system resource state data are collected in real time, dynamic environment characteristics are determined through preprocessing and time sequence analysis, task characteristic data are analyzed at the same time, and a task priority sequence and a resource demand vector are generated through a priority ranking algorithm and a resource evaluation model; and then a state space and an action space are constructed by adopting a reinforcement learning algorithm, an optimal task allocation scheme is generated through strategy iteration and reward function optimization, and if the scheme meets a resource balance threshold, scheduling is executed, and performance indexes are collected. And finally, fusing real-time indexes with historical data, and updating parameters of the reinforcement learning model through experience playback and gradient descent to form a closed-loop optimized improved scheduling strategy. By adopting the method, the accurate mapping of the resource state and the task requirement can be realized, and the problem of insufficient adaptability of the traditional static scheduling to a complex scene is solved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the field of computer science and technology, and particularly relates to a task scheduling optimization method, device and equipment based on reinforcement learning and a medium. BACKGROUND

[0002] In the field of computer science and technology, efficient task scheduling plays a key role in improving the overall performance of the system and optimizing the resource utilization efficiency. Traditional task scheduling methods, such as rule-based scheduling (such as first-in-first-out, shortest job first, etc.) and scheduling methods using linear programming, branch and bound, etc. classic optimization algorithms, can achieve certain results when facing relatively simple and static task scenarios. However, with the development of technology, the types of tasks are becoming more and more complex, and the system running environment is changing frequently, such as the real-time fluctuation of virtual machine load in cloud computing environment, the dynamic increase and decrease of order tasks in industrial automation production line, etc. The traditional method has obvious limitations and cannot achieve precise and efficient task scheduling. Some existing technologies try to improve task scheduling by using machine learning methods, for example, patent CN202210569531.7 uses a deep neural network for job scheduling, which can output a scheduling strategy and a computing node queue according to job information. However, the deep neural network lacks effective learning and adaptive adjustment capabilities for long-term feedback from the environment when dealing with dynamically changing tasks in complex environments, and is prone to local optimal solutions, making it difficult to achieve continuous optimization of task scheduling from a global and long-term perspective. Reinforcement learning, as an advanced technology in the field of machine learning, can learn the optimal strategy by constantly interacting with the environment through an agent, with the goal of maximizing cumulative rewards, providing a new idea and effective way to solve the problem of task scheduling in complex dynamic environments. SUMMARY

[0003] Therefore, it is necessary to provide a task scheduling optimization method, device, equipment and medium based on reinforcement learning, which can realize precise mapping of resource state and task demand and overcome the adaptability of traditional static scheduling to complex scenarios.

[0004] In a first aspect, the application provides a task scheduling optimization method based on reinforcement learning, comprising:

[0005] Obtaining real-time system resource state data, determining the current dynamic environment characteristics and analyzing the task characteristic data to obtain the task priority and resource demand vector.

[0006] Using a reinforcement learning algorithm to process the task priority and resource demand vector, and determining the optimal task allocation scheme through policy iteration.

[0007] If the task allocation scheme meets the preset resource balance threshold, the task scheduling is executed and the scheduling execution result is obtained.

[0008] The performance index data is extracted from the scheduling execution result, and the reinforcement learning model is updated in combination with historical scheduling data to obtain an improved scheduling strategy.

[0009] In one of the embodiments, real-time system resource state data is acquired, current dynamic environment characteristics are determined, and task characteristic data is analyzed to obtain a task priority and a resource demand vector, including:

[0010] Real-time system resource state data is acquired, and key indicators including memory usage, processor occupancy, and network bandwidth attributes are extracted from the system resource state data.

[0011] Based on the preprocessed system resource state data, a time series analysis method is used to determine the current dynamic environment characteristics.

[0012] Based on the dynamic environment characteristics, a resource availability vector is calculated in combination with the key indicators to obtain a resource allocation weight.

[0013] Task characteristic data is acquired, and task attributes including task computational complexity, data transmission volume, and execution time limit are extracted from the task characteristic data.

[0014] A priority sorting algorithm is used for the task attributes to obtain a task priority sequence.

[0015] The task priority sequence and the resource allocation weight are fused to generate a task resource demand vector.

[0016] In one of the embodiments, the resource allocation weight is obtained through the following steps:

[0017] The key indicator data is standardized, and information entropy measuring the uncertainty of the indicators is calculated.

[0018] Based on the information entropy, an indicator entropy weight reflecting the weight proportion of the indicators in the resource state is calculated.

[0019] A grey correlation analysis is used to calculate the correlation degree of the indicators and the ideal resource state, and the resource allocation weight is obtained by integrating the indicator entropy weight and the correlation degree.

[0020] The resource allocation weight is calculated using the following formula:

[0021]

[0022] wherein, R j represents the resource allocation weight of the jth key indicator, W j represents the entropy weight of the jth key indicator, γ j represents the grey correlation degree of the jth key indicator and the ideal resource state, and m represents the total number of key indicators.

[0023] In one of the embodiments, a reinforcement learning algorithm is used to process the task priority and resource demand vectors, and an optimal task allocation scheme is determined through policy iteration, including:

[0024] A task feature matrix is extracted from the task priority and resource demand vectors; the task feature matrix includes task time limit and computing resource demand.

[0025] A state space and an action space of reinforcement learning are constructed according to the task feature matrix, and an initial policy model is obtained by parameterizing a deep neural network.

[0026] The resource utilization of the initial policy model is calculated by a reward function, and a reward value vector is obtained.

[0027] A proximal policy optimization algorithm is used to evaluate the iterative reward value vector, and an optimized task allocation constraint is generated.

[0028] According to the task allocation constraint, the resource demand and resource scheduling are adjusted, the node load balancing optimization resource allocation is calculated, and the optimal task allocation scheme is obtained.

[0029] In one of the embodiments, the resource utilization is calculated by the following formula:

[0030]

[0031] wherein η represents the resource utilization, m represents the total number of resource types, ω k represents the importance weight of the kth resource, n represents the total number of nodes in the system, u i,k represents the utilization of the kth resource on the ith node, represents the used amount, represents the total capacity, and H represents the Shannon entropy of system resource utilization.

[0032] In one of the embodiments, performance indicator data is extracted from the scheduling execution result, and the reinforcement learning model is updated with historical scheduling data to obtain an improved scheduling strategy, including:

[0033] Performance indicator data is extracted from the scheduling execution result; the performance indicator data includes task completion time and resource occupation.

[0034] Task execution efficiency is calculated according to the performance indicator data; the task execution efficiency is determined by the ratio of task completion time to standard time.

[0035] A priority replay buffer is constructed according to the task execution efficiency and historical scheduling data, and the network parameters of the initial policy model of reinforcement learning are updated using an experience replay mechanism and a gradient descent algorithm to obtain an improved scheduling strategy that integrates historical experience.

[0036] In one of the embodiments, the task execution efficiency is calculated by the following formula:

[0037]

[0038] wherein e i represents the task execution efficiency of the i-th task after considering time balance and dynamic adjustment of system load, e' i represents the basic execution efficiency, t actual,i represents the actual completion time of the i-th task, t standard,i represents the standard reference time of the i-th task, S t represents the time entropy, and a represents a dynamic reference adjustment coefficient.

[0039] In a second aspect, the present application further provides a task scheduling optimization device based on reinforcement learning, which comprises:

[0040] a task analysis module, configured to acquire real-time system resource state data, determine current dynamic environment characteristics and analyze task characteristic data, and obtain a task priority and a resource demand vector.

[0041] a distribution decision module, configured to process the task priority and the resource demand vector by using a reinforcement learning algorithm, judge an optimal task distribution scheme by policy iteration, and execute task scheduling and acquire a scheduling execution result if the task distribution scheme meets a preset resource balance threshold.

[0042] a strategy optimization module, configured to extract performance index data from the scheduling execution result, update a reinforcement learning model in combination with historical scheduling data, and obtain an improved scheduling strategy.

[0043] In a third aspect, the present application further provides a computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the foregoing method when executing the computer program.

[0044] In a fourth aspect, the present application further provides a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the foregoing method.

[0045] The task scheduling optimization method, device, computer device and storage medium based on reinforcement learning first acquire system resource state data in real time, determine dynamic environment characteristics through preprocessing and time series analysis, analyze task characteristic data, generate a task priority sequence and a resource demand vector by using a priority sorting algorithm and a resource evaluation model. Then, a state space and an action space are constructed by using a reinforcement learning algorithm, and an optimal task allocation scheme is generated by using policy iteration and a reward function optimization. If the scheme meets a resource balance threshold, the scheduling is performed and performance indicators are collected. Finally, real-time indicators and historical data are fused, the reinforcement learning model parameters are updated by using experience replay and gradient descent, and an improved scheduling strategy is formed in a closed loop optimization. Precise mapping of resource states and task demands is achieved, and the problem of insufficient adaptability of traditional static scheduling to complex scenarios is overcome. BRIEF DESCRIPTION OF DRAWINGS

[0046] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the related art, the drawings needed to be used in the embodiments or the related art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0047] Figure 1 The flowchart of the task scheduling optimization method based on reinforcement learning provided by the embodiments of the present application is shown in

[0048] Figure 2 The structural block diagram of the task scheduling optimization device based on reinforcement learning provided by the embodiments of the present application is shown in DETAILED DESCRIPTION

[0049] In order to make the purpose, technical solutions and advantages of the present application more clear, the present application will be further described in detail below in combination with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and do not limit the present application.

[0050] In one of the embodiments, as shown in Figure 1 The task scheduling optimization method based on reinforcement learning provided by the present application can include the following steps:

[0051] In step S101, real-time system resource state data is acquired, current dynamic environment characteristics are determined, and task characteristic data is analyzed to obtain task priority and resource demand vector.

[0052] First, the real-time system resource state data such as CPU usage, memory remaining capacity, storage read-write rate, network bandwidth, etc. are collected through sensors, system monitoring tools, etc. and the original data is preprocessed such as denoising, standardization, etc. Based on the preprocessed resource data, time series analysis, sliding window statistics, etc. are used to identify dynamic environmental characteristics such as fluctuation trend, peak characteristics, etc. of resource usage. At the same time, the characteristics data such as computational complexity, data transmission volume, execution time limit, priority label, etc. of the task are obtained, and the emergency level and resource demand intensity of the task are quantified through analytic hierarchy process, weighted priority sorting algorithm, etc. Finally, structured data containing task priority sequence and resource demand vector are generated.

[0053] Step S102, the task priority and resource demand vector are processed by using reinforcement learning algorithm, and the optimal task allocation scheme is determined by policy iteration.

[0054] Specifically, the task priority sequence and resource demand vector obtained in the first step are mapped to the state space of reinforcement learning (such as resource load matrix, task attribute tensor), and the action space is defined as the allocation combination of tasks and computing nodes. A deep neural network (such as Actor-Critic architecture) is constructed to parameterize the initial policy model, and a reward function containing resource utilization, task completion rate, load balancing degree, etc. is designed to make the agent constantly try and error in the interaction with the environment. The policy is iteratively optimized by using policy gradient, experience replay, etc. The probability distribution of task allocation is adjusted by evaluating the reward value vector of the current policy, and gradually converges to the optimal task allocation scheme that maximizes the long-term cumulative reward in the dynamic environment.

[0055] Step S103, if the task allocation scheme meets the preset resource balance threshold, the task scheduling is executed and the scheduling execution result is obtained.

[0056] Specifically, the resource balance evaluation indicators (such as node load standard deviation, resource utilization variance, etc.) and corresponding thresholds are set to verify the task allocation scheme generated in the second step. If the scheme meets the balance requirement (such as the difference in node load is not more than 20%), the task scheduling module is triggered to allocate the task to the target computing node for execution. During the scheduling process, the execution data such as task completion time, actual resource occupation, node response delay, etc. are collected in real time to form the scheduling execution result containing timestamp, task ID, resource consumption details.

[0057] Step S104, the performance indicator data is extracted from the scheduling execution result and combined with the historical scheduling data to update the reinforcement learning model, and the improved scheduling strategy is obtained.

[0058] Specifically, the performance indicators such as task execution efficiency (actual time consumption / standard time consumption), resource utilization (used resources / total resources), and load balancing degree are extracted from the execution log of the third step, and are aligned with the same type of task data in the historical scheduling database. A priority experience replay mechanism is adopted to include high-value samples (such as timeout tasks and cases with significant resource waste) into the training set according to the weight, and the network parameters of the reinforcement learning model (such as the policy output layer of the Actor network and the value function parameters of the Critic network) are updated through the gradient descent algorithm. After iterative optimization, an improved scheduling strategy is generated by combining the current environmental characteristics and historical experience.

[0059] The above-mentioned task scheduling optimization method based on reinforcement learning first acquires real-time system resource state data, determines the dynamic environmental characteristics through preprocessing and time series analysis, analyzes the task characteristic data, and generates a task priority sequence and a resource demand vector using a priority sorting algorithm and a resource evaluation model. Then, a state space and an action space are constructed using a reinforcement learning algorithm, and an optimal task allocation scheme is generated through policy iteration and reward function optimization. If the scheme meets the resource balancing threshold, the scheduling is performed and performance indicators are collected. Finally, real-time indicators and historical data are fused, and the reinforcement learning model parameters are updated through experience replay and gradient descent to form an improved scheduling strategy in a closed-loop optimization. The precise mapping of resource state and task demand is achieved, and the problem of insufficient adaptability of traditional static scheduling to complex scenarios is overcome.

[0060] In one of the embodiments, acquiring real-time system resource state data, determining the current dynamic environmental characteristics and analyzing the task characteristic data to obtain the task priority and the resource demand vector can include the following steps:

[0061] Step S201, acquiring real-time system resource state data, and extracting key indicators including memory usage, processor occupancy and network bandwidth attributes from the system resource state data.

[0062] Step S202, determining the current dynamic environmental characteristics based on the preprocessed system resource state data using a time series analysis method.

[0063] Preferably, the extracted key indicator data such as memory usage, processor occupancy and network bandwidth are first sorted in time sequence to construct a multi-dimensional time series dataset. For example, 2 hours of resource data are continuously collected at a sampling interval of 1 minute to form a sequence containing 120 time steps. Subsequently, local feature extraction is performed on the data using a sliding window algorithm, such as calculating the mean, variance and change slope of resource occupancy within each 10-minute window. At the same time, random noise in the data is eliminated by exponential smoothing to highlight the long-term trend of resource usage. Finally, the fluctuation amplitude, frequency of change and trend direction of resource occupancy are quantified as dynamic environmental feature vectors reflecting the current system load pressure and resource fluctuation pattern.

[0064] Step S203, calculate the resource availability vector based on the dynamic environmental features and key indicators to obtain the resource allocation weight.

[0065] Preferably, taking a cloud computing cluster scenario as an example, after obtaining the dynamic environmental features of periodic fluctuations in processor load and persistent tension in memory resources, the processor occupancy, memory usage and other key indicator data are combined to calculate using an entropy weight-gray correlation analysis composite model. First, the entropy weight of each indicator is calculated by the information entropy formula to quantify the uncertainty of the indicator data. For example, if the memory usage data has a high degree of dispersion, the corresponding entropy weight is larger, indicating that it has a greater impact on resource allocation decisions. Next, the gray correlation analysis is used to calculate the correlation degree of each indicator with the ideal resource state (e.g. processor occupancy below 60% and memory usage below 70%). Finally, the entropy weight and correlation degree are normalized and weighted to obtain the resource allocation weight vector of processor, memory and network bandwidth [0.4, 0.35, 0.25].

[0066] Step S204, obtain task characteristic data and extract task attributes including task computational complexity, data transmission volume and execution time limit from the task characteristic data.

[0067] Step S205, apply a priority sorting algorithm to the task attributes to obtain a task priority sequence.

[0068] Step S206, fuse the task priority sequence and resource allocation weight to generate a task resource demand vector.

[0069] Specifically, first, the real-time resource state data such as CPU, memory, storage, etc. are obtained through system monitoring tools, from which key indicators such as memory usage, processor occupancy, network bandwidth, etc. are extracted, and preprocessed such as denoising and standardization. Based on the preprocessed data, time series analysis methods such as sliding window and exponential smoothing are used to identify the fluctuation trend and peak characteristics of resource usage, and to determine the dynamic environment characteristics. Then, combining the environment characteristics and key indicators, the resource availability vector is calculated through the entropy weight-gray correlation analysis composite model to obtain the allocation weight of each resource type. At the same time, task characteristic data is obtained, and attributes such as computational complexity, data transmission volume, and execution time limit are extracted. The analytic hierarchy process or weighted priority algorithm is used to generate a task priority sequence, and finally the priority sequence and resource allocation weight are fused to form a structured vector containing task resource requirements.

[0070] This embodiment realizes accurate characterization of dynamic environment characteristics through multi-dimensional resource indicator extraction and time series analysis, avoiding the problem of insufficient adaptability of traditional static analysis to resource fluctuations. The resource weight calculation method based on entropy weight-gray correlation combines data uncertainty and ideal state matching degree, and compared with a single weight model, it can better reflect the actual availability of resources. The priority sorting of task attributes and the fusion mechanism of resource weights make the generated resource requirement vector take into account both the urgency of the task and the scarcity of resources. This method is especially suitable for resource dynamic change scenarios such as cloud computing and edge computing, and can effectively improve the rationality of resource allocation and the accuracy of task scheduling.

[0071] In one of the embodiments, the resource allocation weight can be obtained by the following steps:

[0072] Step S301, standardizing the key indicator data and calculating the information entropy to measure the uncertainty of the indicator.

[0073] Step S302, calculating the index entropy weight reflecting the weight proportion of the index in the resource state based on the information entropy.

[0074] Step S303, calculating the correlation degree of the index and the ideal resource state using gray correlation analysis, and obtaining the resource allocation weight by integrating the index entropy weight and the correlation degree.

[0075] The resource allocation weight is calculated using the following formula:

[0076]

[0077] where R j is the resource allocation weight of the jth key indicator, W j is the entropy weight of the jth key indicator, γ j is the gray correlation degree of the jth key indicator and the ideal resource state, and m represents the total number of key indicators.

[0078] Preferably, the correlation degree of the indicators and the ideal resource state is calculated by using grey correlation analysis, and the resource allocation weight is obtained by comprehensively considering the entropy weight. Taking the CPU occupancy rate, memory usage rate, and network bandwidth utilization rate of a certain edge computing node as examples, the ideal resource state is defined as: CPU occupancy rate ≤ 50%, memory usage rate ≤ 60%, and network bandwidth utilization rate ≤ 70%. For the currently collected indicator data (such as CPU occupancy rate 65%, memory usage rate 75%, and network bandwidth utilization rate 40%), the absolute difference between each indicator and the ideal state is calculated to obtain the sequence {15, 15, 30}. Then, the resolution coefficient ρ = 0.5 is determined, the minimum absolute difference minΔ = 15 and the maximum absolute difference maxΔ = 30 are calculated, and the grey correlation degree formula is used to obtain the correlation degrees of CPU, memory, and network as follows: Assuming that the entropy weight vector obtained through steps S301-S302 is [0.35, 0.4, 0.25], the resource allocation weight formula is substituted, the numerators are calculated as 0.35 × 1 = 0.35, 0.4 × 1 = 0.4, and 0.25 × 0.667 ≈ 0.167, and the denominator is 0.35 + 0.4 + 0.167 ≈ 0.917. Finally, the resource allocation weight vector is obtained as [0.35 / 0.917 ≈ 0.382, 0.4 / 0.917 ≈ 0.436, 0.167 / 0.917 ≈ 0.182], that is, the weights of CPU, memory, and network are 38.2%, 43.6%, and 18.2% respectively, which reflects that the availability of memory and CPU is more critical to the scheduling decision under the current resource state.

[0079] Specifically, first, the key indicator data such as memory usage rate and processor occupancy rate are standardized to eliminate the influence of dimensional differences, for example, the data is converted into a distribution with a mean of 0 and a standard deviation of 1 by Z-Score standardization. Based on the standardized data, the information entropy of each indicator is calculated to measure its uncertainty, and then the entropy weight of the indicator is calculated according to the information entropy, and the size of the entropy weight reflects the proportion of the weight of the indicator in the resource state. Finally, the correlation degree of each indicator and the ideal resource state is calculated by using grey correlation analysis, and the final resource allocation weight vector is obtained by comprehensively considering the entropy weight and the correlation degree through the formula.

[0080] ​The embodiment solves the problem of inconsistent dimensions of different indicators through standardization processing, ensuring the consistency of weight calculation. The introduction of information entropy and entropy weight objectively quantifies the importance of indicators based on the uncertainty of the data itself, avoiding the bias of subjective weighting. The gray correlation analysis measures the degree of resource availability by comparing the differences between actual indicators and ideal states, making the weight calculation take into account both data characteristics and business objectives. The composite model of entropy weight and correlation degree utilizes the objectivity of data-driven and incorporates subjective expectations of ideal resource states, making it more adaptable to dynamic resource environments than single weight methods. This weight calculation mechanism can accurately reflect the actual importance of resources such as memory and processor under different load scenarios, providing a scientific and quantitative basis for resource allocation in task scheduling, effectively improving resource utilization efficiency and system balance.

[0081] In one embodiment, a reinforcement learning algorithm is used to process the task priority and resource demand vector, and the optimal task allocation scheme is determined through policy iteration, which can include the following steps:

[0082] Step S401, extract the task feature matrix from the task priority and resource demand vector; the task feature matrix includes task time limit and computing resource demand.

[0083] Step S402, construct the state space and action space of reinforcement learning according to the task feature matrix, and obtain the initial policy model through deep neural network parameterization.

[0084] Preferably, assuming that the task feature matrix contains the time limit (unit: seconds) and computing resource demand (CPU core number) of 3 tasks, such as matrix where each row corresponds to the [remaining time limit, CPU core demand] of a task. The state space is defined as the concatenation tensor of the CPU utilization of each node in the current cluster (such as [0.6, 0.4, 0.5] representing the load of 3 nodes) and the task feature matrix, with a dimension of 3 (number of nodes) + 3x2 (task features) = 9. The action space is defined as the allocation mapping of tasks to nodes, with a total of 3 tasks x 3 nodes = 9 possible actions, represented by one-hot encoding (such as [1, 0, 0, 0, 0, 0, 0, 0, 0] representing task 1 allocated to node 1). An initial policy model is parameterized by building a deep neural network containing 2 fully connected layers (input layer 9 dimensions, hidden layer 64 dimensions, output layer 9 dimensions), and the network weights are randomly initialized. Using the task feature matrix and node load data as input, the output is the probability distribution of each action, for example, the output vector [0.2, 0.1, 0.1, 0.3, 0.1, 0.05, 0.1, 0.05, 0.0], corresponding to the selection probability of 9 actions, thus forming an initial task allocation strategy model.

[0085] Step S403, calculate the resource utilization of the initial strategy model by the reward function to obtain a reward value vector.

[0086] Step S404, evaluate the iterative reward value vector by using a proximal policy optimization algorithm to generate an optimized task allocation constraint.

[0087] Preferably, assuming that the reward value vector obtained by the initial strategy model is [0.3, 0.5, -0.2], corresponding to the reward feedback of three task allocation attempts respectively. The PPO algorithm first calculates the ratio r(θ) of the action probabilities under the new and old strategies, for example, the probability of a certain task allocation action under the current strategy θ is 0.4, while the probability under the updated strategy is 0.3, then Next, the policy update range is limited by the clipping mechanism, and the objective function LCLIP(θ) is defined as min(r(θ)·A, clip(r(θ), 1-∈, 1+∈)·A), where A represents the advantage function value (measuring the relative value of actions), and ∈ represents the clipping parameter (usually 0.2). If A = 0.6, then clip(1.33, 0.8, 1.2) = 1.2, and the final objective function value is min(1.33×0.6, 1.2×0.6) = 0.72. By iteratively calculating the objective function and updating the deep neural network parameters using stochastic gradient descent, the policy is gradually optimized. When the cumulative gain of the reward value vector reaches a preset threshold (such as an average reward increase of more than 5%) or the number of iterations reaches an upper limit (such as 100 times), the optimization is stopped, and the task allocation constraint containing the task allocation priority, resource allocation limit, etc. is output, such as "task 1 needs to be executed on node 2 and the CPU resource should not be less than 3 cores".

[0088] Step S405, adjust the resource demand and resource scheduling according to the task allocation constraint, calculate the node load balancing optimized resource allocation, and obtain the optimal task allocation scheme.

[0089] Specifically, first, the task time limit, computing resource demand, and other key information are extracted from the task priority and resource demand vector to construct a structured task feature matrix. For example, the remaining execution time, required CPU core number, memory capacity, and other data of the task are arranged in the form of a tensor. Based on the matrix, the state space of reinforcement learning is defined as the joint representation of resource load and task attributes, the action space is the allocation combination of tasks to computing nodes, the policy is parameterized through a deep neural network (such as a multi-layer perceptron or Transformer architecture) to obtain an initial policy model. Subsequently, a reward function including resource utilization, task completion rate, and other indicators is used to quantitatively evaluate the allocation scheme output by the model, generating a reward value vector. The proximal policy optimization (PPO) algorithm is used to iteratively update the reward value vector, and the performance degradation is avoided by limiting the policy update step, generating an optimized task allocation constraint condition. Finally, the resource demand is dynamically adjusted according to the constraint condition, and the resource is reallocated in combination with the node load balancing algorithm, thereby obtaining an optimal task allocation scheme that meets the system efficiency and balance requirements.

[0090] The present embodiment converts complex task and resource information into structured data by constructing a task feature matrix, providing standardized input for reinforcement learning. The policy model based on deep neural network can automatically learn the complex mapping relationship in high-dimensional data, and has stronger generalization ability than traditional heuristic algorithms. Combined with the iterative optimization mechanism of the PPO algorithm, the performance fluctuation in policy update can be effectively avoided, ensuring the stability and convergence of the allocation scheme. Through the reward function, the resource utilization and task completion rate are comprehensively considered, and the system efficiency and task timeliness are taken into account; and based on the load balancing of resource allocation optimization, node overload or resource waste can be avoided, and the overall throughput of the system can be improved.

[0091] In one of the embodiments, the resource utilization can be calculated by the following formula:

[0092]

[0093] wherein η represents the resource utilization, m represents the total number of resource types, ω k represents the importance weight of the kth resource, n represents the total number of nodes in the system, u i,k represents the utilization rate of the kth resource on the ith node, represents the used amount, represents the total capacity, and H represents the Shannon entropy of system resource utilization.

[0094] Preferably,

[0095] The embodiment reflects the overall use efficiency of various resources by weighted average utilization rate. The introduction of weight can be differentiated according to the importance of resources and adapt to the needs of different business scenarios. The application of Shannon entropy quantifies the balance degree of resource allocation between nodes, avoiding the problem of overloading of some nodes and idling of some nodes due to excessive concentration of resources. By combining the calculation method of efficiency and balance, compared with single dimension evaluation, the system resource state can be more comprehensively reflected. When the index is used for the design of the reward function of reinforcement learning, it can guide the policy model to pursue resource efficient utilization while considering load balancing, optimize task allocation scheme, and improve the overall performance and stability of the system, which is suitable for multi-resource collaborative scheduling scenarios such as cloud computing and data center.

[0096] In one of the embodiments, the performance index data is extracted from the scheduling execution result to update the reinforcement learning model in combination with the historical scheduling data to obtain an improved scheduling strategy, which can include the following steps:

[0097] Step S501, extracting performance index data from the scheduling execution result; the performance index data includes task completion time and resource occupation.

[0098] Step S502, calculating task execution efficiency according to the performance index data; the task execution efficiency is determined by the ratio of task completion time to standard time.

[0099] Step S503, constructing a priority playback buffer according to the task execution efficiency and historical scheduling data, updating the network parameters of the initial policy model of reinforcement learning by using the experience playback mechanism and gradient descent algorithm to obtain an improved scheduling strategy that integrates historical experience.

[0100] Preferably, assuming that the current scheduling execution is completed, the execution efficiency values of 10 tasks are calculated by step S502, which are E=[1.2, 0.8, 1.5, 0.9, 1.1, 0.7, 1.3, 0.6, 1.4, 1.0], wherein the values less than 1 represent task timeout completion. At the same time, 100 records of the same type of current task are retrieved from the historical scheduling database, including the execution efficiency, resource usage and scheduling strategy parameters of the corresponding task. After merging the current data and historical data, the priority score of each sample is calculated according to the task execution efficiency and resource utilization deviation. For example, the priority formula is defined as P i = |1-E i | + resource utilization deviation i , wherein E iThe execution efficiency of the i-th task is represented, and the resource utilization deviation is measured by the difference between the actual utilization and the average utilization. After calculation, the priority vector P = [0.2, 0.2, 0.5, 0.1, 0, 0.3, 0.3, 0.4, 0.4, 0] of each sample is obtained, which is used as a weight to construct a priority playback buffer. Then, the experience replay mechanism is used to sample data from the buffer according to the priority, and 32 samples are extracted each time to form a training batch. The gradient of the loss function (such as mean square error) with respect to the initial policy model network parameters of the reinforcement learning is calculated by the gradient descent algorithm, and the model is updated. After 100 iterations of training, the model parameters converge, and the improved scheduling strategy that integrates historical experience is output. The processing capacity of this strategy for overdue tasks is improved by 20% compared with the initial model.

[0101] Specifically, first, the performance indicator data such as task completion time and actual resource occupation are extracted from the scheduling execution log. The task completion time is the actual time consumed from submission to completion of the task, and the resource occupation includes the actual use of resources such as CPU cores and memory capacity. Then, the task execution efficiency is calculated according to the performance indicators, specifically the ratio of the actual completion time of the task to the standard reference time, i.e. efficiency value = standard time / actual time. If this value is greater than 1, it means that the task is completed ahead of schedule. Finally, the current task execution efficiency data is associated with the same type of task data in the historical scheduling database, and a priority playback buffer is constructed according to the time efficiency difference and resource utilization deviation. High-value samples (such as overdue tasks or resource waste cases) are given higher weights, and samples are sampled through the experience replay mechanism and combined with the gradient descent algorithm to update the network parameters of the initial policy model of reinforcement learning, thereby obtaining an improved scheduling strategy that integrates historical experience.

[0102] This embodiment provides a clear evaluation benchmark for scheduling strategy optimization by quantifying task execution efficiency, avoiding the bias caused by subjective judgment. The construction mechanism of the priority playback buffer dynamically adjusts the sample weights based on data value, enabling the model to more efficiently learn from historical successes and failures, significantly improving learning efficiency compared to random sampling. Combined with the experience replay and gradient descent parameter update method, the historical scheduling knowledge is effectively reused, enabling the improved strategy to adapt more quickly to similar scenarios and reducing the cost of repeated trial and error. This method is particularly suitable for scenarios with diverse task types and abundant historical data. Through continuous accumulation and iteration, the adaptability of the scheduling strategy to complex environments can be gradually improved, optimizing resource utilization efficiency while ensuring task timeliness and achieving continuous evolution of system performance.

[0103] In one embodiment, the task execution efficiency can be calculated by the following formula:

[0104]

[0105] where ei represents the task execution efficiency of the ith task considering time balance and system load dynamic adjustment, e i represents the basic execution efficiency, t actual,i represents the actual completion time of the ith task, t standard,i represents the standard reference time of the ith task, S t represents the time entropy, and a represents the dynamic reference adjustment coefficient.

[0106] Preferably, where h represents the number of system resource types, load k represents the current load rate of the kth type of resource, load k ,max represents the maximum allowable load rate of the kth type of resource, represents an adjustment parameter (usually 1), which controls the influence strength of the load on the reference.

[0107] The embodiment breaks through the traditional single time ratio evaluation mode, and realizes double dynamic adjustment of efficiency evaluation through time entropy and dynamic reference coefficient. The introduction of time entropy makes the evaluation result reflect the fluctuation characteristics of task execution time, avoids the evaluation deviation caused by individual task time anomaly, and enhances the stability of the efficiency index; the dynamic reference adjustment coefficient adjusts the evaluation standard in real time according to the system resource load, automatically relaxes the requirement when the resource is tight, and improves the standard when the resource is sufficient, so as to ensure that the evaluation result meets the actual running condition. Compared with the traditional method, the formula can more accurately depict the real efficiency of task execution, provide more scientific quantitative basis for scheduling strategy optimization, and is especially suitable for scenes with dynamic resource changes and high task timeliness requirements, effectively improving the rationality of scheduling decision and the overall performance of the system.

[0108] In one embodiment, as shown in Figure 2 The application also provides a task scheduling optimization device based on reinforcement learning, which can include:

[0109] The task analysis module 601 is used to acquire real-time system resource state data, determine the current dynamic environment characteristics and analyze the task characteristic data, obtain the task priority and resource demand vector.

[0110] The allocation decision module 602 is used to process the task priority and resource demand vector by using a reinforcement learning algorithm, judge the optimal task allocation scheme through policy iteration, and execute task scheduling and obtain scheduling execution results if the task allocation scheme meets the preset resource balance threshold.

[0111] The strategy optimization module 603 is used to extract performance index data from the scheduling execution results, update the reinforcement learning model combined with historical scheduling data, and obtain an improved scheduling strategy.

[0112] The task analysis module determines dynamic environment characteristics through time series analysis of real-time collected system resource state data, and analyzes task characteristic data, and generates a task priority sequence and a resource demand vector using a priority sorting algorithm and a resource evaluation model. The allocation decision module inputs the above vectors into a reinforcement learning algorithm, constructs a state space and an action space, and generates an optimal task allocation scheme through policy iteration. If the scheme meets a resource balance threshold, the scheme is executed and performance indicators are collected. The policy optimization module extracts task completion time, resource occupation, and other data from the execution results, updates reinforcement learning model parameters in combination with historical scheduling data, and forms an improved scheduling policy. The method solves the problem of insufficient adaptability to dynamic environments in traditional methods, effectively balances system load, improves resource utilization and task completion rate, and realizes continuous optimization of scheduling policy and steady improvement of system performance.

[0113] It should be understood that, although each step in the flowchart involved in each embodiment as described above is displayed in sequence according to the arrow, these steps are not necessarily executed in sequence according to the arrow. Unless otherwise specified herein, the execution of these steps is not strictly limited in sequence, and these steps can be executed in other sequences. Moreover, at least part of the steps in the flowchart involved in each embodiment as described above can include multiple steps or stages, which are not necessarily executed at the same time, but can be executed at different times, and the execution sequence of these steps or stages is not necessarily sequential, but can be alternately or alternately executed with at least part of other steps or steps or stages in other steps.

[0114] In one embodiment, a computer device is provided, comprising a memory and a processor, the memory stores a computer program, and the processor implements the steps of the reinforcement learning-based task scheduling optimization method, device, equipment and medium as described above when executing the computer program.

[0115] In one embodiment, a computer readable storage medium is provided, which stores a computer program, and the computer program is executed by a processor to implement the steps in each method embodiment described above.

[0116] For the device embodiment, since it basically corresponds to the method embodiment, the relevant part can be seen from the part of the method embodiment. The device embodiment described above is only schematic, wherein the components shown as separate components can or can not be physically separate, and the components shown as a unit can or can not be a physical unit, i.e., can be located in one place, or can be distributed to multiple network units. Part or all of the modules can be selected according to actual needs to achieve the purpose of the present disclosure. Those skilled in the art can understand and implement it without creative labor.

[0117] The above-described embodiments only express several implementation manners of the present application, which are described in detail, but cannot be understood as a limitation on the patent scope of the application. It should be pointed out that, for those skilled in the art, without departing from the concept of the present application, several modifications and improvements can be made, which are all within the protection scope of the present application.

Claims

1. A task scheduling optimization method based on reinforcement learning, characterized in that: The method comprises: Obtain real-time system resource status data, determine the current dynamic environment characteristics and analyze task characteristic data to obtain task priority and resource demand vector; Using a reinforcement learning algorithm to process the task priorities and resource demand vectors, and determining the optimal task allocation solution through strategy iteration; If the task allocation plan meets the preset resource balancing threshold, the task scheduling is performed and the scheduling execution result is obtained; The performance indicator data is extracted from the scheduling execution result and combined with the historical scheduling data to update the reinforcement learning model to obtain an improved scheduling strategy.

2. The method according to claim 1, characterized in that The method of acquiring real-time system resource status data, determining current dynamic environment characteristics and analyzing task characteristic data to obtain task priority and resource requirement vectors includes: Acquire real-time system resource status data, and extract key indicators including memory usage, processor occupancy, and network bandwidth attributes from the system resource status data; Determining current dynamic environment characteristics based on the pre-processed system resource status data using a time series analysis method; Calculating a resource availability vector based on the dynamic environment characteristics and the key indicators to obtain a resource allocation weight; Acquiring task characteristic data, and extracting task attributes including task computation complexity, data transmission volume, and execution time limit from the task characteristic data; Applying a priority sorting algorithm to the task attributes to obtain a task priority sequence; The task priority sequence and resource allocation weights are integrated to generate a task resource requirement vector.

3. The method according to claim 2, characterized in that The resource allocation weights are obtained by the following steps: Standardize the data of each key indicator and calculate the information entropy of the uncertainty of the indicator; Calculate the indicator entropy weight reflecting the weight ratio of the indicator in the resource status based on the information entropy; Grey correlation analysis is used to calculate the correlation between the indicator and the ideal resource state, and the resource allocation weight is obtained by combining the entropy weight of the indicator and the correlation; The resource allocation weights are calculated using the following formula: Among them, R j represents the resource allocation weight of the jth key indicator, W j represents the entropy weight of the jth key indicator, γ j It represents the grey correlation between the jth key indicator and the ideal resource state, and m represents the total number of key indicators.

4. The method according to claim 1, wherein The process of using a reinforcement learning algorithm to process the task priorities and resource requirement vectors and determining the optimal task allocation solution through policy iteration includes: Extracting a task feature matrix from the task priority and resource requirement vector; the task feature matrix includes task time limit and computing resource requirement; Constructing the state space and action space of reinforcement learning based on the task feature matrix, and obtaining the initial strategy model through deep neural network parameterization; Calculating the resource utilization of the initial strategy model using a reward function to obtain a reward value vector; Using a proximal policy optimization algorithm to evaluate and iterate the reward value vector to generate an optimized task allocation constraint; According to the task allocation constraints, resource demand and resource scheduling are adjusted, and resource allocation is optimized by computing node load balancing to obtain the optimal task allocation solution.

5. The method according to claim 4, characterized in that The resource utilization is calculated using the following formula: Among them, η represents resource utilization, m represents the total number of resource types, ω k represents the importance weight of the k-th resource, n represents the total number of nodes in the system, u i,k represents the utilization rate of the k-th resource on the i-th node, Indicates the amount used. represents the total capacity, and H represents the Shannon entropy of system resource utilization.

6. The method according to claim 1, characterized in that The step of extracting performance indicator data from the scheduling execution results and combining it with historical scheduling data to update the reinforcement learning model to obtain an improved scheduling strategy includes: Extracting performance indicator data from the scheduling execution result; the performance indicator data includes task completion time and resource occupancy; Calculating the task execution efficiency based on the performance indicator data; the task execution efficiency is determined by the ratio of the task completion time to the standard time; A priority replay buffer is constructed according to the task execution efficiency and historical scheduling data, and the network parameters of the reinforcement learning initial strategy model are updated using the experience replay mechanism and gradient descent algorithm to obtain an improved scheduling strategy that integrates historical experience.

7. The method according to claim 6, characterized in that The task execution efficiency is calculated using the following formula: Among them, e i represents the execution efficiency of the i-th task after considering time balance and dynamic adjustment of system load, e' i Indicates the basic execution efficiency, t actual,i represents the actual completion time of the i-th task, t standard,i represents the standard reference time of the i-th task, S t represents time entropy, and α represents the dynamic benchmark adjustment coefficient.

8. A task scheduling optimization device based on reinforcement learning, characterized in that: The device comprises: The task analysis module is used to obtain real-time system resource status data, determine the current dynamic environment characteristics and analyze task characteristic data to obtain task priority and resource requirement vector; An allocation decision module is configured to process the task priorities and resource demand vectors using a reinforcement learning algorithm, determine the optimal task allocation solution through policy iteration, and execute task scheduling and obtain a scheduling execution result if the task allocation solution meets a preset resource balancing threshold; The strategy optimization module is used to extract performance indicator data from the scheduling execution results and update the reinforcement learning model in combination with historical scheduling data to obtain an improved scheduling strategy.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • A job scheduling method, apparatus, and device based on reinforcement learning

    CN114675975B

Cited By

  • Multi-angle model operation automatic configuration optimization method and system based on big data

    CN121210138A

  • Big data-based multi-angle model job automatic configuration optimization method and system

    CN121210138B

  • Artificial intelligence-based task scheduling method and related device thereof

    CN121255470A

  • Optimized scheduling method based on automatic production scheduling of furniture production and manufacturing

    CN121352384A

  • Computer scheduling method in multi-task operation environment state

    CN121433906A