A big data dynamic allocation and optimization scheduling method based on reinforcement learning

By adopting a hierarchical model based on reinforcement learning in the big data processing system for dynamic resource allocation and scheduling, the problems of unbalanced resource allocation and low scheduling efficiency in the existing technology are solved, and efficient and adaptive resource utilization and system performance improvement are achieved.

CN119311407BActive Publication Date: 2025-05-23BEIJING XINDAYI TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411347525.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-26
Publication Date
2025-05-23
Estimated Expiration
2044-09-26

AI Technical Summary

Technical Problem

Resource allocation and scheduling in existing big data processing systems rely on static or preset rules, and it is difficult to cope with dynamic load fluctuations and environmental changes, resulting in unbalanced resource allocation, low scheduling efficiency and system performance degradation.

Method used

Adopting a hierarchical model based on reinforcement learning, high-level strategies are responsible for global task planning and resource allocation, and low-level strategies are responsible for node-level task execution and resource optimization. Strategy updates are carried out through real-time monitoring and feedback data to achieve dynamic resource allocation and scheduling.

Benefits of technology

It significantly improves the system's resource utilization, scheduling efficiency and adaptability, avoids the problems of idle or overload of node resources, realizes the optimization of global load balancing, and improves the stability and robustness of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119311407B_ABST
    Figure CN119311407B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for dynamic allocation and optimization scheduling of big data based on reinforcement learning, S1, constructing a real-time operation data set; S2, preprocessing the real-time operation data set to construct a reinforcement learning environment model; S3, initializing a hierarchical reinforcement learning model; S4, generating an optimal data allocation and scheduling strategy based on the hierarchical reinforcement learning model; S5, adjusting the data allocation and task scheduling of each computing node in real time to balance the system load and maximize resource utilization; S6, the hierarchical reinforcement learning model continuously monitors the system state and dynamically updates the strategy, and adaptively adjusts the resource allocation and scheduling in time when the system load or environment changes; S7, through multiple rounds of iterative optimization, the hierarchical reinforcement learning model gradually converges to the optimal strategy. The present invention realizes the automatic optimization of resource allocation and scheduling strategies through the application of reinforcement learning, and significantly improves the resource utilization, scheduling efficiency and adaptability of the system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of big data technology, and in particular to a big data dynamic allocation and optimization scheduling method based on reinforcement learning. Background Art

[0002] In the prior art, resource allocation and scheduling in big data processing systems usually rely on static or preset rules. The static or preset rules are set when the system is initialized and are rarely adjusted dynamically during operation. The static resource allocation method can meet system requirements when the data load is relatively stable, but with the explosive growth of big data volume and the complexity of computing tasks, static rules are difficult to cope with frequent load fluctuations, resulting in uneven system resource allocation. Some nodes may have performance degradation due to overload, while other nodes may have idle resources due to insufficient load, resulting in reduced overall system efficiency. In addition, static scheduling rules are also unable to adapt to dynamic changes in the computing environment, especially in distributed computing environments. Some nodes affect the overall performance due to communication delays or resource bottlenecks, further increasing the bottleneck problem of the scheduling system.

[0003] In order to solve the above problems, some existing technologies have proposed algorithm-based resource optimization methods, such as reallocating resources through load balancing algorithms to improve the resource utilization and response speed of the system. However, the load balancing algorithm is usually only triggered when the system load fluctuates significantly, and it is unable to continuously track and predict future load changes, which easily leads to resource adjustment delays and ultimately leads to system performance degradation. In addition, most existing scheduling algorithms lack adaptive learning capabilities and cannot automatically adjust resource allocation strategies according to environmental changes. Especially in complex and changeable environments, in distributed computing or cloud computing environments, static rules or simple optimization algorithms are difficult to cope with dynamically changing task requirements.

[0004] The existing technology still has scheduling bottleneck problems, especially in high-load or multi-tasking environments. Nodes in distributed computing environments often need to process a large number of tasks. When some nodes are overloaded, the scheduling system cannot perceive it in time and make effective scheduling adjustments, which can easily cause some nodes of the system to be in an overloaded state for a long time, further affecting the overall performance.

[0005] To sum up, the existing technology has the following main shortcomings: first, the resource allocation method is fixed and it is difficult to cope with dynamic load fluctuations; second, the scheduling strategy lacks adaptive adjustment capabilities and cannot be optimized according to real-time status; finally, the existing technology has low scheduling efficiency in complex environments, which is easy to form bottlenecks and cannot give full play to the system's resource utilization and performance. Summary of the invention

[0006] One purpose of the present invention is to propose a big data dynamic allocation and optimization scheduling method based on reinforcement learning. Through the application of reinforcement learning, the present invention realizes the automatic optimization of resource allocation and scheduling strategies, and significantly improves the resource utilization, scheduling efficiency and adaptability of the system.

[0007] A big data dynamic allocation and optimization scheduling method based on reinforcement learning according to an embodiment of the present invention includes the following steps:

[0008] S1. Obtain the load, data flow, storage space and network delay parameters of each computing node in the big data processing system, and build a real-time operation data set;

[0009] S2. Preprocess the real-time running data set, build a reinforcement learning environment model, use the resource status of each computing node as the state input, the data allocation and scheduling strategy as the action output, and the system performance index as the reward function;

[0010] S3, initializing the hierarchical reinforcement learning model, including high-level strategies and low-level strategies;

[0011] S4, Generate optimal data allocation and scheduling strategies based on hierarchical reinforcement learning models. High-level strategies determine sub-goals based on global states, and low-level strategies formulate specific actions based on sub-goals.

[0012] S5, real-time adjustment of data allocation and task scheduling of each computing node to balance system load and maximize resource utilization;

[0013] S6. During big data processing, the hierarchical reinforcement learning model continuously monitors the system status and dynamically updates the strategy, making adaptive adjustments to resource allocation and scheduling in a timely manner when the system load or environment changes;

[0014] S7. Through multiple rounds of iterative optimization, the hierarchical reinforcement learning model gradually converges to the optimal strategy.

[0015] Optionally, the S1 includes the following steps:

[0016] S11, real-time collection of load parameters of each computing node of the big data processing system, the load parameters including CPU utilization U cpu (t) and memory usage U mem (t), where t is the current time point, U cpu (t) represents the CPU load ratio of each computing node at time point t, U mem (t) represents the memory usage ratio of each computing node at time point t;

[0017] S12. Obtain the data flow parameter F of each computing node data(t), the data flow parameter represents the data transmission rate per unit time of each node at time point t, measured in bytes per second, and is used to reflect the transmission load during data processing;

[0018] S13. Record the storage space parameter S of each computing node space (t), the storage space parameter represents the available storage capacity of each node at time point t, in GB, which is used to monitor the storage usage status of each node;

[0019] S14. Collect the network delay parameters L of each computing node net (t), the network delay parameter represents the network delay from the data source to the target node at time t, in milliseconds, which is used to evaluate the network communication performance;

[0020] S15, the CPU utilization U cpu (t), memory usage U mem (t), data flow parameter F data (t), storage space parameter S space (t) and network delay parameter L net (t) to construct a real-time running data set D(t):

[0021] D(t)={U cpu (t), U mem (t), F data (t), S space (t), L net (t)}.

[0022] Optionally, S2 includes the following steps:

[0023] S21, performing data standardization processing on the real-time operation data set D / t), normalizing the parameters of each dimension, and obtaining a standardized real-time operation data set D′ / t);

[0024] S22, input the standardized real-time operation data set D′(t) into the reinforcement learning environment model to construct the system state space S(t);

[0025] S23. Define the action space A / t). The action space represents the decision of the system to dynamically allocate and schedule. The action space includes task allocation and scheduling strategies:

[0026]

[0027] Among them, π(A|S) is the policy function, which represents the probability of taking action A in state S, R(t) is the reward function, V(S t+1 ) is the value function at the next moment, γ is the discount factor, ai Assign and schedule specific tasks to the i-th node;

[0028] S24. Construct a reward function R / t) for evaluating the allocation and scheduling strategies of the system. The reward function is designed based on system performance indicators, including resource utilization, load balancing, and system latency. The expression of the reward function is:

[0029] R(t)=α 1 ·U total (t)+α 2 ·B load (t)-α 3 ·L total (t);

[0030] Among them, U total (t) is the overall resource utilization of the system, B load (t) is the balance degree of system load, L total (t) is the total network delay of the system, α 1 , α 2 and α 3 To adjust the weight coefficient of each indicator;

[0031] S25. The system is evaluated using the reward function R(t), and the proximal strategy optimization algorithm is used for iterative optimization. Through the collaboration of high-level agents and low-level agents, the global task allocation and node scheduling are gradually optimized:

[0032]

[0033] Among them, r t (θ) is the strategy ratio, which means the ratio of the new strategy to the old strategy. is the advantage function, ∈ is the coefficient controlling the update amplitude

[0034] S26. After each task allocation and scheduling, the reinforcement learning environment model updates the strategy based on the new state S / t+1) and reward R(t) combined with the historical state. The historical feedback is:

[0035]

[0036] Among them, θ t is the current strategy parameter, η 1 is the learning rate, λ i is the attenuation factor of historical feedback, and k is the maximum time window of historical state.

[0037] Optionally, S3 includes the following steps:

[0038] S31. Construct a hierarchical reinforcement learning model. The hierarchical reinforcement learning model has high-level strategies and low-level strategies. The high-level strategy is responsible for strategic decision-making on global task planning, task allocation, and resource scheduling. The low-level strategy is responsible for executing specific node-level task scheduling and optimization. The high-level strategy and the low-level strategy interact through a collaborative mechanism.

[0039] S32, define the global planning mechanism of high-level strategies, high-level strategies are based on the global state S t Evaluate the system's resource status and task load, and generate global task planning goals and resource allocation plans:

[0040]

[0041] Among them, π H (a H |S t ) is the high-level strategy, G t Planning goals for global tasks generated by high-level is the high-level strategy discount factor, R H (S t , G t ) is the high-level reward function;

[0042] S33, the global task planning target G generated by the low-level strategy in the high-level strategy t Based on the above, task scheduling and resource optimization are performed on specific nodes. The low-level strategy executes the scheduling order of tasks at the node level and adjusts the task execution plan in real time:

[0043]

[0044] Among them, π L (a L |S t , G t ) indicates that the lower-level strategy is in the global state S t and the global mission planning goal G t The specific execution action selected under the guidance of a L , is the low-level discount factor, R L (S t , a L ) is the local reward function of the low-level strategy;

[0045] S34. The high-level strategy and the low-level strategy jointly realize the connection between global planning and local execution through the collaborative optimization mechanism. The high-level strategy is responsible for deciding the strategic direction of the global task, and the low-level strategy executes specific tasks according to this direction. The collaborative optimization formula of the hierarchical reinforcement learning model is:

[0046]

[0047] Among them, η H and η L are the parameters of high-level strategy and low-level strategy respectively, θ H and θ L are the learning rates of high-level and low-level strategies, respectively. and Represent the proximal strategy optimization loss functions of high-level strategy and low-level strategy respectively;

[0048] S35. After each task is executed, the high-level strategy and the low-level strategy will be updated based on the system feedback. The high-level strategy updates the global strategy, and the low-level strategy updates the scheduling strategy within the node.

[0049] Optionally, S4 includes the following steps:

[0050] S41. In the process of big data processing, high-level strategies evaluate the global state S in real time. t Combine the current task load and system resource usage to generate the optimal global task allocation target G t , to cope with load fluctuations at different nodes;

[0051] S42, the optimal global task allocation target G generated by the low-level strategy based on the high-level strategy t , combined with the current local state S local (t) Generate specific task scheduling and resource allocation actions a L (t):

[0052]

[0053] Among them, a L (t) represents the task execution action generated by the low-level strategy, R L (S local (t), a L (t)) is the local reward function, is the discount factor for the low-level strategy;

[0054] S43, as the system status changes during task execution, the low-level strategy updates the task allocation and scheduling strategy in real time to maximize the resource utilization of each computing node and reduce scheduling delays;

[0055] S44, high-level strategy adjusts the optimal global task allocation target G based on system feedback t+1 , and re-evaluate the load balancing and resource optimization of global tasks;

[0056] S445. Through the collaborative work of high-level strategies and low-level strategies, dynamically adapt to different loads and state changes to generate optimal data allocation and scheduling strategies.

[0057] Optionally, the S7 includes the following steps:

[0058] S71. In the initial stage, the hierarchical reinforcement learning model is based on the current global state S t and the local state S local (t) Generate initial strategy π 0 ,The initial strategy is used to perform initial task allocation and scheduling on each node;

[0059] S72. Through multiple rounds of task execution, the system continuously collects the task execution results and system feedback of each node to generate an experience pool E. The data in the experience pool includes the global state S t , local state S local (t), task execution action a L (t) and the corresponding reward function value S local (t), a L (t);

[0060] S73, batch sampling the data in the experience pool and using the proximal strategy optimization algorithm to update the strategy;

[0061] S74. In each round of iterative optimization, the low-level strategy and the high-level strategy adjust their respective strategy parameters respectively, and gradually optimize the task scheduling and resource allocation strategies by maximizing the reward function value:

[0062]

[0063] in, is the policy parameter of the high-level policy at time t, η H is the learning rate of the high-level strategy, is the weight coefficient of the high-level strategy for different sub-goals, is the gradient of the high-level strategy to the parameters, R H (S t , G t ) is the global reward function of the high-level strategy, R L (S local,j (t), a L,j (t)) is the local reward function of the jth low-level strategy;

[0064] S75, the low-level strategy adjusts its own strategy by combining the feedback of the high-level strategy to optimize the resource allocation and task execution of each node:

[0065]

[0066] in, is the policy parameter of the jth low-level policy, η L is the learning rate of the low-level strategy, is the weight of the low-level strategy in different local task allocations, R H,h (S t , G t ) is the global reward feedback from the high-level strategy to the low-level strategy;

[0067] S76. When the convergence conditions of the hierarchical reinforcement learning model strategy meet the following convergence criteria, the iteration is stopped and the optimal strategy is finally generated:

[0068]

[0069] Among them, ∈ 1 is the preset convergence threshold, T is the total number of iterations, when the high-level strategy θ H With the low-level strategy θ L,j The change in 1 When , the model strategy is considered to have converged to the optimal strategy.

[0070] The beneficial effects of the present invention are:

[0071] (1) The present invention adopts a hierarchical reinforcement learning model. The high-level strategy is responsible for global task planning and resource allocation, and the low-level strategy is responsible for node-level task execution. Different from traditional static or preset rules, the present invention can make dynamic adjustments based on the real-time load of the system, network delay, and multi-dimensional state of storage space. By continuously collecting feedback data from each node, the reinforcement learning model can update the strategy in real time during task execution, making resource allocation more accurate, avoiding the problem of idle or overloaded resources at certain nodes, and realizing the optimization of global load balancing, which greatly improves the resource utilization and adaptability of the system.

[0072] (2) The present invention adopts a multi-round iterative optimization mechanism and utilizes the proximal strategy optimization algorithm in reinforcement learning to continuously optimize high-level strategies and low-level strategies. Through the continuous updating of the experience pool and the iterative optimization of strategies, the system can gradually converge to the optimal task allocation and scheduling strategy. Compared with the simple load balancing or task scheduling algorithm in the prior art, the present invention can not only perform real-time optimization according to the current system status, but also predict future load change trends through iterative learning and adjust resource allocation in advance, thereby reducing the occurrence of scheduling bottlenecks and improving the system response speed and overall performance.

[0073] (3) The hierarchical reinforcement learning model of the present invention realizes a close combination of global optimization and local execution through the collaborative work of high-level strategies and low-level strategies. The high-level agent generates a global task allocation target according to the global state. Under the guidance of the high-level strategy, the low-level agent generates a specific task scheduling action according to the local state. Through multi-level strategy design, it can flexibly respond to complex and changeable computing environments, and is particularly suitable for complex scenarios such as distributed computing or cloud computing. It can not only improve the accuracy of task scheduling, but also dynamically optimize resource utilization, ensuring that the system still runs efficiently under different load conditions, greatly improving the stability and robustness of the scheduling system. BRIEF DESCRIPTION OF THE DRAWINGS

[0074] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:

[0075] Figure 1 This is a flow chart of a big data dynamic allocation and optimization scheduling method based on reinforcement learning proposed by the present invention;

[0076] Figure 2 This is the division of labor and cooperation relationship between high-level intelligent agents and low-level intelligent agents in a big data dynamic allocation and optimization scheduling method based on reinforcement learning proposed by the present invention. DETAILED DESCRIPTION

[0077] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, which only illustrate the basic structure of the present invention in a schematic manner, and therefore only show the components related to the present invention.

[0078] refer to Figure 1-2 , a big data dynamic allocation and optimization scheduling method based on reinforcement learning, comprising the following steps:

[0079] S1. Obtain the load, data flow, storage space and network delay parameters of each computing node in the big data processing system, and build a real-time operation data set;

[0080] S2. Preprocess the real-time running data set, build a reinforcement learning environment model, use the resource status of each computing node as the state input, the data allocation and scheduling strategy as the action output, and the system performance index as the reward function;

[0081] S3, initializing the hierarchical reinforcement learning model, including high-level strategies and low-level strategies;

[0082] S4, Generate optimal data allocation and scheduling strategies based on hierarchical reinforcement learning models. High-level strategies determine sub-goals based on global states, and low-level strategies formulate specific actions based on sub-goals.

[0083] S5, real-time adjustment of data allocation and task scheduling of each computing node to balance system load and maximize resource utilization;

[0084] S6. During big data processing, the hierarchical reinforcement learning model continuously monitors the system status and dynamically updates the strategy, making adaptive adjustments to resource allocation and scheduling in a timely manner when the system load or environment changes;

[0085] S7. Through multiple rounds of iterative optimization, the hierarchical reinforcement learning model gradually converges to the optimal strategy.

[0086] In this implementation, S1 includes the following steps:

[0087] S11, real-time collection of load parameters of each computing node of the big data processing system, the load parameters include CPU utilization U cpu (t) and memory usage U mem (t), where t is the current time point, U cpu (t) represents the CPU load ratio of each computing node at time point t, U mem (t) represents the memory usage ratio of each computing node at time point t;

[0088] S12. Obtain the data flow parameter F of each computing node data (t), the data flow parameter represents the data transmission rate per unit time of each node at time point t, measured in bytes per second, and is used to reflect the transmission load during data processing;

[0089] S13. Record the storage space parameter S of each computing node space (t), the storage space parameter represents the available storage capacity of each node at time point t, in GB, which is used to monitor the storage usage status of each node;

[0090] S14. Collect the network delay parameters L of each computing node net (t), the network delay parameter represents the network delay from the data source to the target node at time t, in milliseconds, which is used to evaluate the network communication performance;

[0091] S15. Set the CPU utilization U cpu (t), memory usage U mem (t), data flow parameter F data (t), storage space parameter S space (t) and network delay parameter L net (t) to construct a real-time running data set D(t):

[0092] D(t)={U cpu(t), U mem (t), F data (t), S space (t), L net (t)}.

[0093] In this implementation, S2 includes the following steps:

[0094] S21, performing data standardization processing on the real-time operation data set D / t), normalizing the parameters of each dimension, and obtaining a standardized real-time operation data set D′ / t);

[0095] S22, input the standardized real-time operation data set D′(t) into the reinforcement learning environment model to construct the system state space S / t);

[0096] S23. Define the action space A(t). The action space represents the decision of the system for dynamic allocation and scheduling. The action space includes task allocation and scheduling strategies:

[0097]

[0098] Among them, π(A|S) is the policy function, which represents the probability of taking action A in state S, R(t) is the reward function, V(S t+1 ) is the value function at the next moment, γ is the discount factor, a i Assign and schedule specific tasks to the i-th node;

[0099] S24. Construct a reward function R / t) for evaluating the allocation and scheduling strategies of the system. The reward function is designed based on system performance indicators, including resource utilization, load balancing, and system latency. The expression of the reward function is:

[0100] R(t)=α 1 ·U total (t)+α 2 ·B load (t)-α 3 ·L total (t);

[0101] Among them, U total (t) is the overall resource utilization of the system, B load (t) is the balance degree of system load, L total (t) is the total network delay of the system, α 1 , α 2 and α 3 To adjust the weight coefficient of each indicator;

[0102] S25. Use the reward function R(t) to evaluate the system, and use the proximal strategy optimization algorithm for iterative optimization. Through the collaboration of high-level agents and low-level agents, the global task allocation and node scheduling are gradually optimized:

[0103]

[0104] Among them, r t (θ) is the strategy ratio, which means the ratio of the new strategy to the old strategy. is the advantage function, ∈ is the coefficient controlling the update amplitude

[0105] S26. After each task allocation and scheduling, the reinforcement learning environment model updates the strategy based on the new state S / t+1) and reward R(t) combined with the historical state. The historical feedback is:

[0106]

[0107] Among them, θ t is the current strategy parameter, η 1 is the learning rate, λ i is the attenuation factor of historical feedback, and k is the maximum time window of historical state.

[0108] In this implementation, S3 includes the following steps:

[0109] S31. Construct a hierarchical reinforcement learning model. The hierarchical reinforcement learning model has high-level strategies and low-level strategies. The high-level strategy is responsible for strategic decision-making on global task planning, task allocation, and resource scheduling. The low-level strategy is responsible for executing specific node-level task scheduling and optimization. The high-level strategy and the low-level strategy interact through a collaborative mechanism.

[0110] S32, define the global planning mechanism of high-level strategies, high-level strategies are based on the global state S t Evaluate the system's resource status and task load, and generate global task planning goals and resource allocation plans:

[0111]

[0112] Among them, π H (a H |S t ) is the high-level strategy, G t Planning goals for global tasks generated by high-level is the high-level strategy discount factor, R H (S t , G t ) is the high-level reward function;

[0113] S33, the global task planning target G generated by the low-level strategy in the high-level strategyt Based on the above, task scheduling and resource optimization are performed on specific nodes. The low-level strategy executes the scheduling order of tasks at the node level and adjusts the task execution plan in real time:

[0114]

[0115] Among them, π L (a L |S t , G t ) indicates that the lower-level strategy is in the global state S t and the global mission planning target G t The specific execution action selected under the guidance of a L , is the low-level discount factor, R L (S t , a L ) is the local reward function of the low-level strategy;

[0116] S34. The high-level strategy and the low-level strategy jointly realize the connection between global planning and local execution through the collaborative optimization mechanism. The high-level strategy is responsible for deciding the strategic direction of the global task, and the low-level strategy executes specific tasks according to this direction. The collaborative optimization formula of the hierarchical reinforcement learning model is:

[0117]

[0118] Among them, η H and η L are the parameters of high-level strategy and low-level strategy respectively, θ H and θ L are the learning rates of high-level and low-level strategies, respectively. and Represent the proximal strategy optimization loss functions of high-level strategy and low-level strategy respectively;

[0119] S35. After each task is executed, the high-level strategy and the low-level strategy will be updated based on the system feedback. The high-level strategy updates the global strategy, and the low-level strategy updates the scheduling strategy within the node.

[0120] In this implementation, S4 includes the following steps:

[0121] S41. In the process of big data processing, high-level strategies evaluate the global state S in real time. t Combine the current task load and system resource usage to generate the optimal global task allocation target G t , to cope with load fluctuations at different nodes;

[0122] S42, the optimal global task allocation target G generated by the low-level strategy based on the high-level strategy t, combined with the current local state S local (t) Generate specific task scheduling and resource allocation actions a L (t):

[0123]

[0124] Among them, a L (t) represents the task execution action generated by the low-level strategy, R L (S local (t), a L (t)) is the local reward function, is the discount factor for the low-level strategy;

[0125] S43, as the system status changes during task execution, the low-level strategy updates the task allocation and scheduling strategy in real time to maximize the resource utilization of each computing node and reduce scheduling delays;

[0126] S44, high-level strategy adjusts the optimal global task allocation target G based on system feedback t+1 , and re-evaluate the load balancing and resource optimization of global tasks;

[0127] S445. Through the collaborative work of high-level strategies and low-level strategies, dynamically adapt to different loads and state changes to generate optimal data allocation and scheduling strategies.

[0128] In this implementation, S7 includes the following steps:

[0129] S71. In the initial stage, the hierarchical reinforcement learning model is based on the current global state S t and the local state S local (t) Generate initial strategy π 0 ,The initial strategy is used to perform initial task allocation and scheduling on each node;

[0130] S72. Through multiple rounds of task execution, the system continuously collects the task execution results and system feedback of each node to generate an experience pool E. The data in the experience pool includes the global state S t , local state S local (t), task execution action a L (t) and the corresponding reward function value S local (t), a L (t);

[0131] S73, batch sampling the data in the experience pool and using the proximal strategy optimization algorithm to update the strategy;

[0132] S74. In each round of iterative optimization, the low-level strategy and the high-level strategy adjust their respective strategy parameters respectively, and gradually optimize the task scheduling and resource allocation strategies by maximizing the reward function value:

[0133]

[0134] in, is the policy parameter of the high-level policy at time t, η H is the learning rate of the high-level strategy, is the weight coefficient of the high-level strategy for different sub-goals, is the gradient of the high-level strategy to the parameters, R H (S t , G t ) is the global reward function of the high-level strategy, R L (S local,j (t), a L,j (t)) is the local reward function of the jth low-level strategy;

[0135] S75, the low-level strategy adjusts its own strategy by combining the feedback of the high-level strategy to optimize the resource allocation and task execution of each node:

[0136]

[0137] in, is the policy parameter of the jth low-level policy, η L is the learning rate of the low-level strategy, is the weight of the low-level strategy in different local task allocations, R H,h (S t , G t ) is the global reward feedback from the high-level strategy to the low-level strategy;

[0138] S76. When the convergence conditions of the hierarchical reinforcement learning model strategy meet the following convergence criteria, the iteration is stopped and the optimal strategy is finally generated:

[0139]

[0140] Among them, ∈ 1 is the preset convergence threshold, T is the total number of iterations, when the high-level strategy θ H With the low-level strategy θ L,j The change in 1 When , the model strategy is considered to have converged to the optimal strategy.

[0141] Embodiment 1:

[0142] In this embodiment 1, we use a cloud computing data center as the experimental scenario. The data center is responsible for processing user data of multiple large e-commerce platforms, including real-time user behavior analysis, order processing, and online updates of recommendation systems. During the daily peak hours (7pm to 9pm), user access will surge rapidly, causing the system load pressure to increase significantly. The 100 computing nodes of the data center are distributed in different areas, processing different types of tasks respectively. The system needs to dynamically allocate computing resources in a short time to ensure that all tasks can be completed efficiently without causing overload of certain nodes or waste of resources.

[0143] At 7:05 p.m. one day, the system administrator discovered through the monitoring platform that the CPU utilization of node "Node_A_23" located in area A suddenly surged from 60% to 95%. At the same time, its network delay also increased from 30ms to 90ms, showing obvious overload characteristics. The traditional static scheduling mechanism could not respond immediately in this case, resulting in a backlog of tasks at the node, a significant decrease in the speed of user order processing, and some users began to experience operation delays.

[0144] To solve this problem, the data center enabled the reinforcement learning-based big data dynamic allocation and optimization scheduling method of the present invention. First, the system automatically collected the CPU utilization, memory usage, network delay and storage space parameters of each computing node, and analyzed and made decisions through the reinforcement learning model. The high-level intelligent agent found that the load pressure of the nodes in area A was too large based on the global state evaluation. The system immediately generated a task allocation adjustment strategy and migrated some tasks of "Node_A_23" to the nodes with lower loads "Node_B_11" and "Node_C_14".

[0145] At this point, the system enters the automatic scheduling mode. The lower-level agents start to schedule resources for "Node_B_11" and "Node_C_14". The model allocates 50% of the migration tasks based on the real-time status of "Node_B_11" and the remaining tasks to "Node_C_14". The scheduling process is completed within 30 seconds. During the entire migration process, there is no significant delay in the user's real-time order processing, and the system processing efficiency is effectively guaranteed.

[0146] Then, at 7:30, another node "Node_D_05" had an abnormality, with CPU utilization exceeding 85% for a long time, network delay rising to 100ms, and memory utilization reaching 90%. Unlike traditional methods, the reinforcement learning model of the present invention can predict the resource usage trend of "Node_D_05" in advance. Before task scheduling, the system generates a forecast report indicating that the node's resources may reach a critical point within the next 5 minutes. Based on the forecast, the system migrates part of the data processing tasks to the idle node "Node_E_03" in advance, avoiding the occurrence of node overload problems.

[0147] At 8:10 p.m., a user's order processing task caused a sharp increase in data traffic, resulting in a network delay of 150ms at the node originating from "192.168.3.45". The order processing time increased from the normal 3 seconds to 15 seconds. The system generated a warning in real time, showing abnormal fluctuations in network delay and processing time. The reinforcement learning model automatically split the task and reallocated it to the nodes "Node_F_12" and "Node_G_18" with lower loads. Through collaboration between nodes, the order processing time was quickly restored to normal levels.

[0148] The following is a comparison of the specific data between the method of the present invention and the traditional method:

[0149] Task processing efficiency: When using the traditional static scheduling method, the average processing time for each batch of orders during peak hours is 12 seconds, while after using the reinforcement learning dynamic allocation method of the present invention, the order processing time is shortened to 8 seconds, and the efficiency is improved by 33%. When an abnormality occurs at the "192.168.3.45" node, it only takes 30 seconds to return to normal through real-time task scheduling adjustment.

[0150] Node overload situation: Under the traditional method, about 15 nodes crash due to overload every hour, and the system needs at least 3 minutes to restart or migrate tasks. In the system of the present invention, only 2 overloads occur per hour, and the system can complete resource adjustment within 10 seconds without manual intervention.

[0151] Network delay: Under the traditional scheduling mode, the average network delay of the system is 60ms, and the delay of some nodes exceeds 120ms, which seriously affects the user experience. After using the method of the present invention, the average network delay is reduced to 35ms, the maximum delay does not exceed 80ms, and the task processing is more stable.

[0152] Resource utilization: In the traditional method, the CPU utilization is about 60%, and a large number of node resources are idle. After using the present invention, the overall CPU utilization of the system is increased to 85%, and the load between nodes is more balanced.

[0153] User experience: When using the traditional scheduling method, the user's operation response time is an average of 10 seconds during peak hours. After using the method of the present invention, the response time is shortened to 6 seconds, and the user experience is significantly improved. In particular, in the recommendation system, the system can update the recommendation results within 1 second, and user feedback satisfaction has increased by 20%.

[0154] Through the detailed scenario description and data comparison of this embodiment, we have demonstrated the significant advantages of the present invention in the dynamic allocation and scheduling of big data. The reinforcement learning model can quickly respond to changes in system load through real-time monitoring and prediction, predict potential resource bottlenecks in advance, and ensure the optimal allocation of computing resources. Even during peak hours, the system can still maintain efficient operation, reduce task delays, and improve the overall stability of the system and user experience.

[0155] The present invention adopts a hierarchical reinforcement learning model. The high-level strategy is responsible for global task planning and resource allocation, and the low-level strategy is responsible for node-level task execution. Different from traditional static or preset rules, the present invention can be dynamically adjusted according to the real-time load of the system, network delay and multi-dimensional state of storage space. By continuously collecting feedback data from each node, the reinforcement learning model can update the strategy in real time during task execution, making resource allocation more accurate, avoiding the problem of idle or overloaded resources of some nodes, realizing the optimization of global load balancing, and greatly improving the resource utilization and adaptability of the system.

[0156] The present invention adopts a multi-round iterative optimization mechanism and utilizes the proximal strategy optimization algorithm in reinforcement learning to continuously optimize high-level strategies and low-level strategies. Through the continuous updating of the experience pool and the iterative optimization of strategies, the system can gradually converge to the optimal task allocation and scheduling strategy. Compared with the simple load balancing or task scheduling algorithm in the prior art, the present invention can not only perform real-time optimization according to the current system status, but also predict future load change trends through iterative learning and adjust resource allocation in advance, thereby reducing the occurrence of scheduling bottlenecks and improving the system response speed and overall performance.

[0157] The hierarchical reinforcement learning model of the present invention realizes the close combination of global optimization and local execution through the collaborative work of high-level strategies and low-level strategies. The high-level intelligent agent generates the global task allocation target according to the global state, and the low-level intelligent agent generates the specific task scheduling action according to the local state under the guidance of the high-level strategy. Through the multi-level strategy design, it can flexibly cope with the complex and changeable computing environment, and is particularly suitable for complex scenarios such as distributed computing or cloud computing. It can not only improve the accuracy of task scheduling, but also dynamically optimize resource utilization, ensuring that the system still runs efficiently under different load conditions, which greatly improves the stability and robustness of the scheduling system.

[0158] The above description is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can make equivalent replacements or changes according to the technical scheme and inventive concept of the present invention within the technical scope disclosed by the present invention, which should be covered by the protection scope of the present invention.

Claims

1. A big data dynamic allocation and optimization scheduling method based on reinforcement learning, characterized in that: The steps include: S1. Obtain the load, data flow, storage space and network delay parameters of each computing node in the big data processing system, and build a real-time operation data set; S2. Preprocess the real-time running data set and build a hierarchical reinforcement learning model, taking the resource status of each computing node as the state input, the data allocation and scheduling strategy as the action output, and the system performance index as the reward function; S3, initializing the hierarchical reinforcement learning model, including high-level strategies and low-level strategies; S4, Generate optimal data allocation and scheduling strategies based on hierarchical reinforcement learning models. High-level strategies determine sub-goals based on global states, and low-level strategies formulate specific actions based on sub-goals. S5, real-time adjustment of data allocation and task scheduling of each computing node to balance system load and maximize resource utilization; S6. During big data processing, the hierarchical reinforcement learning model continuously monitors the system status and dynamically updates the strategy, making adaptive adjustments to resource allocation and scheduling in a timely manner when the system load or environment changes; S7. Through multiple rounds of iterative optimization, the hierarchical reinforcement learning model gradually converges to the optimal strategy; The S3 comprises the following steps: S31. Initialize the hierarchical reinforcement learning model, including high-level strategies and low-level strategies. The high-level strategy is responsible for strategic decision-making on global task planning, task allocation, and resource scheduling, while the low-level strategy is responsible for executing specific node-level task scheduling and optimization. The high-level strategy and the low-level strategy interact through a collaborative mechanism. S32, define the global planning mechanism of high-level strategies, high-level strategies are based on the global state S t Evaluate the system's resource status and task load, and generate global task planning goals and resource allocation plans: Among them, π H (a H |S t ) is the high-level strategy, G t Planning goals for global tasks generated by high-level is the high-level strategy discount factor, R H (S t ,G t ) is the high-level reward function; S33, the global task planning target G generated by the low-level strategy in the high-level strategy t Based on the above, task scheduling and resource optimization are performed on specific nodes. The low-level strategy executes the scheduling order of tasks at the node level and adjusts the task execution plan in real time: Among them, π L (a L |S t ,G t ) indicates that the lower-level strategy is in the global state S t and the global mission planning target G t The specific execution action selected under the guidance of a L , is the low-level discount factor, R L (S t ,a L ) is the local reward function of the low-level strategy; S34. The high-level strategy and the low-level strategy jointly realize the connection between global planning and local execution through the collaborative optimization mechanism. The high-level strategy is responsible for deciding the strategic direction of the global task, and the low-level strategy executes specific tasks according to this direction. The collaborative optimization formula of the hierarchical reinforcement learning model is: Among them, η H and η L are the parameters of high-level strategy and low-level strategy respectively, θ H and θ L are the learning rates of high-level and low-level strategies, respectively. and Represent the proximal strategy optimization loss functions of high-level strategy and low-level strategy respectively; S35. After each task is executed, the high-level strategy and the low-level strategy will be updated based on the system feedback. The high-level strategy updates the global strategy, and the low-level strategy updates the scheduling strategy within the node.

2. According to claim 1, a big data dynamic allocation and optimization scheduling method based on reinforcement learning is characterized in that: The S1 comprises the following steps: S11, real-time collection of load parameters of each computing node of the big data processing system, the load parameters including CPU utilization U cpu (t) and memory usage U mem (t), where t is the current time point, U cpu (t) represents the CPU load ratio of each computing node at time point t, U mem (t) represents the memory usage ratio of each computing node at time point t; S12. Obtain the data flow parameter F of each computing node data (t), the data flow parameter represents the data transmission rate per unit time of each node at time point t, measured in bytes per second, and is used to reflect the transmission load during data processing; S13. Record the storage space parameter S of each computing node space (t), the storage space parameter represents the available storage capacity of each node at time point t, in GB, which is used to monitor the storage usage status of each node; S14. Collect the network delay parameters L of each computing node net (t), the network delay parameter represents the network delay from the data source to the target node at time t, in milliseconds, which is used to evaluate the network communication performance; S15, the CPU utilization U cpu (t), memory usage U mem (t), data flow parameter F data (t), storage space parameter S space (t) and network delay parameter L net (t) to construct a real-time running data set D(t): D(t)={U cpu (t),U mem (t),F data (t),S space (t),L net (t)}。 3. The big data dynamic allocation and optimization scheduling method based on reinforcement learning according to claim 2 is characterized in that: The S2 comprises the following steps: S21, performing data standardization processing on the real-time operation data set D(t), normalizing the parameters of each dimension, and obtaining a standardized real-time operation data set D'(t); S22, input the standardized real-time operation data set D'(t) into the hierarchical reinforcement learning model to construct the system state space S(t); S23. Define the action space A(t). The action space represents the decision of the system for dynamic allocation and scheduling. The action space includes task allocation and scheduling strategies: Among them, π(A|S) is the policy function, which represents the probability of taking action A in state S, R(t) is the reward function, V(S t+1 ) is the value function at the next moment, γ is the discount factor, a i Assign and schedule specific tasks to the i-th node; S24. Construct a reward function R(t) for evaluating the allocation and scheduling strategies of the system. The reward function is designed based on system performance indicators, including resource utilization, load balancing, and system delay. The expression of the reward function is: R(t)=α1·U total (t)+α2·B load (t)-α3·L total (t); Among them, U total (t) is the overall resource utilization of the system, B load (t) is the balance degree of system load, L total (t) is the total network delay of the system, α1, α2 and α3 are the weight coefficients for adjusting each index; S25. The system is evaluated using the reward function R(t), and the proximal strategy optimization algorithm is used for iterative optimization. Through the collaboration of high-level agents and low-level agents, the global task allocation and node scheduling are gradually optimized: Among them, r t (θ) is the strategy ratio, which means the ratio of the new strategy to the old strategy. is the advantage function, ∈ is the coefficient controlling the update amplitude S26. After each task allocation and scheduling, the hierarchical reinforcement learning model updates the strategy based on the new state S(t+1) and reward R(t) combined with the historical state. The historical feedback is: Among them, θ t is the current policy parameter, η1 is the learning rate, λ i is the attenuation factor of historical feedback, and k is the maximum time window of historical state.

4. The big data dynamic allocation and optimization scheduling method based on reinforcement learning according to claim 1 is characterized in that: The S4 comprises the following steps: S41. In the process of big data processing, high-level strategies evaluate the global state S in real time. t Combine the current task load and system resource usage to generate the optimal global task allocation target G t , to cope with load fluctuations at different nodes; S42, the optimal global task allocation target G generated by the low-level strategy based on the high-level strategy t , combined with the current local state S local (t) Generate specific task scheduling and resource allocation actions a L (t): Among them, a L (t) represents the task execution action generated by the low-level strategy, R L (S local (t),a L (t)) is the local reward function, is the discount factor for the low-level strategy; S43, as the system status changes during task execution, the low-level strategy updates the task allocation and scheduling strategy in real time to maximize the resource utilization of each computing node and reduce scheduling delays; S44, high-level strategy adjusts the optimal global task allocation target G based on system feedback t+1 , and re-evaluate the load balancing and resource optimization of global tasks; S45. Through the collaborative work of high-level strategies and low-level strategies, dynamically adapt to different loads and state changes to generate the optimal data allocation and scheduling strategy.

5. The big data dynamic allocation and optimization scheduling method based on reinforcement learning according to claim 1 is characterized in that: The S7 comprises the following steps: S71. In the initial stage, the hierarchical reinforcement learning model is based on the current global state S t and the local state S local (t) Generate an initial strategy π0, which is used to perform initial task allocation and scheduling on each node; S72. Through multiple rounds of task execution, the system continuously collects the task execution results and system feedback of each node to generate an experience pool E. The data in the experience pool includes the global state S t , local state S local (t), task execution action a L (t) and the corresponding reward function value S local (t),a L (t); S73, batch sampling the data in the experience pool and using the proximal strategy optimization algorithm to update the strategy; S74. In each round of iterative optimization, the low-level strategy and the high-level strategy adjust their respective strategy parameters respectively, and gradually optimize the task scheduling and resource allocation strategies by maximizing the reward function value: in, is the policy parameter of the high-level policy at time t, η H is the learning rate of the high-level strategy, is the weight coefficient of the high-level strategy for different sub-goals, is the gradient of the high-level strategy to the parameters, R H (S t ,G t ) is the global reward function of the high-level strategy, R L (S local,j (t),a L,j (t)) is the local reward function of the jth low-level strategy; S75, the low-level strategy adjusts its own strategy by combining the feedback of the high-level strategy to optimize the resource allocation and task execution of each node: in, is the policy parameter of the jth low-level policy, η L is the learning rate of the low-level strategy, is the weight of the low-level strategy in different local task allocations, R H,h (S t ,G t ) is the global reward feedback from the high-level strategy to the low-level strategy; S76. When the convergence conditions of the hierarchical reinforcement learning model strategy meet the following convergence criteria, the iteration is stopped and the optimal strategy is finally generated: Among them, ∈1 is the preset convergence threshold, T1 is the total number of iterations, when the high-level strategy θ H With the low-level strategy θ L,j When the changes in are all less than ∈1, the model strategy is considered to have converged to the optimal strategy.

Citation Information

Patent Citations

  • Target-oriented multi-agent coordination method

    CN116468097A

  • Time-varying task scheduling method and system based on constraint near-end strategy optimization

    CN117851056A