Virtual machine scheduling method in distributed environment based on deep reinforcement learning

Through deep reinforcement learning, a hybrid action space and multi-agent framework is built, which solves the multi-objective optimization problem of virtual machine scheduling in a distributed cloud computing environment, realizes efficient resource utilization and dynamic environment adaptability, and improves scheduling accuracy and system performance.

CN120540776APending Publication Date: 2025-08-26INFORMATION & TELECOMM COMPANY SICHUAN ELECTRIC POWER
View PDF 0 Cites 10 Cited by

Patent Information

Application Number
CN202510617039.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-14
Publication Date
2025-08-26

AI Technical Summary

Technical Problem

In a distributed cloud computing environment, traditional virtual machine scheduling methods are difficult to achieve multi-objective optimization (energy consumption, load balancing, SLA guarantee), and it is difficult to make effective collaborative decisions between discrete node selection and continuous resource allocation in a dynamic environment, resulting in resource fragmentation and response delays.

Method used

Using a method based on deep reinforcement learning, a hybrid action space is built, topology and timing features are captured through graph neural networks and long-term memory networks, a hierarchical reward function and adaptive weight strategy are designed, and a multi-agent collaborative framework is combined for virtual machine scheduling.

Benefits of technology

It realizes efficient resource utilization in a distributed environment, improves scheduling flexibility and system adaptability, reduces energy consumption and ensures service quality and load balancing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0005401361530000021
    Figure BDA0005401361530000021
  • Figure BDA0005401361530000022
    Figure BDA0005401361530000022
  • Figure BDA0005401361530000033
    Figure BDA0005401361530000033
Patent Text Reader

Abstract

The invention discloses a virtual machine scheduling method in a distributed environment based on deep reinforcement learning, and belongs to the technical field of cloud computing resource scheduling. According to the method, the defects of a traditional method in multi-objective optimization and mixed action space collaborative decision-making are overcome by constructing a mixed action space joint decision-making mechanism. The method specifically comprises the following steps: establishing a mixed action space containing discrete node selection and continuous resource allocation, filtering invalid nodes by adopting a dynamic mask mechanism, and ensuring resource ratio constraint through projection gradient descent; designing a hierarchical reward function to realize multi-target dynamic balancing, and dynamically adjusting the priorities of energy consumption, load balancing and SLA guarantee based on an adaptive weight strategy; a multi-agent collaborative framework is provided, cross-node topological dependence is captured by using a graph attention network, and dynamic fusion of spatio-temporal characteristics is realized through cross attention in combination with LSTM coding time sequence load characteristics; a course learning strategy and a priority experience playback mechanism are introduced to improve training efficiency and strategy robustness.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of intelligent cloud computing resource scheduling, and specifically relates to a virtual machine scheduling method in a distributed environment based on deep reinforcement learning. Background Art

[0002] With the rapid development of cloud computing technology, distributed data centers have become the core infrastructure supporting modern digital services. Virtualization technology significantly improves hardware resource utilization and system flexibility by abstracting physical resources into dynamically allocated virtual units. However, in ultra-large-scale distributed environments, the scheduling process of virtual machines (VMs) faces the challenge of collaborative optimization of multiple objectives, including key indicators such as load balancing, energy efficiency, and service level agreement (SLA) guarantees. Traditional scheduling strategies often use static rules or heuristic algorithms, which are difficult to adapt to dynamically changing business loads and heterogeneous resource environments, resulting in frequent problems such as resource fragmentation, surge in energy consumption, and response delays.

[0003] The current mainstream virtual machine scheduling methods mainly focus on short-term local optimization, such as threshold-based load migration mechanisms or resource allocation strategies driven by greedy algorithms. Although such methods have low computational overhead in specific scenarios, their decision-making process lacks comprehensive consideration of the global system state and long-term benefits, and is prone to falling into local optimal solutions. In recent years, some studies have attempted to introduce supervised learning or traditional reinforcement learning (RL) techniques to train models with historical data to predict resource requirements. However, such methods still have bottlenecks such as slow policy convergence and low cross-node collaboration efficiency when dealing with high-dimensional state spaces, sparse reward signals, and complex coupling relationships between distributed nodes. In addition, existing research focuses on a single optimization objective (such as minimizing energy consumption), and research on multi-objective dynamic trade-off mechanisms is still insufficient, making it difficult to meet the needs of multi-dimensional constraint parallelism in actual production environments.

[0004] In particular, in a distributed, heterogeneous node environment, virtual machine scheduling must simultaneously handle a mixed decision space of discrete actions (such as node selection) and continuous actions (such as resource quota allocation). Due to limitations in policy network design, traditional deep reinforcement learning (DRL) frameworks often require a discretized approximation of the action space, resulting in reduced scheduling accuracy and flexibility. Furthermore, communication latency and partial observability between distributed nodes further complicate policy training, making it difficult for existing methods to effectively aggregate cross-node state information and achieve real-time response. Summary of the Invention

[0005] The technical problem solved by this invention is: how to achieve efficient scheduling of virtual machines through deep reinforcement learning in a distributed cloud computing environment, overcome the shortcomings of traditional methods in multi-objective optimization (energy consumption, load balancing, SLA guarantee), dynamic environment adaptability and hybrid action space (discrete node selection and continuous resource allocation) collaborative decision-making, and improve global resource utilization and scheduling flexibility.

[0006] In order to achieve the above-mentioned purpose, the present invention adopts the following technical means:

[0007] The present invention provides a virtual machine scheduling method in a distributed environment based on deep reinforcement learning, comprising the following steps:

[0008] S1. Modeling the distributed virtual machine scheduling environment: Define the state vectors of physical nodes and virtual machine requests, construct a hybrid action space containing discrete node selection actions and continuous resource allocation actions, and fit a multi-objective optimization function;

[0009] S2. Multi-agent collaborative framework design: This adopts a layered architecture. The global coordinator captures cross-node topological dependencies through a graph attention network. Local executors encode temporal load dynamics through LSTM and capture global dynamic changes based on an attention mechanism.

[0010] S3. Topological Feature Extraction and State Encoding: Graph neural networks are used to aggregate node topological features, combined with LSTM to encode temporal load dynamics. The topological and temporal features are fused through a cross-attention mechanism to generate a joint encoding state.

[0011] S4. Hybrid spatial decomposition strategy: In the first stage, a dynamic action mask is generated based on physical constraints to filter out invalid nodes. In the second stage, a normalized weight generator is used on selected nodes to output a continuous resource allocation, and projected gradient descent is used to ensure the total resource constraint.

[0012] S5. Hierarchical reward function design: Basic rewards optimize energy consumption and service quality, high-level rewards improve load balancing and fairness, and adaptive weighting strategies are combined to dynamically adjust the priorities of multiple objectives.

[0013] S6. Curriculum Learning and Experience Replay Optimization: Migrate from simple environments to complex environments through progressive curriculum learning, and use a prioritized experience replay mechanism based on TD error to improve training efficiency.

[0014] In the above scheme, step s1 includes the following steps:

[0015] S11: Define the physical node state vector and the virtual machine request feature vector, where the physical node state vector is:

[0016]

[0017] The virtual machine request feature vector is:

[0018]

[0019] And define the global topological state S G The connection relationship between nodes and the cross-region communication cost matrix are obtained by aggregating the graph structure;

[0020] The CPU util Indicates the node CPU utilization, RAM util Indicates the node memory usage, Disk io Indicates node disk usage, Net lat Indicates the node network bandwidth;

[0021] CPU req Indicates the proportion of CPU and RAM that the virtual machine intends to occupy req Indicates the proportion of memory that the virtual machine intends to occupy, Qos level Indicates the service quality requirement level, and Security indicates the security requirement;

[0022] S12: Construct a mixed action space, including discrete actions A d ∈{0,1} N and continuous action A c ∈[0, 1] k , through a dual-branch policy network to output discrete and continuous action distributions respectively, share the state encoding layer and filter invalid node options through a dynamic mask mechanism;

[0023] S13 fitting multi-objective optimization function F = α t E+β t L+γ t S, where:

[0024] Energy consumption u i is the weighted resource utilization of node i;

[0025] Load balancing L = Var({u1, u2, ..., u N}),

[0026] SLA default rate The default condition is that the response time is greater than T SLA Or resource allocation delay is greater than D SLA .

[0027] α t , β t , γ t Represents the dynamic weight coefficient;

[0028] Core energy consumption weight α tDynamic adjustment based on SLA default rate:

[0029]

[0030] in represents the average SLA default rate over the past W time windows, τ is the SLA default rate threshold, and k represents the sensitivity coefficient;

[0031] Load balancing weight β t and energy consumption weight γ t Dynamic allocation based on target achievement rate:

[0032] γ t =1-α t -β t

[0033] where η L is the load balancing target achievement rate, η E Energy consumption target achievement rate is β t is the weight.

[0034] In the above solution, in step s12,

[0035] Discrete Action A d ∈{0, 1} N The target node selection for virtual machine placement is handled by Gumbel-Softmax relaxation for discrete action sampling;

[0036] Continuous Action A c ∈[0, 1] k Represents the normalized ratio of CPU, memory, and disk resources, satisfying ∑A c =1.

[0037] In the above scheme, the multi-agent collaborative framework design in step S2 includes:

[0038] Step S21: Construct a hierarchical architecture including a global coordinator and local executors, where:

[0039] The global coordinator integrates the graph attention network to generate a global state embedding h based on the node topology graph G. g ,dynamically capture the dependency between cross-node bandwidth delay and communication cost;

[0040] The local executor uses a long short-term memory network to encode the load sequence of the local node in the past T time steps and outputs the time series features. And report to the global coordinator;

[0041] Step S22: Calculate the dynamic weight a between nodes based on the scaled dot product attention mechanism ij , aggregate global features Specific satisfaction:

[0042]

[0043] Where Q, K, V are the learnable parameter matrices in the graph attention network, represents the global state embedding of node i, represents the global state embedding of node j, and d represents the dimension normalization factor;

[0044] Step S23: local time series features With global features Splice into joint state s i , the input policy network drives the scheduling decision, the formula is:

[0045]

[0046] In the above solution, the step S3 of topological feature extraction and state encoding includes:

[0047] Step S31: Hierarchical aggregation of node topology features:

[0048] The rack-level, region-level, and data center-level topological dependencies are extracted sequentially through a three-layer graph convolutional network (GCN), where the output of the kth layer is:

[0049]

[0050] in, Represented as an adjacency matrix with self-loops, is a degree matrix. The first layer aggregates the features of adjacent nodes in the same rack, the second layer aggregates the features of nodes across racks in the same region, and the third layer aggregates the features of nodes across regions.

[0051] Perform mean pooling on the 3-layer GCN output to get the node topology embedding:

[0052]

[0053] Represents the feature vector of node i after the kth layer of graph convolutional network (GCN). Each layer corresponds to a different topological relationship:

[0054] The first-level GCN aggregates the features of adjacent nodes within a rack and only includes the connection relationships between nodes within the same rack (e.g., server nodes are in the same physical rack);

[0055] The second layer GCN aggregates the features of nodes across racks within a region, including connections between nodes in different racks within the same region (e.g., a region consisting of multiple racks within a data center);

[0056] The third layer, GCN, aggregates the features of cross-region nodes, including connections between cross-region nodes (e.g., wide area network connections between different data centers or edge nodes);

[0057] MEAN represents the element-by-element arithmetic average of the three feature vectors. Mean pooling can balance the contribution of each layer of GCN and avoid the dominance of single-level features in the embedding representation;

[0058] Step S32: Extracting timing load characteristics:

[0059] Use LSTM network to encode the node's historical T time step load sequence X t =[x t-T ,…,x t ], output time series hidden state Where W lstm Represents the LSTM network parameters, namely the input gate, forget gate, and output gate weights, The output encodes the hidden state of the historical load dynamics;

[0060] Step S33: Fusion of topological and temporal features:

[0061] Dynamically associate topology and load characteristics based on the cross-attention mechanism to meet the following requirements:

[0062]

[0063] Generate a joint encoding state for subsequent policy decision-making, where CrossAttn(·) represents the cross attention mechanism.

[0064] In the above solution, the hybrid space decomposition strategy in step S4 includes:

[0065] Step S41: After receiving the virtual machine request, the system scans the remaining resources of all physical nodes and generates a dynamic action mask in real time based on the physical constraints. The mask calculation process is as follows: for each node i, if the remaining resources are Less than the requested resource Mark m i =0, indicating that it is not selectable, otherwise m i =1 can be selected to generate a mask vector M = [m1, m2, ..., m N ]Filter invalid node options;

[0066] S41. Dynamic action mask generation: When a virtual machine request is received, scan the remaining resources of all physical nodes and generate a dynamic action mask based on physical constraints;

[0067] For each node i, if its remaining resources Less than the virtual machine's requested resources Then mark the mask m i = 0 to filter invalid node options, otherwise mark m i =1 allows the node to be selected;

[0068] S42. Continuous resource allocation decision: At the selected node, the policy network outputs the non-normalized resource weight ω = W c ·s i +b c , where W c represents the weight matrix, b c Indicates the generation of original weights. The controlled softmax function is normalized to generate a continuous resource allocation weight w′:

[0069]

[0070] The temperature coefficient The initial value is 1.0 and gradually decreases to 0.1 during training to balance exploration and utilization;

[0071] S43. Strengthening the total resource constraint: Use the projected gradient descent algorithm to project the normalized weights into the legal space, and solve the following convex optimization problem to ensure that the resource allocation meets ∑ w ′=1,and w′>0:

[0072] min=||vw′|| 2 st∑v=1,v≥0

[0073] where v=(v1,v2,...,v n ) is a vector representing the legal normalized weight after the projection operation in resource allocation, and st represents the constraint condition.

[0074] In the above solution, the step S5 of designing the hierarchical reward function includes the following steps:

[0075] S51, directly optimizes energy consumption and service quality rewards Qos through basic rewards. The energy consumption reward is:

[0076] r e =-log(P i ·t vm +∈)

[0077] The quality of service reward Qos is defined as:

[0078] r q =exp(-max(0,latency-SLA max ))

[0079] Among them, P i is the node dynamic power consumption, t vm is the virtual machine running time, ∈ is a smoothing constant;

[0080] Advanced reward calculation:

[0081] The fairness of resource allocation is measured by the Gini coefficient G:

[0082]

[0083] The load balancing reward is defined as the negative variance of resource utilization among nodes:

[0084] r b = -Var({u1, u2, ..., u N})

[0085] The total reward function is:

[0086] R total =αr e +βr e +γr b +ζr f

[0087] Among them, α, β, γ, ζ represent the weights of energy consumption reward, Qos reward, fairness reward, load balancing reward, and fairness reward r respectively. f =-G.

[0088] In the above solution, step S6 of course learning and experience replay optimization includes the following steps:

[0089] S61. Progressive course learning:

[0090] Level 1: Train the basic placement strategy in a single-area static load environment. When the average reward fluctuation for 10 consecutive training runs is less than 5%, this stage is terminated and the process enters Level 2.

[0091] Level 2: In a multi-region cyclical load environment, the underlying parameters of the graph neural network (GNN) and long short-term memory network (LSTM) are frozen, and only the high-level parameters of the policy network are fine-tuned to adapt to cyclical load fluctuations.

[0092] Level 3: Unfreeze all parameters in a full-scale dynamic environment and introduce real-time random requests and node failure simulation to enhance policy robustness;

[0093] S62. Priority Experience Replay:

[0094] Calculate each experience (s, a, R total , s′)1’s absolute value of the timing difference (TD) error:

[0095] δ=|R total +γmaxQ(s′,a′)-Q(s,a)|

[0096] Where γ is the discount factor, Q(s, a) represents the Q-value estimate of the current state-action pair, maxQ(s′, a′) represents the maximum Q-value of all possible actions in the next state s′, that is, the estimate of future rewards, s, a, s′ represent the current state, current action, and next time series state respectively;

[0097] The experience pool is divided into three levels according to the TD error: high error (first 20%), medium error (middle 50%), and low error (bottom 30%), and the sampling probability ratio is set to high: medium: low = 6:3:1;

[0098] Assign importance weight ω to each experience i :

[0099]

[0100] Where N is the total number of experience pools, p i is the sampling probability of different error levels, β is the annealing parameter (usually β∈[0,1]), the initial value is close to 0 (ignoring weight), and gradually increases to 1 (completely compensating for the deviation), max j ω j Indicates the maximum importance weight among all samples and is used for normalization.

[0101] Because the present invention adopts the above technical solution, it has the following beneficial effects:

[0102] 1. This paper constructs a joint decision-making mechanism in a hybrid action space. Through topology-aware dynamic mask filtering and projected gradient constraint technology, it achieves precise collaborative optimization of node selection and resource allocation while ensuring the legitimacy of physical resources.

[0103] 2. To address the challenges of multi-objective dynamic optimization in cloud computing environments, this paper designs an adaptive weight adjustment mechanism based on sliding window feedback. By monitoring system-level indicators in real time and dynamically balancing the priorities of multiple objectives such as energy consumption, service quality, and load balancing, this paper combines a hierarchical reward function to guide the policy network to gradually optimize the global benefit.

[0104] 3. At the state modeling level, this invention deeply integrates graph neural networks (GNNs) and long short-term memory networks (LSTMs), achieving the first joint encoding of topological spatial dependencies and temporal load dynamics. By extracting multi-level topological relationships and modeling periodic load patterns, combined with a cross-attention mechanism to dynamically focus on key spatiotemporal features, it provides high-resolution contextual awareness for global optimization decision-making in complex distributed environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0105] Figure 1 It is a flow chart of the present invention. DETAILED DESCRIPTION

[0106] The following is a detailed description of the embodiments of the present invention. Although the present invention will be described and illustrated in conjunction with certain specific embodiments, it should be noted that the present invention is not limited to these embodiments. On the contrary, modifications or equivalent substitutions of the present invention are intended to fall within the scope of the claims of the present invention.

[0107] In addition, in order to better illustrate the present invention, numerous specific details are given in the following detailed description. It will be understood by those skilled in the art that the present invention can also be implemented without these specific details.

[0108] This application proposes a distributed virtual machine scheduling method based on deep reinforcement learning. By constructing a multi-agent collaborative framework and a hybrid action space modeling mechanism, it realizes global scheduling decisions for multi-objective optimization in a dynamic environment. This method combines graph neural networks (GNNs) with attention mechanisms to effectively capture the topological dependencies between distributed nodes; at the same time, it designs a hierarchical reward function to ensure service quality and resource fairness while reducing energy consumption. Compared with other virtual machine placement algorithms, we not only solve the limitations of discretized approximation or single action modeling, but also significantly enhance the system's adaptability in dynamic environments. At the same time, combined with the cross-attention mechanism to dynamically focus on key spatiotemporal features, it significantly improves the accuracy of cross-regional scheduling and resource prediction, breaking through the decision-making bottleneck from a local perspective.

[0109] The present invention aims to provide a method for scheduling virtual machines in a distributed environment based on deep reinforcement learning, which can effectively reduce energy consumption and resource waste in cloud environments. This objective of the present invention is achieved through the following technical solutions:

[0110] S1, distributed virtual machine scheduling environment modeling: First, define the state vectors of physical nodes and virtual machine requests, then establish a hybrid action space (joint modeling of discrete placement and continuous resource allocation), and finally fit a multi-objective optimization function.

[0111] S2, multi-agent collaborative framework design: The hierarchical architecture adopts a global coordinator topology to perceive basic information, uses local executors to perceive timing loads, and captures global dynamic change states based on the attention mechanism.

[0112] S3, topological feature extraction and state encoding: Use graph neural network (GNN) to aggregate node topological features, LSTM to encode time series load dynamics, and jointly encode topological and time series features to capture the spatial dependencies of distributed systems.

[0113] S4, hybrid space decomposition strategy: In the first stage, discrete node selection actions are generated through a topology-aware network, dynamic action masks are generated in real time based on physical constraints, and invalid node options are filtered out; in the second stage, based on the selected nodes, a normalized weight generator is used to output continuous resource allocation parameters, and the total resource constraint is ensured through projected gradient descent.

[0114] S5, Hierarchical Reward Function Design: This approach directly optimizes energy consumption and QoS through basic rewards, while also introducing higher-level rewards to improve system-level metrics such as fairness and load balancing. Based on an adaptive weighting strategy, the system dynamically adjusts target priorities, achieving global optimization while ensuring core performance.

[0115] S6, curriculum learning and experience replay optimization: First, through progressive curriculum learning, the system gradually migrates from simple environments to complex environments, reducing the difficulty of exploration in the early stages of training and accelerating strategy convergence; second, a priority experience replay mechanism is adopted to dynamically screen key samples based on TD error, giving priority to learning high-value experience, thereby improving training efficiency.

[0116] Specifically, the process of modeling the distributed virtual machine scheduling environment of the function point in step S1 includes the following sub-steps:

[0117] S11, state vector definition.

[0118] Construct a multi-dimensional vector to describe the real-time status of the node, including CPU utilization, memory usage, disk usage, and network bandwidth, and define the physical node state vector:

[0119]

[0120] Define the virtual machine request feature vector:

[0121]

[0122] Includes resource requirements such as CPU, memory, quality of service level, and security requirements.

[0123] Global topology state S G :The connection relationship between nodes (bandwidth delay, energy transmission loss) and cross-regional communication cost matrix are aggregated through graph structure.

[0124] S12, in the construction of the hybrid action space, two types of actions are defined: discrete actions and continuous actions, which together constitute the complete action space. Specifically, the discrete action A d ∈{0, 1} N Indicates the target node selection for virtual machine placement, where N is the total number of physical nodes, and Gumbel-Softmax relaxation is used to process discrete action sampling. Each binary value indicates whether the corresponding physical node is selected, for example, Ad =[0, 1, 0..., 0] means selecting the second physical node. Resource allocation decision A c ∈[0, 1] k It is used to adjust the resource allocation ratio, corresponding to the ratio of k resources such as CPU, memory (RAM), disk (Disk), etc., which needs to meet ∑A c =1.

[0125] By constructing a dual-branch policy network, the discrete and continuous branches share the state encoding layer but output action distributions separately, and joint decision-making under physical constraints is achieved through a dynamic mask mechanism.

[0126] S13, fitting the multi-objective function F = α t E+β t L+γ t S.

[0127] in Represents energy consumption, P i is the dynamic power consumption of node i, t vm Indicates the total running time of the virtual machine from startup to completion (dynamic value), P i =P idle +(P max -P idle )·u i , P idle is the power consumption of the node when it is idle (fixed value), P max is the power consumption when the node is fully loaded (fixed value), u i is the current resource utilization of the node (CPU + memory weighted average), the total running time of the virtual machine from startup to completion (dynamic value). L = Var({u1, u2, ..., u N}) represents the load imbalance, which reflects the cluster load balancing status by calculating the variance of resource utilization of all nodes. It also integrates multiple resource types (cpu, memory, disk, bandwidth), where u i The weighted variance is calculated:

[0128] u i =w cpu ·u cpu +w ram ·u ram +w mem ·u mem +w bw ·u bw

[0129] w cpu +w ram +w mem +w bw =1.

[0130] where w cpu 、w ram 、w mem 、w bw Respectively represent the weights of cpu, memory, disk, and bandwidth:

[0131] u cpu 、u ram 、u mem 、u bw Respectively represent CPU utilization, memory utilization, disk utilization, and bandwidth utilization

[0132] Indicates the SLA breach rate. When the response time of a request is greater than T SLA Or resource allocation delay is greater than D SLA When breach of contract.

[0133] Based on the adaptive weight strategy, the system can dynamically adjust the target priority, where the core energy consumption weight adjustment formula is:

[0134]

[0135] represents the average SLA default rate over the past W time windows. τ is the SLA default rate threshold, which can be set based on the specific business and is set to 0.05 by default. k represents the sensitivity coefficient, which controls the rate of weight change. The larger k is, the more sensitive the weight is to changes in the default rate. To ensure that the total weight is constant, the other weight adjustment formulas are as follows:

[0136] γ t =1-α t -β t

[0137] η L is the load balancing target achievement rate, η E Energy consumption target achievement rate. The historical target achievement rate is counted through a sliding window. Every T time steps, the following indicators are counted: Energy consumption achievement rate If η E >1 indicates that energy consumption does not exceed expectations; load balancing achievement rate If η L >1, indicating that the load balancing is better than the target; SLA achievement rate When dynamically adjusting priorities, SLA takes precedence. When α t Exponentially increasing, forcing the strategy to prioritize reducing the default rate, if η E <1, that is, energy consumption exceeds the standard, then increase β t Weight, strengthen load balancing to disperse hot spots, if η E <1, that is, the balance is degraded, then increase γt Weight, giving priority to ensuring service quality, where E target Indicates the expected target energy consumption, E actual Indicates actual energy consumption, L target Indicates the expected target load balancing value, L actual Indicates the actual load balancing value.

[0138] Specifically, the process of designing a multi-agent collaborative framework for function point description in step S2 includes the following sub-steps:

[0139] S21: adopts a layered architecture, including a global coordinator and local executors. The global coordinator is deployed in the central node, integrating the graph attention network (GAT) to capture cross-node topological dependencies, realize dynamic information transmission between nodes, and generate the global state embedding h g =GAT(G;W g ), where G is the node topology; local executors are distributed in each physical node, and LSTM is used to encode the local time series load sequence Dynamically reported to the coordinator, where W g Represents the GAT parameters in the global coordinator, represents the input timing sequence of the local executor, W l Represents the parameters of the LSTM in the local executor.

[0140] S22: Based on the attention mechanism to capture the global dynamic change state, the cross-node attention implementation relies on the coordinator to calculate the attention weights between nodes Aggregate global features The local executor will and After splicing, input the policy network, the formula is:

[0141]

[0142] Realize state awareness in dynamic environments.

[0143] Where Q, K, V are the learnable parameter matrices in the graph attention network. represents the global state embedding of node i, represents the global state embedding of node j, and d represents the dimension normalization factor;

[0144] Specifically, the process of topological feature extraction and state encoding in step S3 includes the following sub-steps:

[0145] S31: Design a node feature aggregation mechanism. Use a graph neural network (GNN) to aggregate node features layer by layer, and use a graph convolutional layer (GCN) to iteratively update node features. The output of the kth layer is:

[0146]

[0147] Where W (k) represents the trainable weight matrix, σ(·) represents the nonlinear activation function, Represented as an adjacency matrix with self-loops, is the degree matrix. Each layer of GCN aggregates the features of one-hop neighbors, and multi-hop information transmission is achieved by stacking multiple layers. Stacking 3 layers of GCN captures the three-level topology relationship of rack-region-data center. The first layer aggregates the features of adjacent nodes in the same rack, the second layer aggregates the features of cross-rack nodes in the same region, and the third layer aggregates the features of cross-region nodes. The final node embedding is:

[0148]

[0149] S32: Extract time series features. Use the long short-term memory network (LSTM) to model and encode the historical state sequence, and input the load sequence X of the node in the past T time slices. t =[x t-T ,…,x t ], output hidden state Where W lstm Represents the LSTM network parameters, namely the input gate, forget gate, and output gate weights, The output encodes the hidden state of the historical load dynamics. The fusion of topological and temporal features is achieved through the cross-attention mechanism:

[0150]

[0151] Specifically, the process of mixing the spatial decomposition strategy in step S4 includes the following sub-steps:

[0152] S41: After receiving the virtual machine request, the system scans the remaining resources of all physical nodes and generates a dynamic action mask in real time based on the physical constraints. The mask calculation process is as follows: for each node i, if the remaining resources Less than the requested resource Mark m i =0, indicating that it is not selectable, otherwise m i =1 to select.

[0153] S42: Based on the selected nodes, the normalized weight generator is used to make quota decisions. The policy network outputs the non-normalized resource weight ω = W for the selected nodes. c ·s i +b c , where W c represents the weight matrix, b c Indicates the generation of original weights. Temperature coefficient initial As training progresses, it gradually decreases to 0.1, shifting from exploration to utilization. Finally, the total resource constraint is ensured by projected gradient descent. During gradient updates, the resource weights are forced to satisfy ∑ω′=1 and ω′>0. The weights ω′ after gradient updates are projected into the legal space, which is mathematically equivalent to solving a convex optimization problem:

[0154] min||v-ω′|| 2 st∑v=1,v≥0

[0155] where v=(v1, v2, ..., v n ) is a vector representing the legal normalized weight st after the projection operation in resource allocation, which represents the constraint condition.

[0156] Specifically, the hierarchical reward function design in step S5 includes the following sub-steps:

[0157] S51: Directly optimize energy consumption and service quality reward Qos through basic rewards, energy consumption reward is

[0158] r e =-log(P i ·t vm +∈)

[0159] Qos rewards are:

[0160] r q =exp(-max(0,latency-SLA max )),

[0161] Among them, P i is the node dynamic power consumption, t vm is the virtual machine running time, ∈ is a smoothing constant, latency is the actual response time of the request, SLA max The maximum latency threshold allowed by the service level agreement.

[0162] At the same time, we introduce high-order reward calculation to calculate the Gini coefficient G of resource allocation. The closer it is to 0, the fairer it is. If all nodes have the same resource utilization, then G = 0. At the end of each training cycle, the global Gini coefficient G is calculated to generate the fairness reward r f = -G; Load balancing reward is used to penalize the variance of resource utilization between nodes: r b = -Var({u1, u2, ..., u N}), where N represents the number of nodes, u i represents the resource utilization of node i, xj represents the resource utilization of node j, Indicates the mean utilization of all nodes.

[0163] The total reward function is: R total =αr e +βr e +γr b +ζr f

[0164] Where α, β, γ, and ζ represent the weights of energy consumption reward, QoS reward, fairness reward, and load balancing reward, respectively. Specifically, the course learning and experience replay optimization in step S6 includes the following sub-steps:

[0165] S61: Progressive Curriculum Learning. Level 1: Considering a single-area static load, the basic placement strategy is learned when training the target. This stage is terminated when the average reward fluctuation for 10 consecutive training runs is less than 5%.

[0166] Level 2, taking into account multi-regional cyclical loads, adopts a migration strategy that freezes the underlying parameters of GNN and LSTM and only fine-tunes the high-level policy network to adapt to cyclical load fluctuations. Level 3, on the other hand, considers the full-scale dynamic environment, unfreezes all parameters, and introduces real-time random requests and node failure simulation.

[0167] S62: Priority experience playback. In order to efficiently utilize samples, a priority experience playback strategy is adopted. For each experience (s, a, R total , s′), calculate the absolute value of TD error

[0168] δ=|R total +γmaxQ(s′,a′)-Q(s,a)|

[0169] Where s, a, and s′ represent the current state, current action, and next time series state, respectively.

[0170] Then, the experience pool is divided into three levels according to the TD error: high (first 20%), medium (middle 50%), and low (last 30%). The sampling probability is set to high: medium: low = 6:3:1. The high error samples are replayed first, and the importance weight is set to

[0171] Where N is the total number of experience pools, p i is the sampling probability of different error gears, β is the annealing parameter, the initial value is close to 0, gradually increases to 1, max j ω j Represents the maximum importance weight of all samples, used for normalization, ω j Represents the importance weight of the j-th experience in the experience pool.

[0172] The distributed virtual machine scheduling method based on deep reinforcement learning described in the present invention brings the following significant beneficial effects at the technical level:

[0173] 1. Collaborative Decision-Making Optimization in Hybrid Action Spaces

[0174] Through a topology-aware dynamic action masking mechanism and projected gradient descent constraint technology, the system innovatively achieves joint decision-making for discrete node selection and continuous resource allocation. Dynamic masking effectively filters out invalid node options due to insufficient physical resources, ensuring the legitimacy of the decision space. The projected gradient descent algorithm ensures that the continuous resource allocation strictly meets the total resource constraints, fundamentally avoiding the resource allocation inaccuracies caused by traditional discretization approximations. The synergy between the two enables the system to generate optimal scheduling solutions that conform to physical laws even under complex constraints.

[0175] 2. Improved adaptability in dynamic environments

[0176] By integrating the multi-level topological feature extraction of graph neural networks with the temporal dynamic encoding capabilities of long-short-term memory networks, a feature fusion mechanism with spatiotemporal-aware joint modeling capabilities was constructed. A three-level graph convolutional network captures rack-level, regional-level, and cross-regional topological dependencies. Combined with a cross-attention mechanism, it dynamically correlates node load fluctuations with network communication status, enabling scheduling strategies to perceive topological changes and load migration trends in distributed environments in real time, significantly enhancing responsiveness to dynamic scenarios such as traffic bursts and node failures.

[0177] 3. Intelligent trade-off mechanism for multi-objective conflicts

[0178] The designed hierarchical reward function and sliding window feedback mechanism overcome the limitations of fixed weight coefficients in traditional multi-objective optimization. Basic rewards directly drive the immediate optimization of energy consumption and service quality, while higher-order rewards guide fair allocation of system-wide resources through Gini coefficient and load variance calculations. Combined with an adaptive weight adjustment algorithm based on SLA default rate feedback, this algorithm dynamically adjusts multi-objective priorities based on real-time business status, achieving Pareto optimality in energy consumption and load balancing while ensuring core service quality, effectively resolving the problem of policy oscillation caused by conflicting objectives.

[0179] 4. Enhanced system-level collaborative scheduling capabilities

[0180] The layered architecture of a global coordinator and local executors utilizes a graph attention network to dynamically aggregate cross-node state information and coordinate distributed decision-making. The global coordinator accurately captures topological features such as inter-node bandwidth latency and communication costs, while the local executors deeply explore historical node load patterns. The two complement each other's features through a scaled dot-product attention mechanism. This layered coordination mechanism significantly improves the efficiency of resource state synchronization in ultra-large clusters and provides a highly accurate basis for cross-region virtual machine migration.

[0181] 5. Optimize strategy training efficiency and stability

[0182] An innovative progressive curriculum learning strategy significantly reduces the risk of policy divergence during the initial exploration phase of deep reinforcement learning by layering and progressively training based on environmental complexity. A prioritized experience replay mechanism dynamically selects high-value training samples based on temporal difference error, and combined with importance sampling weight adjustment, effectively accelerates the convergence of the policy network. This training optimization framework enables the system to rapidly adapt to policy migration from simple static environments to complex dynamic ones, significantly improving the reliability of the algorithm's deployment in real-world production environments.

Claims

1. A virtual machine scheduling method in a distributed environment based on deep reinforcement learning, characterized in that: The following steps are involved: S1. Modeling the distributed virtual machine scheduling environment: Define the state vectors of physical nodes and virtual machine requests, construct a hybrid action space containing discrete node selection actions and continuous resource allocation actions, and fit a multi-objective optimization function; S2. Multi-agent collaborative framework design: This adopts a layered architecture. The global coordinator captures cross-node topological dependencies through a graph attention network. Local executors encode temporal load dynamics through LSTM and capture global dynamic changes based on an attention mechanism. S3. Topological Feature Extraction and State Encoding: Graph neural networks are used to aggregate node topological features, combined with LSTM to encode temporal load dynamics. The topological and temporal features are fused through a cross-attention mechanism to generate a joint encoding state. S4. Hybrid spatial decomposition strategy: In the first stage, a dynamic action mask is generated based on physical constraints to filter out invalid nodes. In the second stage, a normalized weight generator is used on selected nodes to output a continuous resource allocation, and projected gradient descent is used to ensure the total resource constraint. S5. Hierarchical reward function design: Basic rewards optimize energy consumption and service quality, high-level rewards improve load balancing and fairness, and adaptive weighting strategies are combined to dynamically adjust the priorities of multiple objectives. S6. Curriculum Learning and Experience Replay Optimization: Migrate from simple environments to complex environments through progressive curriculum learning, and use a prioritized experience replay mechanism based on TD error to improve training efficiency.

2. The method according to claim 1, characterized in that Step s1 includes the following steps: S11: Define the physical node state vector and the virtual machine request feature vector, where the physical node state vector is: The virtual machine request feature vector is: And define the global topological state S G The connection relationship between nodes and the cross-region communication cost matrix are obtained by aggregating the graph structure; The CPU util Indicates the node CPU utilization, RAM util Indicates the node memory usage, Disk io Indicates node disk usage, Net lat Indicates the node network bandwidth; CPU req Indicates the proportion of CPU and RAM that the virtual machine intends to occupy req Indicates the proportion of memory that the virtual machine intends to occupy, Qos level Indicates the service quality requirement level, and Security indicates the security requirement; S12: Construct a mixed action space, including discrete actions A d ∈{0, 1} N and continuous action A c ∈[0, 1] k , through a dual-branch policy network to output discrete and continuous action distributions respectively, share the state encoding layer and filter invalid node options through a dynamic mask mechanism; S13 fitting multi-objective optimization function F = α t E+β t L+γ t S, where: Energy consumption u i is the weighted resource utilization of node i; Load balancing L = Var({u1, u2, ..., u N }), SLA default rate The default condition is that the response time is greater than T SLA Or resource allocation delay is greater than D SLA . α t , β t , γ t Represents the dynamic weight coefficient; Core energy consumption weight α t Dynamic adjustment based on SLA default rate: in represents the average SLA default rate over the past W time windows, τ is the SLA default rate threshold, and k represents the sensitivity coefficient; Load balancing weight β t and energy consumption weight γ t Dynamic allocation based on target achievement rate: where η L is the load balancing target achievement rate, η E Energy consumption target achievement rate is β t is the weight.

3. The method according to claim 1, characterized in that In step s12, Discrete Action A d ∈{0, 1} N The target node selection for virtual machine placement is handled by Gumbel-Softmax relaxation for discrete action sampling; Continuous Action A c ∈[0, 1] k Represents the normalized ratio of CPU, memory, and disk resources, satisfying ∑A c =1.

4. The method according to claim 1, wherein The multi-agent collaborative framework design in step S2 includes: Step S21: Construct a hierarchical architecture including a global coordinator and local executors, where: The global coordinator integrates the graph attention network to generate a global state embedding h based on the node topology graph G. g ,dynamically capture the dependency between cross-node bandwidth delay and communication cost; The local executor uses a long short-term memory network to encode the load sequence of the local node in the past T time steps and outputs the time series features. And report to the global coordinator; Step S22: Calculate the dynamic weight a between nodes based on the scaled dot product attention mechanism ij , aggregate global features Specific satisfaction: Where Q, K, V are the learnable parameter matrices in the graph attention network. represents the global state embedding of node i, represents the global state embedding of node j, and d represents the dimension normalization factor; Step S23: local time series features With global features Connected to joint state s i , the input policy network drives the scheduling decision, the formula is:

5. The method according to claim 1, characterized in that The step S3 of topological feature extraction and state encoding includes: Step S31: Hierarchical aggregation of node topology features: The three-layer graph convolutional network is used to extract the rack-level, regional-level, and data center-level topological dependencies in turn, where the output of the k-th layer is: in, Represented as an adjacency matrix with self-loops, is a degree matrix. The first layer aggregates the features of adjacent nodes in the same rack, the second layer aggregates the features of nodes across racks in the same region, and the third layer aggregates the features of nodes across regions. Perform mean pooling on the 3-layer GCN output to get the node topology embedding: Represents the feature vector of node i after the kth layer of graph convolutional network. Each layer corresponds to a different topological relationship: The first-level GCN aggregates the features of adjacent nodes within a rack and only includes the connection relationships between nodes within the same rack; The second layer GCN aggregates the features of nodes across racks within a region, including connections between nodes in different racks within the same region; The third layer GCN aggregates the features of cross-region nodes, including the connections between cross-region nodes; MEAN represents the element-by-element arithmetic average of the three feature vectors. Mean pooling can balance the contribution of each layer of GCN and avoid the dominance of single-level features in the embedding representation; Step S32: Extracting timing load characteristics: Use LSTM network to encode the node's historical T time step load sequence X t =[x t-T ,…,x t ], output time series hidden state Where W lstm Represents the LSTM network parameters, namely the input gate, forget gate, and output gate weights, The output encodes the hidden state of the historical load dynamics; Step S33: Fusion of topological and temporal features: Dynamically associate topology and load characteristics based on the cross-attention mechanism to meet the following requirements: Generate a joint encoding state for subsequent policy decision-making, where CrossAttn(·) represents the cross attention mechanism.

6. The method according to claim 1, characterized in that The hybrid space decomposition strategy in step S4 includes: Step S41: After receiving the virtual machine request, the system scans the remaining resources of all physical nodes and generates a dynamic action mask in real time based on the physical constraints. The mask calculation process is as follows: for each node i, if the remaining resources are Less than the requested resource Mark m i =0, indicating that it is not selectable, otherwise m i =1 can be selected to generate a mask vector M = [m1, m2, ..., m N ]Filter invalid node options; S41. Dynamic action mask generation: When a virtual machine request is received, scan the remaining resources of all physical nodes and generate a dynamic action mask based on physical constraints; For each node i, if its remaining resources Less than the virtual machine's requested resources Then mark the mask m i = 0 to filter invalid node options, otherwise mark m i =1 allows the node to be selected; S42. Continuous resource allocation decision: At the selected node, the policy network outputs non-normalized resource weights By temperature coefficient The controlled softmax function is normalized to generate a continuous resource allocation weight w′: The temperature coefficient The initial value is 1.0 and gradually decreases to 0.1 during training to balance exploration and utilization; S43. Strengthening the total resource constraint: Use the projected gradient descent algorithm to project the normalized weights into the legal space, and solve the following convex optimization problem to ensure that the resource allocation meets ∑ w ′=1,and w′>0: min=||v-w′|| 2 s.t.∑v=1,v≥0 where v=(v1, v2, ..., v n ) is a vector representing the legal normalized weight after the projection operation in resource allocation, and st represents the constraint condition.

7. The method according to claim 1, characterized in that The step S5 of designing the hierarchical reward function includes the following steps: S51, directly optimizes energy consumption and service quality rewards Qos through basic rewards. The energy consumption reward is: r e =-log(P i ·t vm +∈) The quality of service reward Qos is defined as: r q =exp(-max(0,latency-SLA max )) Among them, P i is the node dynamic power consumption, t vm is the virtual machine running time, ∈ is a smoothing constant; Advanced reward calculation: The fairness of resource allocation is measured by the Gini coefficient G: The load balancing reward is defined as the negative variance of resource utilization among nodes: r b =-Var({u1,u2,…,u N }) The total reward function is: R total =αr e +βr e +γr b +ζr f Among them, α, β, γ, ζ represent the weights of energy consumption reward, Qos reward, fairness reward, load balancing reward, and fairness reward r respectively. f =-G.

8. The method according to claim 1, characterized in that The step S6 course learning and experience replay optimization includes the following steps: S61. Progressive course learning: Level 1: Train the basic placement strategy in a single-area static load environment. When the average reward fluctuation for 10 consecutive training runs is less than 5%, this stage is terminated and the next step is Level 2. Level 2: In a multi-region cyclical load environment, the underlying parameters of the graph neural network and long short-term memory network are frozen, and only the high-level parameters of the policy network are fine-tuned to adapt to cyclical load fluctuations. Level 3: Unfreeze all parameters in a full-scale dynamic environment and introduce real-time random requests and node failure simulation to enhance policy robustness; S62. Priority Experience Replay: Calculate each experience (s, a, R total ,s′)1 the absolute value of the timing difference error: δ=|R total +γmaxQ(s′,a′)-Q(s,a)| Where γ is the discount factor, Q(s,a) represents the Q-value estimate of the current state-action pair, maxQ(s′,a′) represents the maximum Q-value of all possible actions in the next state s′, that is, the estimate of future rewards, s, a, s′ represent the current state, current action, and next time series state respectively; The experience pool is divided into three levels according to the TD error: high error, medium error, and low error, and the sampling probability ratio is set to high: medium: low = 6:3:1; Assign importance weight ω to each experience i : Where N is the total number of experience pools, p i is the sampling probability of different error gears, β is the annealing parameter, the initial value is close to 0, gradually increases to 1, max j ω j Indicates the maximum importance weight among all samples and is used for normalization.

Citation Information

Cited By

  • Mountain land drip irrigation pipe network pressure balance adjusting method based on GAT and HRL

    CN120725411A

  • Advertisement effect evaluation method and system based on artificial intelligence

    CN120746652A

  • An advertisement effect evaluation method and system based on artificial intelligence

    CN120746652B

  • Aircraft dispatching strategy generation method, device, equipment, medium and product

    CN120823730A

  • Resource management method and system for green cloud computing and storage medium

    CN120832248A