Deep reinforcement learning resource allocation method for green mobile edge computing
Through the deep reinforcement learning resource allocation method, the problem of increased energy consumption in modern communication systems is solved, and the effect of improving energy efficiency and throughput is achieved without sacrificing network performance.
Patent Information
- Application Number
- CN202510264568.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-06
- Publication Date
- 2025-06-17
AI Technical Summary
When facing dynamic network environments and strict user service quality requirements, modern communication systems face the problem of increased energy consumption, which hinders the development of key technologies such as 5G.
The deep reinforcement learning resource allocation method is adopted, and the optimization goals and constraints are determined by establishing a communication network model, optimizing the optimization problem into a Markov decision-making process, using PPO and D3QN algorithms to train the agent, and optimizing the resource allocation strategy to improve energy efficiency.
Without sacrificing network performance, reduce energy waste and costs, improve overall system throughput, improve energy efficiency and reduce power consumption, and realize green mobile edge computing.
Smart Images

Figure CN120166458A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of wireless communication technologies, and particularly to a deep reinforcement learning resource allocation method for green mobile edge computing. Background Art
[0002] With the exponential growth of the number of mobile users and computationally intensive applications, modern communication systems are facing a more dynamic network environment and more stringent user service quality requirements. To address this challenge, resource allocation has become one of the fundamental issues for optimizing network performance. However, so far, one of the main obstacles hindering the development of key technologies such as 5G is the significantly increased energy consumption.
[0003] The improvement of energy efficiency has become an important concern in the field of communication technologies in recent years. Energy efficiency refers to minimizing energy consumption while completing the same tasks. By reducing energy waste and costs without sacrificing network performance, energy-efficient communication systems are more sustainable and sound both ecologically and economically than existing communication systems. Therefore, for the energy efficiency problem of communication systems, an innovative solution is needed that can both improve network performance and reduce energy consumption. Summary of the Invention
[0004] To solve the above technical problems, the present invention provides a deep reinforcement learning resource allocation method for green mobile edge computing, including the following steps:
[0005] S1. Establish a communication network model, and initialize the communication environment, the number of base stations, the number of users, and the number of subcarriers;
[0006] S2. Determine the optimization objective and the constraint conditions according to the signal-to-noise ratio of the users and the system capacity;
[0007] S3. Convert the optimization problem into a Markov decision process, determine the agent, the state space, the action space, and the reward function, and use a deep reinforcement learning algorithm to train and allocate an optimal strategy for each agent;
[0008] S4. The agent continuously interacts with the environment through the PPO and D3QN algorithms to optimize and update the network parameters;
[0009] S5. Obtain the optimal resource allocation scheme.
[0010] The further limited technical solution of the present invention is:
[0011] Furthermore, in step S1, the communication environment is a heterogeneous network HetNet that supports mobile edge computing (MEC) in the downlink. The communication network model includes a macro base station and M small base stations, and each base station is equipped with a designated MEC server. The base stations and the MEC servers are connected through a wired network, and the base stations offload computationally intensive tasks to the MEC servers through the wired network. The communication network model also includes L orthogonal frequency division multiplexing (OFDM) subcarriers and two types of users, namely primary users and secondary users. All primary users are directly served by the macro base station, while secondary users are served by their associated small base stations respectively.
[0012] As described above for the deep reinforcement learning resource allocation method for green mobile edge computing, in step S2, the optimization objective is set to maximize the energy efficiency of the HetNet that supports MEC by optimizing the resource allocation strategy.
[0013] The maximum achievable downlink transmission rate of user (m, n) is defined as:
[0014]
[0015] where Nm represents the total number of users of the m-th base station, represents the signal-to-noise ratio of user (m, n); thus the energy efficiency is expressed as:
[0016]
[0017] where, represents the downlink transmit power of user (m, n), represents the static power of the m-th base station; further, the optimization objective is formulated as maximizing the average energy efficiency over a period of time, defined as:
[0018]
[0019] The constraint conditions are as follows:
[0020] (a)
[0021] (b)
[0022] (c)
[0023] where t0 represents an arbitrary time step and T represents the maximum time length; constraint condition (a) means that the transmit power of each user cannot exceed its maximum value P max , and constraint conditions (b) and (c) mean that each user can only be allocated to one subcarrier at each time step.
[0024] For the deep reinforcement learning resource allocation method for green mobile edge computing as described above, in step S3,
[0025] Agent: Take base station m as the agent, where 1 ≤ m ≤ M;
[0026] State space: Define the state space as where represents the channel state information measured by the m-th base station and user (m, n) at time step t, represents the transmit power, c t (m, n) represents the maximum transmission rate;
[0027] Action space: The agent has k subcarriers to choose from, where k ∈ K and K represents the total number of subcarriers; Define action A as the transmit power and subcarrier selected by the agent, where, A t (m, n) represents the one-hot encoding vector for subcarrier allocation, represents a non-negative real number less than P max ;
[0028] Reward function: When the agent executes an action and satisfies the limiting conditions (a), (b), and (c), it will obtain a reward value; According to the observation exchange range defined by the state, define the reward functions for any primary user (0, n) or secondary user (m, n) respectively:
[0029]
[0030] where, represents the static power, θ1 and θ3 represent the coefficients of the dynamic power terms, θ2 and θ4 represent the coefficients of the static power terms, and these four coefficients are used to adjust the relative proportion between the dynamic power and the static power, as well as to control the overall reward proportion.
[0031] For the deep reinforcement learning resource allocation method for green mobile edge computing as described above, in step S4, use a convolutional neural network to simulate the policy function and Q function, and train it using deep learning methods, extend D3QN to a continuous action space or a high-dimensional discrete value; Use the D3QN architecture to allocate subcarriers and the PPO architecture to allocate the transmit power. After the model training is completed, each agent performs distributed execution to obtain its own action output.
[0032] The deep reinforcement learning resource allocation method for green mobile edge computing as described above. In step S4, the D3QN network responsible for sub-band selection is used as the top-level reinforcement learning component, and the PPO algorithm is used at the bottom layer to train the Actor network to be responsible for the selection of transmission power; the top layer uses a discrete action space composed of sub-band indices, and the bottom layer is defined by a continuous action space as A power =[0, 1]; The bottom layer is executed after the top layer, and the defined action is Before the bottom layer Actor network sets the transmission power of the agent, it needs the sub-band selection of the top layer to determine its state input; at the beginning of each time slot, each agent sequentially executes two policies to determine its associated sub-band and transmission power.
[0033] The deep reinforcement learning resource allocation method for green mobile edge computing as described above. In step S4, the Actor networks and Critic networks of all agents are trained in parallel and independently. At the beginning of the training phase, the parameters of all agent Actor networks and Critic networks are independently initialized; during each time step, the agents first exchange their local observations to obtain the states of all agents, and the agents interact with the environment by repeatedly observing states, executing actions, and receiving rewards.
[0034] The deep reinforcement learning resource allocation method for green mobile edge computing as described above. In step S4, in the D3QN algorithm, the target value is calculated as follows:
[0035] y t = r t+1 + γq(s t+1 , argmax a q(s t+1 , a; θ e ); θ t )
[0036] That is, use the evaluation network to obtain the action corresponding to the optimal action value in the s t+1 state, and then use the target network to calculate the action value of this action to obtain the target value; where, θ e and θ t respectively represent the parameters of the evaluation network and the target network, q represents the state-action value function, and γ represents the discount factor.
[0037] The deep reinforcement learning resource allocation method for green mobile edge computing as described above. In step S4, in the PPO algorithm, each agent contains an Actor network and a Critic network, and their network weights are denoted as θ π , and θ Q, the Actor network also includes a parameterized Actor network and an original Actor network. The state S is input into the original Actor network, which uses a dual-head output. The discrete action is passed through the original Actor network and the parameterized Actor network to obtain the parameterized action, and the continuous action directly obtains the probability of the action through the original Actor network;
[0038] Each agent continuously interacts with the environment, inputting the current state s t into the original Actor network, and obtaining the processed action A according to the original Actor network and the parameterized Actor network t , executing this action, and obtaining the feedback r of the environment t , while the environment enters the next state s t+1 ; during this process, the agent obtains a reward value to evaluate the quality of the action; PPO stores the experience samples (s t , A t , r t , s t+1 ) obtained from the interaction with the environment into the experience pool D of the PPO algorithm, and continuously loops this sampling process until the experience pool is full. The sampling data in the experience pool D is used for the agent's training, so as to perform offline updates.
[0039] In the depth reinforcement learning resource allocation method for green mobile edge computing as described above, in step S4, during the learning phase, after obtaining the sample data of the batch sample number, the original Actor network, the parameterized Actor network, and the Critic network respectively use the sample data in the experience pool D to update and optimize their own networks; all the states in the experience pool are input into the Critic network to obtain the value function V of all the states t , and then calculate the advantage function A at each time t = R t - V t , R t = ∑ T>t γ T-t r T , r t represents the immediate reward obtained after executing the action a t , γ is a discount factor, and the loss function of the Critic network adopts the mean square error form, expressed as:
[0040] loss(θ Q ) = mean(A t ) 2
[0041] Then, the loss function is minimized by the gradient method, and the Critic network is updated by backpropagation;
[0042] Obtain the importance sampling ratio according to the ratio of the probabilities of the new and old policies in the PPO algorithm The new policy is obtained based on the real-time interaction between the agent and the environment, and the old policy is obtained based on the data in the sampled experience pool; The Actor network algorithm uses the following loss function:
[0043] loss(θ π ) = E′ t [min(r t (θ π )A t , clip(r t (θ π ), 1 - ∈, 1 + ∈)A t )]
[0044] Among them, clip(r t (θ), 1 - ∈, 1 + ∈) is used to limit r t (θ) within [1 - ∈, 1 + ∈], and ∈ is a hyperparameter representing the range for clip; The parameterized Actor network updates the network according to backpropagation.
[0045] The beneficial effects of the present invention are as follows:
[0046] (1) In the present invention, for the communication scenario where the actual environment changes rapidly, most traditional algorithms require a large amount of data sets including information such as channels and base stations. However, in the actual communication process, due to the fact that cellular heterogeneous networks are often high-density and multi-layered, information such as channel state information is usually difficult to obtain. The deep reinforcement learning algorithm adopted by the present invention can obtain information through interaction with the environment and obtain long-term maximum benefits through continuous learning, which is more realistic;
[0047] (2) In the present invention, the Dueling Double Deep Q Network algorithm incorporates the idea of the Double DQN algorithm on the basis of the Dueling DQN algorithm. The only difference between it and the Dueling DQN algorithm lies in the way of calculating the target value; In the Dueling DQN algorithm, the target value is calculated by using the target network to obtain the action values of all actions in the s t+1 state, and then calculating the target value based on the optimal action value. Due to the maximization operation here, the algorithm has an "overestimation" problem, which affects the accuracy of decision-making; In the D3QN algorithm, the target value is calculated by using the evaluation network to obtain s t+1The action corresponding to the optimal action value in the state is then used to calculate the action value of this action by the target network, thereby obtaining the target value. Through the interaction of the two networks, the "overestimation" problem of the algorithm is effectively avoided;
[0048] (3) In the present invention, the PPO algorithm is a reinforcement learning algorithm based on policy gradients, aiming to solve the problems of training instability and low sampling efficiency existing in traditional policy gradient methods; the core idea of the PPO algorithm is to introduce a clipping term in each update to limit the amplitude of policy updates, ensuring that the changes in the policy are not too large, thereby maintaining relatively stable training; the PPO algorithm also incorporates some advantages of the trust region method, such as being easier to implement and more general; since PPO can optimize the long-term cumulative reward, it performs excellently in mobile edge computing scenarios that require long-term planning and decision-making;
[0049] (4) In the present invention, in the edge computing network, while ensuring the quality of user services, the deep reinforcement learning algorithm is used to achieve the method of adaptively selecting subcarriers and transmit power, improving the overall throughput of the system, enhancing energy efficiency and reducing power consumption. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] Figure 1 is a schematic diagram of the overall process of the present invention;
[0051] Figure 2 is a schematic diagram of the model of the communication system in the embodiment of the present invention;
[0052] Figure 3 is a schematic diagram of the network structure combining the D3QN and PPO algorithms in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0053] A deep reinforcement learning resource allocation method for green mobile edge computing provided in this embodiment, as Figure 1 shown, includes the following steps:
[0054] S1. Establish a communication network model, and initialize the communication environment, the number of base stations, the number of users, and the number of subcarriers.
[0055] As Figure 2As shown in the figure, the scenario mainly considered in this embodiment is a heterogeneous network (HetNet) supporting downlink edge computing MEC. The communication network model includes a macro base station (Macrocell Base Station, MBS) and M small base stations (Small-cell Base Stations, SBSs), where each base station is equipped with a designated MEC server; the base station and the MEC server are connected through a wired network, and through the wired network, the base station can offload computationally intensive tasks to the MEC server; the communication network model also includes L orthogonal frequency division multiplexing (Orthogonal Frequency Division Multiplexing, OFDM) subcarriers and two types of users, namely primary users and secondary users; all primary users are directly served by the macro base station MBS, while secondary users are served by their associated small base stations respectively.
[0056] Each user can only be allocated one subcarrier at each time step. In this embodiment, the method mainly adopted is reinforcement learning to allocate resources for each base station-user pair, and by optimizing the resource allocation strategy, the energy efficiency of the HetNet supporting MEC is maximized.
[0057] S2. Determine the optimization objective and constraints according to the signal-to-noise ratio of the user and the system capacity; in this embodiment, the objective is to maximize the energy efficiency of the HetNet supporting MEC by optimizing the resource allocation strategy; the maximum achievable downlink transmission rate of user (m, n) is defined as:
[0058]
[0059] where N m represents the total number of users of the m-th base station, represents the signal-to-noise ratio of user (m, n).
[0060] Therefore, the energy efficiency is expressed as:
[0061]
[0062] where represents the downlink transmission power of user (m, n), represents the static power of the m-th base station.
[0063] To better achieve the long-term energy-saving operation of the network, the optimization objective is further expressed as maximizing the average energy efficiency over a period of time, defined as:
[0064]
[0065] The constraints are as follows:
[0066] (a)
[0067] (b)
[0068] (c)
[0069] Among them, t0 represents any time step, and T represents the maximum time length; the constraint condition (a) means that the transmission power of each user cannot exceed its maximum value P max , and the constraint conditions (b) and (c) mean that each user can only be assigned to one subcarrier at each time step.
[0070] S3. Convert the optimization problem into a Markov decision process, determine the agent, state space, action space, and reward function, and use the deep reinforcement learning algorithm to train to assign the optimal policy to each agent.
[0071] Agent: Take the base station m as the agent, where 1 ≤ m ≤ M; each agent independently updates its policy, each agent can collect information from its own area and explore the network environment, and each agent can select the subcarrier and transmission power by itself.
[0072] State space (State): Define the state space as Among them represents the channel state information measured by the m-th base station and the user (m, n) at time step t, represents the transmission power, and c t (m, n) represents the maximum transmission rate.
[0073] Action space (Action): The agent has k subcarriers to choose from, where k ∈ K and K represents the total number of subcarriers; define the action A as the transmission power and subcarrier selected by the agent, Among them, A t (m, n) represents the one-hot encoded vector for subcarrier allocation, represents a non-negative real number less than P max .
[0074] Reward function (Reward): When the agent executes an action and meets the restrictive conditions (a), (b), and (c), it will receive a reward value; the reward is a quantitative feedback from the environment, used to evaluate the effectiveness of the executed action based on the observed state; in order to reflect the optimization goal of the resource allocation task, a reward function form similar to the energy efficiency η is designed, where the reward is a quantitative feedback from the environment, used to evaluate the effectiveness of the executed action based on the observed state; in order to reflect the optimization goal of the resource allocation task, a reward function form similar to the energy efficiency η is designed.
[0075] According to the observation exchange range defined by the state, the reward functions for any primary user (0, n) or secondary user (m, n) are defined respectively:
[0076]
[0077] Among them, represents the static power, θ1 and θ3 represent the coefficients of the dynamic power terms, θ2 and θ4 represent the coefficients of the static power terms, and these four coefficients are used to adjust the relative ratio between the dynamic power and the static power, as well as to control the overall reward ratio.
[0078] S4. The agent continuously interacts with the environment through the Proximal Policy Optimization (PPO) and D3QN algorithms to optimize and update the network parameters.
[0079] The deep reinforcement learning algorithm is adopted to enable the agent to perform autonomous learning; the reinforcement learning algorithms include policy-based methods such as the Policy Gradient (PG) algorithm and the Actor-Critic (AC) algorithm; value-based methods such as Q-Learning and DQN; although the above traditional algorithms are simple and easy to implement, in practical applications, they cannot handle the problem of large state spaces, and the convergence speed of using the above algorithms is greatly reduced, and even the situation of unstable training will be encountered.
[0080] Therefore, the combined use of the D3QN and PPO algorithms in this embodiment can solve the deficiencies of the above algorithms. The convolutional neural network is used to simulate the policy function and the Q function, and deep learning methods are used for training. The D3QN is extended to the continuous action space or the high-dimensional discrete values, and the experience replay method in DQN is adopted to interact with the environment, increasing the robustness of the system; in addition, this embodiment also improves it, using the D3QN architecture to allocate subcarriers and the PPO architecture to allocate transmission power. After the model training is completed, each agent performs distributed execution to obtain its own action output, improving the training speed and stability and enhancing the system performance.
[0081] Such as Figure 3As shown, the top-level reinforcement learning component is the D3QN network responsible for sub-band selection, and the bottom layer uses the PPO algorithm to train the Actor network responsible for the selection of transmission power; the top layer uses a discrete action space composed of sub-band indices, and the bottom layer is defined by a continuous action space as A power = [0, 1]; The bottom layer executes after the top layer, and the defined action is The bottom-layer Actor network needs the top-level sub-band selection to determine its state input before setting the transmission power of the agent; at the beginning of each time slot, each agent sequentially executes two policies to determine its associated sub-band and transmission power.
[0082] To conform to the multi-agent setting with information exchange, the Actor networks and Critic networks of all agents are trained in parallel and independently, while a small amount of information is exchanged between the MEC server or between base stations through a wired network; at the beginning of the training phase, the parameters of all agent Actor networks and Critic networks are independently initialized; during each time step, the agents first exchange their local observations to obtain the states of all agents, and the agents interact with the environment by repeatedly observing states, executing actions, and receiving rewards.
[0083] The Dueling Double Deep Q Network (D3QN) algorithm incorporates the idea of the Double DQN algorithm based on the Dueling DQN algorithm. Its only difference from the Dueling DQN algorithm lies in the way of calculating the target value; in the Dueling DQN algorithm, the target value is calculated as follows:
[0084] y t = r t+1 + γ max a q(s t+1 , a; θ t )
[0085] That is, the target network is used to obtain the action values of all actions in the s t+1 state, and then the target value is calculated based on the optimal action value; due to the maximization operation here, the algorithm has an "overestimation" problem, affecting the accuracy of decision-making; in the above formula, ω t represents the target network parameters.
[0086] In the D3QN algorithm, the target value is calculated as follows:
[0087] y t = r t+1 + γ q(s t+1 , argmax a q(s t+1 , a; θ e); θ t )
[0088] That is, the evaluation network is used to obtain the action corresponding to the optimal action value in the s t+1 state, and then the target network is used to calculate the action value of this action, so as to obtain the target value; through the interaction of the two networks, the "overestimation" problem of the algorithm is effectively avoided; among them, θ e and θ t respectively represent the parameters of the evaluation network and the target network, q represents the state-action value function, and γ represents the discount factor.
[0089] In the PPO algorithm, each agent contains an Actor network and a Critic network, and their network weights are respectively denoted as θ π , and θ Q , the Actor network also includes a parameterized Actor network and an original Actor network. The state S is input into the original Actor network, and the original Actor network uses a dual-head output. The discrete action is obtained as a parameterized action through the original Actor network and the parameterized Actor network, and the continuous action directly obtains the action probability through the original Actor network.
[0090] Each agent continuously interacts with the environment, and inputs the current state s t into the original Actor network, and obtains the processed action A t according to the original Actor network and the parameterized Actor network, executes this action, and obtains the feedback r t of the environment. At the same time, the environment enters the next state s t+1 ; during this process, the agent obtains a reward value, which is used to judge the quality of the action.
[0091] PPO improves the online update strategy in the policy gradient algorithm, adopts the idea of the experience pool in DQN, and stores the experience samples (s t , A t , r t , s t+1 ) obtained by interacting with the environment into the experience pool D of the PPO algorithm, and keeps looping this sampling process until the experience pool is full. The sampling data in the experience pool D is used for the agent to train, so as to perform offline update.
[0092] In the learning stage, when the sample data of the batch sample number is obtained, the original Actor network, the parameterized Actor network, and the Critic network respectively use the sample data in the experience pool D to update and optimize their own networks; all the states in the experience pool are input into the Critic network to obtain the value function V t of all states, and then the advantage function A t= R t -V t , R t = ∑ T>t γ T-t r T , r t represents the immediate reward obtained after executing action a t , γ is a discount factor, and the loss function of the Critic network adopts the mean squared error form, expressed as:
[0093] loss(θ Q ) = mean(A t ) 2
[0094] Then, minimize the loss function through the gradient method and update the Critic network by backpropagation.
[0095] Obtain the importance sampling ratio according to the ratio of the new and old policy probabilities in the PPO algorithm The new policy is obtained based on the real-time interaction between the agent and the environment, and the old policy is obtained based on the data in the sampled experience pool. This ratio is used to ensure that the gap between the new and old policies is not too large when the policy is updated; the Actor network algorithm uses the following loss function:
[0096] loss(θ π ) = E′ t [min(r t (θ π )A t , clip(r t (θ π ), 1 - ∈, 1 + ∈)A t )]
[0097] Among them, clip(r t (θ), 1 - ∈, 1 + ∈) is used to limit r t (θ) within [1 - ∈, 1 + ∈], and ∈ is a hyperparameter representing the range for clip; the parameterized Actor network updates the network according to backpropagation.
[0098] S5. Obtain the optimal resource allocation scheme.
[0099] In this embodiment, initialize the green MEC communication environment, which includes 1 macro base station, 4 primary users, and 4 small base stations. Set the number of auxiliary users of all small base stations to 3; the maximum communication distances of the macro base station and the small base stations are 1 km and 0.1 km respectively; all users are randomly distributed within the specified coverage ranges of their assigned base stations.
[0100] Three orthogonal subcarriers are set for users to select, the maximum transmit power is 38 dBm, and the Gaussian white noise σ 2 =-114 dBm. The channel model uses the Jakes model to simulate the small-scale fading factor.
[0101] Initialize the network parameters. The original Actor network, parameterized Actor network, and Critic network of the PPO algorithm are three three-layer fully connected networks, and the number of neurons in the three hidden layers are 128, 64, and 64 respectively; the learning rates of the two Actor networks are both set to 0.00003, the learning rate of the Critic network is set to 0.0001, the learning rate of the DQN network is 0.0001, and the discount factor γ is 0.1.
[0102] Set the number of training rounds T = 10000 for the agent, set the number of training steps step = 100 in each round, optimize the neural network parameters using the Adma optimizer every 50 steps, record the reward situation according to the training process of the agent, and the agent continuously optimizes its own strategy according to the proposed algorithm to finally obtain the optimal resource allocation plan; finally, apply the trained model to the actual scenario, and users can independently select the optimal resource allocation plan.
[0103] Regarding the energy efficiency problem of communication systems, an innovative solution needs to be sought that can not only improve network performance but also reduce energy consumption; such a solution can not only meet users' needs for high-quality services but also reduce the waste of energy resources, contributing to environmental protection and economic sustainable development.
[0104] Mobile Edge Computing (MEC) deploys computing and storage devices close to the user side, which is an extension of cloud computing and allows various computing tasks to be processed on the MEC server associated with the base station; therefore, the computing workload is distributed to dispersed MEC servers, effectively alleviating network latency and congestion; MEC is one of the most promising technologies for enhancing network management capabilities and improving the quality of user experience.
[0105] This embodiment aims to propose an innovative resource allocation algorithm to study the energy efficiency optimization problem in heterogeneous networks (HetNets) supporting green edge computing (MEC); the goal of this embodiment is to maximize the long-term average energy efficiency by designing a resource allocation algorithm based on decentralized multi-agent deep reinforcement learning (MADRL); the algorithm of this embodiment has an observation exchange mechanism, which can promote strategy coordination among multiple agents, thereby effectively optimizing resource allocation; it provides an efficient and sustainable resource management solution for the future development of edge computing and communication networks, thus promoting the application and development of green communication technologies.
[0106] In addition to the above embodiments, the present invention may have other embodiments. Any technical solutions formed by equivalent replacement or equivalent transformation shall fall within the protection scope claimed by the present invention.
Claims
1. A deep reinforcement learning resource allocation method for green mobile edge computing, characterized in that: The following steps are involved: S1. Establish a communication network model, initialize the communication environment, number of base stations, number of users and number of subcarriers; S2. Determine the optimization target and constraints based on the user's signal-to-noise ratio and system capacity; S3. Convert the optimization problem into a Markov decision process, determine the agent, state space, action space, and reward function, use deep reinforcement learning algorithm training, and assign the optimal strategy to each agent; S4, the agent continuously interacts with the environment through PPO and D3QN algorithms to optimize and update network parameters; S5. Get the optimal resource allocation plan.
2. The deep reinforcement learning resource allocation method for green mobile edge computing according to claim 1, characterized in that: In step S1, the communication environment is a heterogeneous network HetNet supporting downlink edge computing MEC, and the communication network model includes a macro base station and M small base stations, each base station is equipped with a designated MEC server; The base station and MEC server are connected through a wired network, and the base station offloads computing-intensive tasks to the MEC server through the wired network; the communication network model also includes L orthogonal frequency division multiplexing subcarriers and two types of users, namely primary users and auxiliary users; all primary users are directly served by macro base stations, while auxiliary users are served by their associated small base stations.
3. The deep reinforcement learning resource allocation method for green mobile edge computing according to claim 1, characterized in that: In the step S2, the optimization goal is set to maximize the energy efficiency of the HetNet supporting MEC by optimizing the resource allocation strategy; The maximum downlink transmission rate that can be achieved by user (m, n) is defined as: Among them, N m represents the number of all users of the mth base station, represents the signal-to-noise ratio of user (m, n); therefore, the energy efficiency is expressed as: in, represents the downlink transmission power of user (m, n), represents the static power of the mth base station; the optimization objective is further expressed as maximizing the average energy efficiency over a period of time, which is defined as: The constraints are as follows: (a) (b) (c) Where t0 represents any time step, T represents the maximum time length; constraint (a) means that the transmission power of each user cannot exceed its maximum value P max ,Constraints (b) and (c) indicate that each user can only be assigned to one subcarrier at each time step.
4. The deep reinforcement learning resource allocation method for green mobile edge computing according to claim 3 is characterized in that: In the step S3, Agent: Base station m is taken as the agent, where 1≤m≤M; State Space: Define the state space as in represents the channel state information measured between the mth base station and user (m, n) at time step t, represents the transmission power, c t (m, n) represents the maximum transmission rate; Action space: The agent has k subcarriers to choose from, where k∈K, K represents the total number of subcarriers; action A is defined as the transmit power and subcarrier selected by the agent, Among them, A t (m, n) represents the one-hot encoding vector for subcarrier allocation, Indicates less than P max A non-negative real number; Reward function: When the agent performs an action and the constraints (a), (b), and (c) are met, a reward value will be obtained. According to the observation exchange range defined by the state, the reward function of any main user (0, n) or auxiliary user (m, n) is defined separately: in, represents static power, α1 and α3 represent the coefficients of dynamic power terms, α2 and α4 represent the coefficients of static power terms, and these four coefficients are used to adjust the relative proportion between dynamic power and static power, as well as to control the overall reward ratio.
5. The deep reinforcement learning resource allocation method for green mobile edge computing according to claim 1, characterized in that: In step S4, a convolutional neural network is used to simulate the policy function and the Q function, and a deep learning method is used for training, so that D3QN is extended to a continuous action space or a high-dimensional discrete value; The D3QN architecture is used to allocate subcarriers, and the PPO architecture is used to allocate transmission power. After the model training is completed, each agent performs distributed execution to obtain its own action output.
6. The deep reinforcement learning resource allocation method for green mobile edge computing according to claim 5, characterized in that: In step S4, the D3QN network responsible for subband selection is used as the top-level reinforcement learning component, and the bottom-level PPO algorithm is used to train the Actor network responsible for the selection of the transmission power; the top-level uses a discrete action space composed of subband indexes, and the bottom-level is defined as a continuous action space A power = [0, 1]; the bottom layer is executed after the top layer, and the action is defined as The bottom-level Actor network requires the top-level subband selection to determine its state input before setting the agent’s transmit power; at the beginning of each time slot, each agent executes two strategies in turn to determine its associated subband and transmit power.
7. The deep reinforcement learning resource allocation method for green mobile edge computing according to claim 6, characterized in that: In step S4, the Actor networks and Critic networks of all agents are trained in parallel and independently. At the beginning of the training phase, the parameters of the Actor networks and Critic networks of all agents are initialized independently. During each time step, the agents first exchange their local observations to obtain the states of all agents. The agents interact with the environment by repeatedly observing states, executing actions, and receiving rewards.
8. The deep reinforcement learning resource allocation method for green mobile edge computing according to claim 7, characterized in that: In step S4, in the D3QN algorithm, the target value is calculated as follows: y t =r t+1 +γq(s t+1 ,argmax a q(s t+1 ,a;a e );a t ) That is, using the evaluation network to obtain s t+1 The action corresponding to the optimal action value in the state is then calculated using the target network to obtain the target value; where α e and α t They represent the parameters of the evaluation network and the target network respectively, q represents the state-action value function, and γ represents the discount factor.
9. The deep reinforcement learning resource allocation method for green mobile edge computing according to claim 1, characterized in that: In step S4, in the PPO algorithm, each agent includes an Actor network and a Critic network, and their network weights are denoted as α π , and α Q , the Actor network also includes a parameterized Actor network and an original Actor network. The state S is input into the original Actor network. The original Actor network adopts a double-headed output. The discrete action is passed through the original Actor network and the parameterized Actor network to obtain the parameterized action. The continuous action directly obtains the probability of the action through the original Actor network; Each agent continuously interacts with the environment and sends the current state s t Input to the original Actor network, and get the processed action A according to the original Actor network and the parameterized Actor network t , execute the action and get feedback from the environment t , and the environment enters the next state s t+1 ; In this process, the agent gets a reward value to judge the quality of the action; PPO will get the experience sample (s t , A t , r t ,s t+1 ) is stored in the experience pool D of the PPO algorithm, and the sampling process is repeated until the experience pool is full. The sampled data in the experience pool D is used for agent training, thereby performing offline updates.
10. The deep reinforcement learning resource allocation method for green mobile edge computing according to claim 1, characterized in that: In step S4, in the learning phase, after obtaining the sample data of the batch sample number, the original Actor network, the parameterized Actor network and the Critic network respectively use the sample data in the experience pool D to update and optimize their own networks; all states in the experience pool are input into the Critic network to obtain the value function V of all states t , and then calculate the advantage function A at each time t =R t -V t , R t =∑ T>t γ T-t r T , r t Represents the execution of action a t The immediate reward obtained after the reward is obtained, γ is a discount coefficient, and the loss function of the Critic network is in the form of mean square error, expressed as: loss(a Q )=mean(A t ) 2 Then the loss function is minimized by the gradient method and the Critic network is updated by back propagation; The importance sampling ratio is obtained according to the ratio of the new and old strategy probabilities in the PPO algorithm. The new strategy is obtained based on the real-time interaction between the agent and the environment, and the old strategy is obtained based on the data in the sampled experience pool; the Actor network algorithm uses the following loss function: loss(a π )=E′ t [min(r t (a π )A t ,clip(r t (a π ),1-∈,1+∈)A t )] Among them, clip(r t (α), 1-∈, 1+∈) is used to convert r t (α) is limited to [1-∈, 1+∈], where ∈ is a hyperparameter representing the range of clipping; the parameterized Actor network is based on Backpropagation updates the network.
Citation Information
Cited By
Intelligent safety capacity improving method and device for guaranteeing deterministic time delay
CN120547627A
Power trusted wireless local area network access point layout design method and system
CN120881595A