A Multi-Agent Reinforcement Learning Method for Internet of Things Data Offloading

Through the multi-agent reinforcement learning method, the beamforming and resource allocation are jointly optimized in the IRS-assisted multi-user wireless Internet of Things system, solving the problem of minimizing system energy consumption and achieving more efficient learning and performance improvement.

CN114222368BActive Publication Date: 2025-06-27SUN YAT SEN UNIVERSITY SHENZHEN +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111442259.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-30
Publication Date
2025-06-27
Estimated Expiration
2041-11-30

AI Technical Summary

Technical Problem

In IRS-assisted Multi-Input Single Output (MISO) multi-user wireless IoT system, how to effectively reduce energy consumption to solve the problem of power supply for IoT devices.

Method used

Multi-agent reinforcement learning method is adopted to construct Markov decision-making process, jointly optimize active, passive beamforming and user resource allocation decisions, and solve global and local decision-making problems in a layered manner to minimize energy consumption.

Benefits of technology

It significantly improves the learning rate and performance, can regulate IRS more efficiently, optimize wireless communication resource allocation, and thus reduce the total power consumption of the base station.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114222368B_ABST
    Figure CN114222368B_ABST
Patent Text Reader

Abstract

The present invention discloses a multi-agent reinforcement learning method for Internet of Things (IoT) data offloading. The method includes: in a multi-terminal scenario of the IoT, jointly optimizing active and passive beamforming and resource allocation decisions of users, and formulating a power minimization problem; constructing a Markov decision process, and solving the power minimization problem based on multi-agent reinforcement learning. By using the present invention, the optimization problem is solved in a hierarchical manner, with improved multi-agent deep reinforcement learning, significantly improving the learning rate and performance. As a multi-agent reinforcement learning method for IoT data offloading, the present invention can be widely applied in the field of wireless communication.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of wireless communication, and particularly to a multi-agent reinforcement learning method for Internet of Things (IoT) data offloading. Background Art

[0002] With the rapid development of wireless communication networks, the number of terminals accessing the network is increasing day by day. IoT devices represented by sensor nodes will exist widely. How to ensure the power supply of these ubiquitous devices is a difficult problem that the IoT urgently needs to solve. A wireless-powered communication network uses radio frequency energy signals to transmit energy to passive terminals, which is an important way to solve the problem of energy limitation of IoT devices. In recent years, intelligent reflecting surface (IRS) has been considered as a technology with important prospects because it can improve the quality and spectral efficiency of wireless communication. Applying machine learning to the regulation of IRS has strong robustness. However, the offline training of deep neural network (DNN) relies on exhaustive search and alternating optimization (AO) method. Although it can learn the optimal strategy from scratch, the learning speed is usually slow. Summary of the Invention

[0003] The object of the present invention is to provide a multi-agent reinforcement learning method for IoT data offloading to solve the problem of minimizing energy consumption in an IRS-assisted multi-input single-output (MISO) multi-user wireless IoT system.

[0004] The first technical solution adopted by the present invention is as follows: A multi-agent reinforcement learning method for IoT data offloading, comprising the following steps:

[0005] In a multi-terminal scenario of the IoT, jointly optimize active and passive beamforming and resource allocation decisions of users, and formulate a power minimization problem;

[0006] Construct a Markov decision process, and solve the power minimization problem based on multi-agent reinforcement learning.

[0007] Furthermore, the framework of multi-agent reinforcement learning includes:

[0008] A high-level controller, used to observe the overall environment and estimate the actions of each agent;

[0009] Low-level user agents, used to observe their own states and their own actions.

[0010] Furthermore, the step of constructing a Markov decision process and solving the power minimization problem based on multi-agent reinforcement learning specifically includes:

[0011] Set the wireless transmitting base station as the controller, and set independent agents at both the controller and each user to obtain a controller agent and user agents;

[0012] Divide the action a t =(θ i , ω i , ρ i , τ i , k i ) into a global action a c,t =θ i and a local action a o,t =(ω i , ρ i , τ i , k i );

[0013] The controller agent optimizes the global action a c,t =θ i , i∈{1,…,N}, based on the deep reinforcement learning method, and estimates some local actions (ω i , ρ i ), i∈{1,…,N} based on the optimization method;

[0014] The user agent optimizes the action a u,i =(τ i , k i ) of this user agent based on the deep reinforcement learning method;

[0015] At the beginning of the iteration, the actor network of the controller agent outputs the global action a c,t =θ i , i∈{1,…,N}, and obtains (ω i , ρ i ), i∈{1,…,N}, the estimated values a u,i , i={1,…,N} of all user agent actions, and the estimated value y' of the lower bound of the target value, and distributes the target estimated value y' and the action estimates a u,-i of other agents to each user agent i;

[0016] The user agent i obtains the action estimates a u,-i of the other n - 1 users from the controller,

[0017] The actor network of the user agent i outputs the action a u,i =(τ i , k i ) of this user agent;

[0018] The target - critic network of the user agent i generates the target value y i ;

[0019] The user agent i obtains an estimated value y' of the lower bound of the target value from the controller;

[0020] The user agent i compares the user's target value y i and the estimated value y' of the lower bound of the target value, and uses the larger value as the target value for training the critic network of the user agent i;

[0021] The user agent i uses the action estimate a u,-i of the controller as an approximate substitute for the action information of other agents;

[0022] After all user agents are trained, the global action and the local action are combined to obtain the complete action a t =(a c,t , a o,t ), and interacts with the environment according to this complete action to obtain the reward value r t .

[0023] The target-critic network of the controller agent outputs the target value y of the controller;

[0024] The complete action a t , the reward value r t and the lower bound y' of the target value are fed back to the controller agent. The target value y of the controller is compared with the lower bound y' of the target value, and the larger target value is selected as the training target value of the critic network of the controller agent. The controller agent trains the actor network according to the policy gradient output by the critic network, and trains its critic network according to the time difference (TD) error between the output of the critic network and the target value;

[0025] When the network output value converges, the iteration stops; otherwise, it enters the next round of iteration.

[0026] The beneficial effects of the method of the present invention are as follows: By solving the optimization problem in layers, the upper layer optimizes and solves the global decision-making problem, that is, the control problem of IRS, and the lower layer uses multi-agent deep reinforcement learning to solve the local decision-making problem, that is, the resource allocation problem of users. Through the multi-agent deep reinforcement learning of the present invention, the learning rate and performance can be significantly improved. Description of the Drawings

[0027] Figure 1 is a schematic diagram of a multi-agent reinforcement learning of the present invention;

[0028] Figure 2 is a step flow chart of a multi-agent reinforcement learning method for Internet of Things data offloading of the present invention;

[0029] Figure 3 This is the IRS-assisted IoT wireless data offloading system in a specific embodiment of the present invention. Specific implementation manner

[0030] The present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments. For the step numbers in the following embodiments, they are only set for the convenience of elaboration and explanation, and no limitation is imposed on the order between the steps. The execution order of each step in the embodiments can be adaptively adjusted according to the understanding of those skilled in the art.

[0031] Referring to Figure 2 , the present invention provides a multi-agent reinforcement learning for IoT data offloading, and the method includes the following steps:

[0032] S1. In the multi-terminal scenario of the Internet of Things, jointly optimize the active and passive beamforming and the resource allocation decision of users, and intelligently allocate the amount of data offloaded by the terminal and the offloading time to achieve edge computing and data offloading of the Internet of Things, and formulate a power minimization problem;

[0033] Specifically, the application scenarios include multi-terminal smart home scenarios, information transmission scenarios of multiple users, data edge computing scenarios of passive intelligent devices, etc.;

[0034] S2. Construct a Markov decision process and solve the power minimization problem based on multi-agent reinforcement learning. Specific embodiment 1:

[0036] For the continuous variable control problem of the agent in a dynamic environment, the multi-agent deep deterministic policy gradient (MADDPG) method is used to solve it: Suppose there are n agents, and we design corresponding action sets a1,..., a n , and the corresponding observation quantities o1,..., o n . The state transition equation includes all states, actions, and observations: where The policy of each agent only includes its own state and action: μ i : o i →a i . Each agent uses a deep critic network to approximate the Q function, that is, the critic network of the i-th agent learns the action value function: where The parameters of the critic network will be updated in each decision period to output a better Q-value approximation. It can be achieved by minimizing the loss function through training a deep neural network (DNN):

[0037]

[0038] Wherein:

[0039]

[0040] represents the output of the target - critic network, is the target policy with a lagged update parameter θ' j of the target policy. represents the TD error between the Q - value and the target value. During the update of the Critic network, the global policy information is unknown, but each agent can learn and estimate the policies of other agents. Using approximate policies, MADDPG can learn effectively without requiring agents to communicate with each other to obtain global information. MADDPG maintains n - 1 policy approximation functions for each agent represents the function approximation of the policy of the j - th agent by the i - th agent. Its approximation cost is the logarithmic cost function, and adding the entropy of the policy, its cost function can be written as:

[0041]

[0042] As long as the above cost function is minimized, the approximation of the policies of other agents can be obtained. Therefore, y can be replaced by:

[0043]

[0044] Before updating update the approximate policies of other agents using a sampled batch from experience replay Meanwhile, each agent also uses a deep actor network to approximate the policy function. The actor network updates the network parameters through gradient estimates to improve the estimation of the Q - function, which only uses its own observation information. Using the deterministic policy gradient theorem, the gradient estimate is simplified as follows:

[0045]

[0046] It can be seen that the critic network borrows global information for learning, and the actor network only uses local observation information. The main advantage of MADDPG is that it can give the optimal action using only local information when applying the learned optimal policy, without the need to know the model of the environment and without special communication requirements between agents. Therefore, MADDPG can be used not only in the case of cooperation between agents but also in the case of confrontation between agents. Specific Embodiment 2:

[0048] Aiming at the problem of learning complexity and target value estimation of highly coupled multivariables, an optimization-driven Deep Deterministic Policy Gradient (DDPG) method is designed:

[0049] Specifically, the action a t =(θ i , ω i , ρ i , τ i , k i ) is divided into the global action a c,t =θ i and the local action a o,t =(ω i , ρ i , τ i , k i );

[0050] Define the Deep Deterministic Policy Gradient (DDPG) module and the optimization module. The DDPG module generates the global action a c,t , and the optimization module generates the local action a o,t .

[0051] At the beginning of the iteration, the DDPG module outputs the global action a c,t =θ i , and inputs it into the optimization model;

[0052] The optimization module first fixes the phase θ i , and solves the active beamforming strategy ω i by solving the equivalent convex problem. Then, fixing the above parameters, an inner iteration is carried out to alternately solve the reflection coefficient ρ i , the time slot division ratio τ i , and the data offloading ratio k i :

[0053] (1) Fix τ i , k i , and calculate the average value of the upper and lower bounds of ρ i as the estimated value of the current ρ i ;

[0054] (2) Fix ρ i , and use a convex optimization solver (CVX) to solve the original problem to obtain τ i , k i .

[0055] Alternately perform the above two steps to solve. When all three parameters converge to stable values, exit the inner iteration to obtain a o,t =(ω i , ρ i , τi , k i ), calculate the lower bound y' of the target value.

[0056] Combine the outputs of the DDPG module and the optimization module into the action set a t = (a c,t , a o,t ), and the action set interacts with the environment to obtain the reward value r t ;

[0057] The target critic network of the DDPG module generates the target value y.

[0058] Feed the action set a t , the reward r t and the lower bound y' of the target value back to the DDPG module for learning and updating the neural network parameters. Compare the two target values y and y', and select the larger target value as the training target value of the critic network of the DDPG module. The DDPG model learns and updates the network parameters and enters the next iteration; when the network output converges, the iteration stops. Specific Embodiment 3:

[0060] For the multi-agent information fusion problem, in the MADDPG framework, each user has an estimated target value y generated by the target critic network and an estimated policy of other users generated by the approximate policy network j≠i, thus completing the multi-agent information fusion. In the initial stage of learning, due to the random initialization of the critic network and the approximate policy network, the estimation of y and the policy estimation are far from the optimal value. This problem can be solved by using an optimization-driven hierarchical reinforcement learning method. Estimate the lower bound of the target value y and the approximate policies of other agents by solving the approximate optimization problem Specifically, divide the system participants into high-level controllers and low-level multi-user agents. The controller agent has a DDPG module and an optimization module. Through the optimization module in Example 2, the action estimation values of all agents can be efficiently obtained. Substitute the action estimation values into the optimization objective function to estimate the lower bound of the target value y, and the action estimation values of the users can be used as the policy estimations for different users. Define y' as the lower bound of the target value determined by the optimization method, and a u,i as the action estimation value obtained by solving the optimization problem. In the initial stage of learning, the model-based optimization method can usually provide a better target value, that is, y'>y. At this time, the lower bound of the target value and the action estimations a of other users u,-iDistribute to the low-level user agents to guide the low-level agents to learn more efficiently. The low-level user agents feedback the learned actions to the high-level controller, forming a complete action set of the low-level user agents. Accordingly, the controller can update its own policy and improve the reward function. Specific Embodiment 4:

[0062] For a specific multi-user wireless power transfer communication system, set the wireless transmitting base station (AP) as the controller, and set independent agents at both the controller and each user;

[0063] Divide the action a t =(θ i , ω i , ρ i , τ i , k i ) into a global action a c,t =θ i and a local action a o,t =(ω i , ρ i , τ i , k i );

[0064] The controller agent uses the DDPG module to optimize the global action a c,t =θ i , i∈{1,…,N}, and uses the optimization module to optimize a part of the local action (ω i , ρ i ), i∈{1,…,N}, and the user agent only uses the DDPG method to optimize the action a u,i =(τ i , k i ) of this user agent;

[0065] In the k-th iteration process, the flow of the algorithm is as follows:

[0066] The actor network of the controller DDPG module outputs the global action a c,t =θ i , i∈{1,…,N}, and inputs it into the optimization module;

[0067] The optimization module first fixes the phase θ i , and solves the active beamforming strategy ω i by solving the equivalent convex problem, and then fixes the above parameters and performs the inner iteration to alternately solve the reflection coefficient ρ i and the time slot division ratio τ i and the data offloading ratio k i :

[0068] (3) Fix τi , k i , calculate ρ i Take the average of the upper and lower bounds of i as the current estimated value of ρ;

[0069] (4) Fix ρ i , and use a convex optimization solver (CVX) to solve the original problem to obtain τ i , k i .

[0070] Alternately perform the above two steps to solve. When all three parameters converge to stable values, exit the inner iteration to obtain a o,t =(ω i , ρ i , τ i , k i ), i ∈ {1, …, N}, and use a part of this action as the estimated value of all user agents' actions a u,i =(τ i , k i ), i = {1, …, N}, and calculate the lower bound y' of the objective value.

[0071] Distribute the lower bound y' of the objective value and the action estimates a u,-i of other agents to the user agents;

[0072] The actor network of user agent i outputs the action a u,i =(τ i , k i ) of this user agent;

[0073] The target-critic network of user agent i generates the estimated objective value y i ;

[0074] User agent i obtains the action estimates a u,-i of the other n - 1 users and the lower bound y' of the objective value from the controller;

[0075] Compare the estimated objective value y i with the lower bound y' of the objective value, and use the larger value as the objective value for training the critic network of user agent i;

[0076] Use the action estimate a u,-i as an approximate substitute for the action information of other agents by a user agent i;

[0077] After all user agents are trained, combine the global action and the local action to obtain the complete action a t =(a c,t , a o,t), and interact with the environment according to this complete action to obtain the reward value r t .

[0078] The target evaluator (target-critic) network of the controller agent outputs the target value y;

[0079] Feed the complete action a t , the reward value r t and the lower bound of the target value y' back to the controller agent, compare the target value y and the lower bound of the target value y', and select the larger target value as the training target value of the critic network of the controller agent. The controller agent trains the actor network according to the policy gradient output by the critic network, and trains its critic network according to the time difference (TD) error between the output of the critic network and the target value;

[0080] Stop the iteration when the value output by the network converges, otherwise enter the next round of iteration.

[0081] The algorithm block diagram refers to Figure 1 .

[0082] Regulation objective: p0 is the fixed transmit power of the base station, represents the beamforming strategy for user i to transmit data uplink. The entire time period is divided into N parts, and the time length of each time slot is the unit time 1, τ i ∈[0,1] is the time slot division coefficient of user i. Our regulation objective is to minimize the total power consumption of the base station, including transmit power consumption and computing power consumption.

[0083] Constraint conditions: As Figure 3 shown is an IRS-assisted single-base station multi-antenna multi-user wireless network communication system. The base station is responsible for power supply and collecting the data to be processed uploaded by users. Under the IRS assistance, the composite channel from the base station (AP) to user i is defined as Our process is divided into downlink energy transmission and uplink data transmission. For this model, there are the following constraints:

[0084] Time constraint: Assume that each user has a certain amount of data L i to be processed, and the local data processing rate is C i . The user will divide the data into two parts and use the parameter k i to divide the data volume, where the total amount is k i L i of the data will be uploaded to the base station for processing. The entire time period is divided into N parts, and the time length of each time slot is the unit time 1, τ i ∈[0,1] is the time slot division coefficient of user i, τ iThe part performs downlink energy harvesting, 1 - τ i The part performs uplink data transmission, and the data transmission rate is p o,i represents the power of the transmission signal of the i-th user.

[0085] After the definition is completed, each user needs to complete data transmission within the uplink sub-slot and complete local data calculation within this time slot, so there are the following constraints:

[0086] (1 - τ i )o i ≥k i L i ≥L i -C i

[0087] Then for p o,i , there are the following constraints:

[0088]

[0089] User energy constraint: Assume that the power of user data processing is P l,i , E h,i represents the harvested energy. Then there are the following energy constraints:

[0090]

[0091] IRS energy constraint: The downlink complex channel matrix from the base station to the IRS is defined as Assume that the IRS has K elements, μ is the power consumption of each reflection unit of the IRS, and η0 is the conversion efficiency of the energy harvested by the IRS. During the IRS energy harvesting phase, the IRS reflects ρ 2 of the energy to the user and harvests the remaining 1 - ρ 2 of the energy for the IRS to operate. The IRS needs to assist both downlink communication and user uplink data transmission. The downlink communication time is τ i , and ω e,i is defined as the beamforming strategy for the i-th user's downlink transmission. Therefore, the IRS energy requirement satisfies the following constraints:

[0092]

[0093] Algorithm application: Our goal is to jointly optimize the active and passive beamforming and the resource allocation decisions of users (θ i , ω i , ρ i , τ i , k i), to minimize the total power consumption of the base station. This is a non-convex problem, and the non-convexity is mainly reflected in the high coupling of the control parameters (ρ, θ, ω). By decoupling each decision variable through a certain method, this problem can be efficiently solved:

[0094] Divide the action a t =(θ i , ω i , ρ i , τ i , k i ) into a global action a c,t =θ i and a local action a o,t =(ω i , ρ i , τ i , k i );

[0095] Layer the system. The upper layer uses the DDPG algorithm to solve for a c,t , and based on the optimization method, solve for a part of the local action (ω i , ρ i ), and at the same time estimate the local action a u,i of each agent;

[0096] Distribute the estimated action values obtained by optimization to the MADDPG algorithm in the lower layer, and use the MADDPG algorithm to solve for the local actions (τ i , k i ) of each user agent;

[0097] Combine the actions of the upper and lower layers to obtain the complete action a t =(a c,t , a o,t ) and apply it to the environment to obtain the reward r t ;

[0098] Feed the results back to the upper layer to enable the upper layer to learn and update. Repeat the above steps.

[0099] In the simulation, we used a fixed network topology to test the learning ability of the optimization-driven MADDPG. In the environment as Figure 3 shown, the distance relationship between the base station, users, and intelligent transmitting surfaces is as Figure 3As shown. The straight-line distance from the base station to the user is 10 meters, the distance from the base station to the intelligent reflecting surface is 3 meters, and the perpendicular distance from the intelligent reflecting surface to the connection line between the base station and the user is 0.5 meters. The signal propagation satisfies the log-distance model, the reference point path loss is 30 dB, the path loss exponent is equal to 2, the energy capture efficiency is set to 0.8, the number of users is set to 3, the time slot for each user is set to 1 ms, and the data volume is set to 1 M bits. Then, we first demonstrate the learning performance of the proposed algorithm, and then study the impact of different parameters on the power consumption of the base station. According to the simulation results, the method of the present invention has better performance, more stable and efficient learning performance and scalability.

[0100] The above is a specific description of the preferred embodiment of the present invention, but the present invention is not limited to the described embodiment. Those skilled in the art can also make various equivalent deformations or substitutions without departing from the spirit of the present invention, and these equivalent deformations or substitutions are all included in the scope defined by the claims of this application.

Claims

1. A multi-agent reinforcement learning method for Internet of Things data offloading, characterized in that, Including the following steps: Under the multi-terminal scenario of the Internet of Things, jointly optimize the active and passive beamforming and the resource allocation decision of users, and formulate a power minimization problem; Construct a Markov decision process and solve the power minimization problem based on multi-agent reinforcement learning; The step of constructing a Markov decision process and solving the power minimization problem based on multi-agent reinforcement learning specifically includes: Set the wireless transmission base station as the controller, and set independent agents at the controller and each user to obtain a controller agent and user agents; Perform action a t = (θ i , ω i , ρ i , τ i , k i ). Divide action a c,t = θ i into global action a o,t = (ω i , ρ i , τ i , k i ); The controller agent optimizes the global action a based on the deep reinforcement learning method c,t = θ i , i ∈ {1, …, N}, estimates partial local actions (ω i , ρ i ), i ∈ {1, …, N}; where θ i represents the phase, ω i represents the active beamforming strategy, ρ i represents the reflection coefficient, τ i represents the time slot division ratio, k i represents the data offloading ratio, and the entire time period is divided into N parts; The user agent optimizes the action a of this user agent based on the deep reinforcement learning method u,i =(τ i , k i ); The iteration starts, and the actor network of the controller agent outputs the global action a c,t = θ i , i ∈ {1, …, N}, and based on the optimization method, (ω i , ρ i ) are obtained, i ∈ {1, …, N}, the estimated values of all user agent actions a u,i , i = {1, …, N} and the estimated value y' of the lower bound of the target value. Then, the estimated value y' of the lower bound of the target value and the action estimates a u,-i of other agents are distributed to each user agent i; The user agent i obtains the action estimations a of the other n-1 users from the controller u,-i , The actor network of user agent i outputs the action a of this user agent u,i =(τ i , k i ); The target-critic network of user agent i generates the target value y of the user agent i ; User agent i obtains an estimated value y′ of the lower bound of the target value from the controller; The user agent i compares the target value y of the user i with the estimated value y' of the lower bound of the target value, and uses the larger value as the target value for training the critic network of the user agent i; The user agent i estimates the action a of the controller u,-i as an approximate substitute for the action information of other agents by the user agent i; All user agents have completed training. Combine the global action and the local action to obtain the complete action a t =(a c,t , a o,t ), and interact with the environment based on this complete action to obtain the reward value r t ; The target critic network of the controller agent outputs the target value y of the controller; The complete action a t , the reward value r t and the estimated value y' of the lower bound of the target value are fed back to the controller agent. The target value y of the controller is compared with the estimated value y' of the lower bound of the target value, and the larger target value is selected as the training target value for the critic network of the controller agent. The controller agent trains the actor network according to the policy gradient output by the critic network, and trains its critic network according to the temporal difference (TD) error between the output of the critic network and the target value; When the network output value converges, stop the iteration; otherwise, enter the next round of iteration.

2. The multi-agent reinforcement learning method for Internet of Things data offloading according to claim 1, wherein The framework of multi-agent reinforcement learning includes: A high-level controller, which is used to observe the overall environment and estimate the actions of each agent; Low-level user agents, which are used to observe their own states and decide their own actions.

Citation Information

Patent Citations

  • Resource allocation method of wireless information and power transfer technology

    CN111212438A

  • Intelligent access control and resource allocation method based on distributed A-C

    CN112887999A