A method for deploying and associating users of a UAV in an uncertain information environment

By optimizing drone deployment and user association strategies using the MADDPG network model, the problems of low system resource utilization and poor user experience caused by the uncertainty of user location and business needs were solved, thereby improving system resource utilization and user fairness.

CN116489664BActive Publication Date: 2026-04-21CHONGQING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CHONGQING UNIV OF POSTS & TELECOMM
Filing Date
2023-04-28
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing drone deployment and user association strategies result in low system resource utilization and poor user experience in scenarios where user location and business needs are random and uncertain.

Method used

The MADDPG network model is adopted to optimize UAV deployment and user association strategies by modeling user location statistics, channel model, UAV downlink transmission rate and system cost function. The trained MADDPG network is used to determine the UAV deployment and user association strategies.

Benefits of technology

In environments where users have uncertain information, this approach aims to improve system resource utilization and user fairness, thereby enhancing the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116489664B_ABST
    Figure CN116489664B_ABST
Patent Text Reader

Abstract

The present application relates to a kind of uncertain information environment unmanned aerial vehicle deployment and user association method, belong to wireless communication technical field.The method includes: S1: modeling communication system, unmanned aerial vehicle deployment area, and user-unmanned aerial vehicle association variable;S2: modeling user position statistical characteristics;S3: modeling channel model;S4: modeling unmanned aerial vehicle downlink transmission rate;S5: modeling user rate requirement;S6: modeling system cost function;S7: modeling user-unmanned aerial vehicle association and unmanned aerial vehicle deployment restriction condition;S8: modeling system state, action, observation and benefit function;S9: modeling and training MADDPG network, finally, using the trained MADDPG network to determine unmanned aerial vehicle deployment and user association strategy.The present application realizes system transmission performance optimization and user QoS promotion by optimizing unmanned aerial vehicle deployment and user association strategy jointly.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of wireless communication technology and relates to a method for deploying and associating unmanned aerial vehicles (UAVs) in uncertain information environments. Background Technology

[0002] In recent years, due to advancements in drone manufacturing technology and cost reductions, drones have been widely used in both civilian and military fields. In wireless communication systems, drones can be used as aerial base stations to provide on-demand communication services to hotspot areas. In drone-assisted wireless communication systems, optimizing drone deployment and user association schemes is a crucial issue affecting user experience and system performance. Existing work primarily focuses on designing drone deployment schemes based on determined user location information, with limited consideration for drone deployment and user association strategies in scenarios where user location and service needs are random and uncertain. Furthermore, existing research often focuses on optimizing system transmission performance to achieve drone deployment and user association, with less consideration for the differences between user service needs and drone service provision capabilities, resulting in limited system resource utilization and user fairness. Therefore, designing efficient drone deployment and user association strategies for environments with uncertain user information, thereby improving system resource utilization and user fairness, has become an urgent research topic. Summary of the Invention

[0003] In view of this, the purpose of this invention is to provide a method for drone deployment and user association in uncertain information environments, solving the problems of low system resource utilization and poor user experience caused by existing drone deployment and user association strategies. In this method, for a system containing multiple drone base stations and multiple users, the negative of the system cost function is modeled as the optimization objective to realize the drone deployment and user association strategy.

[0004] To achieve the above objectives, the present invention provides the following technical solution:

[0005] A method for deploying and associating drones with users in uncertain information environments, specifically including the following steps:

[0006] S1: Model the communication system, the drone deployment area, and the user-drone correlation variables;

[0007] S2: Modeling user location statistics;

[0008] S3: Modeling the channel model;

[0009] S4: Modeling the downlink transmission rate of the drone;

[0010] S5: Modeling user rate requirements;

[0011] S6: Modeling the system cost function;

[0012] S7: Modeling user-drone associations and drone deployment constraints;

[0013] S8: Model the system state, actions, observations, and payoff functions;

[0014] S9: Model and train the MADDPG network; finally, use the trained MADDPG network to determine drone deployment and user association strategies.

[0015] Furthermore, in step S1, constructing a communication system specifically includes: assuming the system consists of multiple ground users and multiple UAVs, let M represent the number of users in the system; deploying UAV base stations to provide downlink data transmission services to ground users, let UAV... k Let K represent the k-th drone, 1≤k≤K, where K represents the number of drone base stations deployed. Assume that the drone base station can access multiple users and that the system has sufficient spectrum resources. No interference is considered within or between base stations. The total bandwidth of the system is evenly divided into N sub-channels, each with a bandwidth of B.

[0016] Modeling the drone deployment area specifically includes: assuming the drone's flight altitude is fixed, discretizing the drone deployment area, and letting... These represent the maximum number of grid cells in the X and Y directions, respectively. Where X1 and Y1 are the lengths of the drone deployment area along the X and Y axes, respectively, and Δ x Δ y These represent the distances between adjacent grid cells along the X and Y axes, respectively. UAV k The location of the deployed grid is at the x-th position on the X-axis. k Column, the y-th column on the Y-axis k Okay, among them,

[0017] Modeling user-drone association variables, specifically including: Let (i,j) represent the grid located in the i-th column of the X-axis and the j-th row of the Y-axis, where, make This represents the user-drone association variable. UAV k Associated with the grid at (i,j), and vice versa.

[0018] Furthermore, in step S2, the statistical characteristics of user location are modeled, specifically including: letting f(x,y) represent the statistical characteristics of the user location at (x,y), and modeling it as a two-dimensional truncated Gaussian distribution. Where η is the normalization constant, μ x μ y , σ x, σ y Let p be the mean and standard deviation of x and y, respectively. i,j The probability of a user existing in the grid at position (i,j) is modeled as: p i,j ≈f(iΔ x ,jΔ y )Δ x Δ y .

[0019] Furthermore, in step S3, the channel model is modeled, specifically including: Let h i,j,k UAV k The average path loss between the grid at (i,j) and the grid at (i,j) is modeled as follows: in, They represent UAV k The line-of-sight (LoS) and non-line-of-sight (NLoS) transmission probabilities for sending data to the grid at position (i,j) are modeled as follows: Where a and b are constants, θ i,j,k UAV k The elevation angle of the link between the grids at (i,j) and (i,j) is modeled as follows: UAV k The Loss link channel gain between the grids at (i,j) and (i,j) is modeled as follows: Among them, f c It is the carrier frequency, c is the speed of light, and ζ is the carrier frequency. L The additional path loss of the Loss of Service (LoS) link caused by shadow fading follows a log-normal distribution with a mean of 0 and a standard deviation of σ. i,j,k UAV k The distance between the grid points at (i,j) and (i,j) is modeled as follows: h represents the drone's flight altitude; UAV k The NLoS link channel gain between the grids at (i,j) and (i,j) is modeled as follows: Where ζ N It is the additional path loss of the NLoS link caused by shadow fading, which follows a log-normal distribution with a mean of 0 and a standard deviation of σ.

[0020] Furthermore, in step S4, modeling the downlink transmission rate of the UAV specifically includes: letting R k UAV k The average transmission rate of the associated users is modeled as follows: Where, p k It is a UAV k The transmit power is N0, and the noise power spectral density is N0.

[0021] Furthermore, in step S5, the user rate requirement is modeled, specifically including: Let q i,j This represents the data transmission rate requirement of a grid user located at (i,j), which follows a mean of μ and a variance of σ. 2 If the UAV is a normally distributed random variable, then... k The average rate requirement Q of the associated users k The model is as follows:

[0022] Furthermore, in step S6, the system cost function is modeled, specifically including: Let F represent the system cost function. Considering the downlink transmission rate of the UAV and the rate requirements of associated users, F is modeled as:

[0023] Furthermore, in step S7, the user-drone association and drone deployment constraints are modeled, specifically including:

[0024] Modeling user-drone association constraints, including:

[0025]

[0026]

[0027] Model the limitations of drone deployment, including:

[0028]

[0029]

[0030] ③ Where d th This indicates the safe distance between adjacent drones.

[0031] Furthermore, in step S8, the system state, actions, observations, and reward function are modeled, specifically including: modeling the system state at step t. in, Let represent the set of drone deployment locations at step t, where Indicates the UAV at step t k Deployment location, Q t ={Q 1,t ,…,Q k,t ,…,Q K,t} represents the set of user rate demands associated with the drone at step t, where Q k,t Indicates the UAV at step t k The average rate requirement of the associated users, h t =[h i,j,k,t [] represents the set of channel gains between the UAV and the grid at step t, where h i,j,k,tIndicates the UAV at step t k The average path loss between the grid at (i,j) and the grid at (i,j); modeling the joint action a of the UAV at step t. t ={G t ,I t}, where G t The action space representing the movement of the drone in step t is... Let represent the set of drone association strategies at step t, where Indicates the UAV at step t k Association strategy with the grid at (i,j); Modeling the joint observations of the UAV at step t. t ={o 1,t ,…,o k,t ,…,o K,t},in, Indicates the UAV at step t k Observed values; modeling the system's profit function r t Let r be the negative of the system cost function at step t, i.e.: t =-F(s) t ,a t ), where F(s) t ,a t ) indicates that in state s t At that time, action a is used. t The corresponding system cost function.

[0032] Furthermore, in step S9, the MADDPG network is modeled and trained, specifically including: initializing the online policy network parameters θ of the UAV. μ Online Q-network parameters θ Q Target policy network parameters θ μ′ and the target Q network parameters θ Q′ Initialize the experience replay buffer; initialize the random process χ, and process the system state s. t Perform initialization; for UAV k It selects actions using its current policy network and stochastic process. Where, μ k UAV k The policy network, Representation of the policy network μ k The parameter, χ k,t Represents random noise; actions are applied to the system environment to obtain the reward value r. t and the next state s t+1 The data is stored in the experience replay buffer D; a batch of samples is drawn from D, and the online Q-network of the UAV is updated based on minimizing the loss function; based on the sample data and the Q-value generated by the online Q-network, the policy gradient update formula is used. Update its online policy network, in which, UAV k policy network μ k In parameters The expectations below Indicates UAV k Parameters of the policy network Differentiate, Indicates UAV k Action a selected by the policy network k Differentiate x = (o1,...,o) K () represents the set of all UAV observations. UAV k The parameters of the online Q-network are updated; the parameters of the target policy network and the target Q-network are updated using a soft update algorithm, specifically as follows: Where ε << 1 represents the soft update parameter of the target network.

[0033] The modeling is based on the MADDPG algorithm to determine drone deployment and user association strategies, specifically including:

[0034] Environmental observations are input into the MADDPG network, and drone deployment and user association policies are determined based on the output of the online policy network.

[0035] The beneficial effects of this invention are as follows: the method of this invention can effectively ensure that, in the case of uncertain information environment for users, based on the principle of maximizing system benefit function, an efficient drone deployment and user association strategy can be designed, which can effectively improve system resource utilization and user fairness, and improve user experience.

[0036] Other advantages, objectives, and features of the invention will be set forth in part in the description which follows, and in part will be apparent to those skilled in the art from the following examination, or may be learned from practice of the invention. The objectives and other advantages of the invention can be realized and obtained through the following description. Attached Figure Description

[0037] To make the objectives, technical solutions, and advantages of the present invention clearer, the preferred embodiments of the present invention will be described in detail below with reference to the accompanying drawings, wherein:

[0038] Figure 1 This is a schematic diagram of the system scenario built according to the present invention;

[0039] Figure 2 This is a flowchart illustrating the method for deploying and associating drones with users in uncertain information environments according to the present invention. Detailed Implementation

[0040] The following specific examples illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of the present invention. Unless otherwise specified, the following embodiments and features can be combined with each other.

[0041] The accompanying drawings are for illustrative purposes only and are schematic diagrams, not actual pictures. They should not be construed as limiting the invention. To better illustrate the embodiments of the invention, some parts in the drawings may be omitted, enlarged, or reduced, and do not represent the actual product dimensions. It is understandable to those skilled in the art that some well-known structures and their descriptions may be omitted in the drawings.

[0042] In the accompanying drawings of the embodiments of the present invention, the same or similar reference numerals correspond to the same or similar components. In the description of the present invention, it should be understood that if terms such as "upper," "lower," "left," "right," "front," and "rear" indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, they are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, the terms used to describe positional relationships in the drawings are only for illustrative purposes and should not be construed as limiting the present invention. For those skilled in the art, the specific meaning of the above terms can be understood according to the specific circumstances.

[0043] Please see Figures 1-2 , Figure 1 This is a schematic diagram of the system scenario built according to the present invention, such as... Figure 1 As shown, the system contains multiple drone base stations and multiple users. By jointly designing the optimal drone deployment and user association strategy, the negative number of the system cost function can be maximized.

[0044] Figure 2 This is a flowchart illustrating the method for deploying and associating drones with users in uncertain information environments according to the present invention. Figure 2 As shown in this embodiment, the method for deploying drones and associating users in uncertain information environments specifically includes the following steps:

[0045] S1, modeling the communication system, the drone deployment area, and user-drone correlation variables.

[0046] 1) Modeling communication systems.

[0047] Assume the system consists of multiple ground users and multiple drones, let M represent the number of users in the system; deploy drone base stations to provide downlink data transmission services to ground users, let UAVs... k Let K represent the k-th drone, 1≤k≤K, where K represents the number of drone base stations deployed. Assume that the drone base station can access multiple users and that the system has sufficient spectrum resources. No interference is considered within or between base stations. The total bandwidth of the system is evenly divided into N sub-channels, each with a bandwidth of B.

[0048] 2) Model the deployment area of ​​the drone.

[0049] Assuming the drone's flight altitude is fixed, the deployment area of ​​the drone is discretized, let... These represent the maximum number of grid cells in the X and Y directions, respectively. Where X1 and Y1 are the lengths of the drone deployment area along the X and Y axes, respectively, and Δ x Δ y Let these represent the distances between adjacent grid cells along the X and Y axes, respectively; UAV k The location of the deployed grid is at the x-th position on the X-axis. k Column, the y-th column on the Y-axis k Okay, among them,

[0050] 3) Model user-drone association variables.

[0051] Let (i,j) represent the grid located in the i-th column of the X-axis and the j-th row of the Y-axis, where, make This represents the user-drone association variable. UAV k Associated with the grid at (i,j), and vice versa. 1≤k≤K.

[0052] S2. Model user location statistics.

[0053] Let f(x,y) represent the statistical characteristics of the user's location at (x,y), which can be modeled as a two-dimensional truncated Gaussian distribution: Where η is the normalization constant, μ x μ y , σ x , σ y Let p be the mean and standard deviation of x and y, respectively. i,j The probability that a user exists in the grid at position (i,j) can be modeled as: p i,j ≈f(iΔ x ,jΔ y )Δ x Δy .

[0054] S3, Modeling the channel model.

[0055] Let h i,j,k UAV k The average path loss between the grid at (i,j) and the grid at (i,j) is modeled as follows:

[0056] in, They represent UAV k The line-of-sight (LoS) and non-line-of-sight (NLoS) transmission probabilities for sending data to the grid at (i,j) can be modeled as follows: Where a and b are constants, θ i,j,k UAV k The elevation angle of the link between the grids at (i,j) and (i,j) can be modeled as: UAV k The Loss-of-Stake (LoS) link channel gain between the grids at (i,j) and (i,j) can be modeled as: Among them, f c It is the carrier frequency, c is the speed of light, and ζ is the carrier frequency. L The additional path loss of the Loss of Service (LoS) link caused by shadow fading follows a log-normal distribution with a mean of 0 and a standard deviation of σ. i,j,k UAV k The distance between the grid points at (i,j) and (i,j) can be modeled as: h represents the drone's flight altitude; UAV k The NLoS link channel gain between the grids at (i,j) and (i,j) can be modeled as: Where ζ N It is the additional path loss of the NLoS link caused by shadow fading, which follows a log-normal distribution with a mean of 0 and a standard deviation of σ.

[0057] S4, Modeling the downlink transmission rate of the drone.

[0058] Let R k UAV k The average transmission rate of the associated users is modeled as follows:

[0059] Where, p k It is a UAV k The transmit power is N0, and the noise power spectral density is N0.

[0060] S5, Model user rate requirements.

[0061] Let q i,jThis represents the data transmission rate requirement of a grid user located at (i,j), which follows a mean of μ and a variance of σ. 2 If the UAV is a normally distributed random variable, then... k The average rate requirement Q of the associated users k The model is as follows:

[0062] S6, Modeling system cost function.

[0063] Let F denote the system cost function. Considering the downlink transmission rate of the UAV and the rate requirements of associated users, F is modeled as:

[0064]

[0065] S7, Modeling User-Drone Associations and Drone Deployment Restrictions.

[0066] 1) Model user-drone association constraints, including:

[0067]

[0068]

[0069] 2) Model the deployment constraints of drones, including:

[0070]

[0071]

[0072] ③ Where d th This indicates the safe distance between adjacent drones.

[0073] S8. Model the system state, actions, observations, and payoff function.

[0074] Modeling the state of the system at step t in, Let represent the set of drone deployment locations at step t, where Indicates the UAV at step t k Deployment location, Q t ={Q 1,t ,…,Q k,t ,…,Q K,t} represents the set of user rate demands associated with the drone at step t, where Q k,t Indicates the UAV at step t k The average rate requirement of the associated users, h t =[h i,j,k,t [] represents the set of channel gains between the UAV and the grid at step t, where hi,j,k,t Indicates the UAV at step t k The average path loss between the grid at (i,j) and the grid at (i,j); modeling the joint action a of the UAV at step t. t ={G t ,I t}, where G t The action space representing the movement of the drone in step t is... Let represent the set of drone association strategies at step t, where Indicates the UAV at step t k Association strategy with the grid at (i,j); Modeling the joint observations of the UAV at step t. t ={o 1,t ,…,o k,t ,…,o K,t},in, Indicates the UAV at step t k Observed values; modeling the system's profit function r t Let r be the negative of the system cost function at step t, i.e.: t =-F(s) t ,a t ), where F(s) t ,a t ) indicates that in state s t At that time, action a is used. t The corresponding system cost function.

[0075] S9. Model and train the MADDPG network; finally, use the trained MADDPG algorithm network to determine the drone deployment and user association strategy.

[0076] Initialize the online policy network parameters θ of the drone μ Online Q-network parameters θ Q Target policy network parameters θ μ and target Q network parameters θ Q′ Initialize the experience replay buffer; initialize the UAV's online policy network parameters θ. μ Online Q-network parameters θ Q Target policy network parameters θ μ′ and the target Q network parameters θ Q′ Initialize the experience replay buffer; initialize the random process χ, and process the system state s. t Perform initialization; for UAV k It selects actions using its current policy network and stochastic process. Where, μ k UAV k The policy network, Representation of the policy network μk The parameter, χ k,t Represents random noise; actions are applied to the system environment to obtain the reward value r. t and the next state s t+1 The data is stored in the experience replay buffer D; a batch of samples is drawn from D, and the online Q-network of the UAV is updated based on minimizing the loss function; based on the sample data and the Q-value generated by the online Q-network, the policy gradient update formula is used. Update its online policy network, in which, UAV k policy network μ k In parameters The expectations below Indicates UAV k Parameters of the policy network Differentiate, Indicates UAV k Action a selected by the policy network k Differentiate x = (o1,...,o) K () represents the set of all UAV observations. UAV k The parameters of the online Q-network are updated; the parameters of the target policy network and the target Q-network are updated using a soft update algorithm, specifically as follows: Where ε << 1 represents the soft update parameter of the target network.

[0077] The trained MADDPG algorithm network is used to determine drone deployment and user association strategies, including:

[0078] Environmental observations are input into the MADDPG network, and drone deployment and user association policies are determined based on the output of the online policy network.

[0079] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A method for deploying and associating unmanned aerial vehicles (UAVs) in uncertain information environments, characterized in that, include: S1: Modeling the communication system, UAV deployment area, and user-UAV relationship variables; S2: Modeling user location statistics; S3: Modeling the channel model; S4: Modeling the downlink transmission rate of the drone; S5: Modeling user rate requirements; S6: Modeling the system cost function; S7: Modeling user-drone associations and drone deployment constraints; S8: Model the system state, actions, observations, and payoff functions; S9: Model and train the MADDPG network; finally, use the trained MADDPG network to determine drone deployment and user association strategies. In step S1, the communication system is constructed, specifically including: assuming the system consists of multiple ground users and multiple drones, let... This indicates the number of users in the system; deploying drone base stations provides downlink data transmission services to ground users, enabling... Indicates the first k One drone, , This represents the number of drone base stations deployed; it is assumed that each drone base station can connect to multiple users, and that the system has sufficient spectrum resources, with no interference considered within or between base stations; the total system bandwidth is evenly divided into... There are 1 sub-channel, and the bandwidth of each sub-channel is 1. ; Modeling the drone deployment area specifically includes: assuming the drone's flight altitude is fixed, discretizing the drone deployment area, and letting... , These represent the maximum number of grid cells in the X and Y directions, respectively. , ,in, , These are the lengths of the drone deployment area along the X and Y axes, respectively. , Let these represent the distances between adjacent grid cells along the X and Y axes, respectively; express The location of the deployed grid is on the X-axis. Column, Y-axis Okay, among them, , ; Modeling user-drone association variables, specifically including: Let Indicates the position located on the X-axis. Column, Y-axis A grid of rows, where , ;make This represents the user-drone association variable. express and The mesh association at the location, and vice versa. , ; In step S2, the statistical characteristics of user location are modeled, specifically including: Let express The statistical characteristics of the user's location are modeled as a two-dimensional truncated Gaussian distribution: ,in, The normalization constant is , , , They are and The mean and standard deviation of ; let . express The probability of a user existing in a grid cell is modeled as follows: ; In step S3, the channel model is modeled, specifically including: Let express and The average path loss between grid cells is modeled as follows: ,in, , They represent Send data to The line-of-sight and non-line-of-sight transmission probabilities of the grid are modeled as follows: , ,in, and It is a constant. express and The elevation angle of the link between the grids is modeled as follows: ,in, express and The distance between the grid points is modeled as follows: , h Indicates the drone's flight altitude; express and The line-of-sight link channel gain between grid cells is modeled as follows: ,in, It is the carrier frequency. c It's the speed of light. This is the additional path loss in line-of-sight links caused by shadow fading, with a mean of 0 and a standard deviation of [missing value]. The log-normal distribution; express and The non-line-of-sight link channel gain between grids is modeled as follows: ,in This is the additional path loss in non-line-of-sight links caused by shadow fading, with a mean of 0 and a standard deviation of [missing value]. The log-normal distribution; In step S4, the downlink transmission rate of the UAV is modeled, specifically including: Let express The average transmission rate of the associated users is modeled as follows: ,in, yes The transmission power, It is the noise power spectral density; In step S5, modeling user rate requirements specifically includes: Let Indicates that it is located at The data transmission rate requirements of grid users follow the average value. The variance is For a normally distributed random variable, then Average rate requirement of associated users The model is as follows: ; In step S6, modeling the system cost function specifically includes: Let The system cost function is represented by the downlink transmission rate of the UAV and the rate requirements of associated users, and is modeled accordingly. for: ; In step S7, the user-drone association and drone deployment constraints are modeled, specifically including: Modeling user-drone association constraints, including: ① ; ② ; Model the limitations of drone deployment, including: ① ; ② ; ③ ,in Indicates the safe distance between adjacent drones; In step S8, the system state, actions, observations, and payoff function are modeled, specifically including: modeling the system in the... Step state ,in, Indicates the first A collection of drone deployment locations, including Indicates the first step Deployment location, Indicates the first The set of user rate demands associated with drones, among which Indicates the first step The average rate requirement of the associated users Indicates the first The set of channel gains between the drone and the grid, where, Indicates the first step and Average path loss between grids; modeling the first Joint operations of infantry and unmanned aerial vehicles ,in, Indicates the first The action space for the movement of a step-by-step drone, i.e. , Indicates the first The set of drone association strategies, among which, Indicates the first step and The association strategy of the grid at the location; modeling the first Joint observations from unmanned aerial vehicles ,in, Indicates the first step Observed values; modeling the system's profit function For the first The negative of the system cost function, i.e.: ,in, Indicates the state At that time, take action The corresponding system cost function.

2. The method for deploying and associating unmanned aerial vehicles (UAVs) in uncertain information environments according to claim 1, characterized in that, In step S9, the MADDPG network is modeled and trained, specifically including: initializing the online policy network parameters of the UAV. Online Q network parameters Target policy network parameters and target Q network parameters Initialize the experience playback buffer; initialize the random process. and the system status Perform initialization; for It selects actions using its current policy network and stochastic process. ,in, express The policy network, Representational Policy Network The parameters, Represents random noise; applies actions to the system environment to obtain the reward value. and the next step The data is stored in the experience replay buffer D; a batch of samples is drawn from D, and the online Q-network of the UAV is updated based on minimizing the loss function; based on the sample data and the Q-value generated by the online Q-network, the policy gradient update formula is used. Update its online policy network, in which, express Policy network In parameters The expectations below Indicates to Parameters of the policy network Differentiate, Indicates to Actions selected by the policy network Differentiate, Represents the set of all UAV observations. express The parameters of the online Q-network are updated; the parameters of the target policy network and the target Q-network are updated using a soft update algorithm, specifically as follows: , ,in These are the soft update parameters for the target network.

Citation Information

Patent Citations

  • Unmanned aerial vehicle base station deployment and user association method based on Q learning algorithm

    CN113286314A

  • Multi-beam satellite communication system resource allocation method based on MADDPG algorithm

    CN115441939A