User association and uav trajectory optimization method based on multi-intelligent reflecting surface assisted multi-uav communication network

By employing phase alignment and an improved A3C algorithm for joint optimization in a multi-UAV communication network assisted by multiple intelligent reflectors, the problems of dynamic user association and UAV trajectory optimization are solved, and an efficient and reliable communication system in urban environments is realized.

CN118612767BActive Publication Date: 2025-11-28CHONGQING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410673173.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-05-28
Publication Date
2025-11-28
Estimated Expiration
2044-05-28

AI Technical Summary

Technical Problem

In multi-intelligent reflector-assisted multi-UAV communication systems, dynamic user association decisions are difficult, especially in urban environments where line-of-sight link channel obstruction is not fully considered, resulting in insufficient communication performance of the communication system.

Method used

This paper proposes a user association and UAV trajectory optimization method based on a multi-intelligent reflector-assisted multi-UAV communication network. By designing a phase-aligned intelligent reflector phase control strategy and combining it with an improved A3C algorithm for joint optimization, the method maximizes the hybrid channel gain between UAVs and users.

Benefits of technology

It achieves efficient and reliable communication for multi-UAV communication networks in urban environments with obstacles, and improves the average system transmission rate through dynamic service deployment and line-of-sight link judgment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118612767B_ABST
    Figure CN118612767B_ABST
Patent Text Reader

Abstract

The application discloses a user association and unmanned aerial vehicle trajectory optimization method based on a multi-intelligent reflecting surface assisted multi-unmanned aerial vehicle communication network. First, a phase control strategy of an intelligent reflecting surface based on phase alignment is designed. According to the phase control strategy designed by the application, the lake and channel gain between the users of the multi-unmanned aerial vehicle are derived. In order to maximize the average transmission power of the multi-unmanned aerial vehicle, user association, line-of-sight link judgment and multi-unmanned aerial vehicle estimation optimization are jointly optimized, and an improved A3C algorithm is designed to solve the joint optimization problem, so as to ensure the long-term optimization of the average transmission rate between the unmanned aerial vehicles and the users in the system. In order to ensure the stability of the system, the application maximizes the average transmission rate of the system, and realizes efficient and reliable communication of the multi-intelligent reflecting surface assisted multi-unmanned aerial vehicle communication network. The application provides a new method for user association and unmanned aerial vehicle trajectory optimization of the multi-intelligent reflecting surface assisted multi-unmanned aerial vehicle communication network.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application relates to a method for jointly dynamically allocating unmanned aerial vehicles and intelligent reflecting surfaces in a multi-intelligent reflecting surface assisted multi-unmanned aerial vehicle communication network in an academic field, in particular to a joint optimization method based on an asynchronous actor-critic network of deep reinforcement learning. BACKGROUND

[0002] Unmanned aerial vehicle communication technology and intelligent reflecting surface technology are widely used in wireless communication systems due to their flexibility, low cost and ease of deployment. Unmanned aerial vehicle technology not only can establish a temporary network for areas with weak network coverage or congestion, but also can achieve better communication performance by shortening the distance between the user and the unmanned aerial vehicle. However, in cities with high-rise buildings, the line-of-sight link channel between the unmanned aerial vehicle and the user will still be blocked. Intelligent reflecting surfaces can achieve passive beam shaping through phase shift adjustment to improve the transmission rate of the unmanned aerial vehicle communication system. Therefore, the integration of intelligent reflecting surfaces and unmanned aerial vehicles has good prospects in the enhancement of wireless communication. A major problem is that the communication system of the unmanned aerial vehicle is dynamic, resulting in a time-varying association between the intelligent reflecting surface and the user, especially in a multi-intelligent reflecting surface assisted multi-unmanned aerial vehicle communication system, the dynamic user association decision becomes more difficult. Therefore, it is necessary to further explore how to make full use of the reflecting units of the intelligent reflecting surface in a mobile scenario. In addition, current research on joint optimization of unmanned aerial vehicles and intelligent reflecting surfaces does not consider the problem of line-of-sight link channel blocking caused by unmanned aerial vehicle position changes in urban scenarios. Therefore, further research is needed on high-performance joint optimization methods. SUMMARY

[0003] The purpose of the application is mainly aimed at some deficiencies of the existing research, and a user association and unmanned aerial vehicle trajectory optimization method based on a multi-intelligent reflecting surface assisted multi-unmanned aerial vehicle communication network is proposed. The technical problem of jointly dynamically allocating unmanned aerial vehicles and intelligent reflecting surfaces in a multi-intelligent reflecting surface assisted multi-unmanned aerial vehicle communication network is solved. In order to maximize the average transmission rate of the system in a multi-unmanned aerial vehicle and obstacle existing scenario, efficient and reliable communication under a multi-intelligent reflecting surface assisted multi-unmanned aerial vehicle communication network is realized. The application designs an online learning strategy based on an A3C algorithm to solve the joint user association and unmanned aerial vehicle trajectory optimization problem in the time-varying channel.

[0004] Therefore, the technical scheme adopted by the application is that the user association and unmanned aerial vehicle trajectory optimization method based on a multi-intelligent reflecting surface assisted multi-unmanned aerial vehicle communication network comprises the following steps:

[0005] Step 1: Construct a system model, confirm a transmission channel model and a transmission rate model, and construct an optimization problem with the optimization objective of maximizing the average system total transmission rate as:

[0006]

[0007] K represents the number of UAVs, N represents the number of users, I represents the number of IRSs, R k [t] represents the sum of transmission rates of each user group of the UAVs at time slot t, q k [t] represents the coordinates of the UAV k at time slot t, r represents the number of columns of the IRS array, q represents the position of the UAV, tau represents the time slot, and T represents the number of time slots.

[0008] Constraint C1 defines a binary variable beta i,n [t] takes values 0 and 1; constraint C2 indicates that each user n has and only has one IRS providing service; constraint C3 indicates that the number of users served by each IRS cannot exceed the number of its rows; constraint C4 indicates that each UAV flies uniformly at a speed v, and the moving distance in each time slot is fixed; constraint C5 indicates that the line-of-sight link judgment variable mu k,n [t] takes values 0 and 1; constraint C6 limits the phase shift variable phi n,i [t] takes values between 0 and 2pi;

[0009] Step 2: The optimization problem of maximizing the average system total transmission rate in step 1 is solved by using beam alignment to obtain the optimal phase shift strategy and transforming the problem, and then the transformed optimization problem is modeled as a Markov decision process;

[0010] Step 3: The transformed optimization problem in step 2 is solved by using an improved A3C algorithm.

[0011] The beneficial effects of the present application include:

[0012] The present application constructs a dynamic service deployment framework for realizing high-rate communication in a communication network with multiple intelligent reflecting surfaces assisting multiple unmanned aerial vehicles. In order to ensure high-rate communication of the unmanned aerial vehicle communication network, the present application first proposes an intelligent reflecting surface phase control strategy based on phase alignment, and derives the hybrid channel gain between the unmanned aerial vehicle and the user according to the strategy. Then the joint user association, line-of-sight link judgment and multi-UAV trajectory optimization problem is transformed into a Markov decision process, and an improved A3C algorithm is designed to realize the maximum average system transmission rate. The experimental results prove the efficiency of the present application in terms of average system transmission rate. The present application provides a new user association and UAV trajectory optimization method applied to a multi-intelligent reflecting surface assisted multi-unmanned aerial vehicle communication network.

[0013] Step one constructs an optimization problem for multi-intelligent reflector-assisted multi-UAV communication in obstacle-prone scenarios. This optimization problem maximizes the system's average transmission rate by jointly optimizing joint user scheduling, line-of-sight link determination, phase shift matrix, and UAV trajectory optimization. Step two obtains the optimal phase shift strategy using beam alignment and transforms the optimization problem to derive the hybrid channel gain between UAVs and users, ensuring that the optimal hybrid channel gain can be found regardless of the positions of UAVs and users. Step three transforms the optimization problem into a Markov decision process, modeling it as a continuous decision process, and obtains the optimal strategy through interaction with the environment. Furthermore, based on the improved A3C algorithm, a noisy network is introduced into the policy network. This method adds noise weights to the network, requiring the selection of policy actions based on the currently introduced noisy policy and obtaining a reward. Compared to direct exploration in the policy space, this method saves on artificial entropy loss in policy selection. Attached Figure Description

[0014] Figure 1 A model diagram of a multi-intelligent reflector-assisted multi-UAV communication system;

[0015] Figure 2 This is a comparison chart of the convergence of the proposed algorithm MUMR with the comparative algorithms TDQ and FUA.

[0016] Figure 3 The average transmission rate of the proposed algorithm MUMR compared with the comparative algorithms TDQ and FUA under different numbers of reflective elements on the smart reflective surface;

[0017] Figure 4 The average transmission power of the proposed algorithm MUMR and the comparative algorithms TDQ and FUA under different transmit powers is shown.

[0018] Figure 5 The average transmission power of the proposed algorithm MUMR and the comparative algorithms TDQ and FUA under different UAV network scales is shown.

[0019] Figure 6 The average transmission power of the proposed algorithm MUMR is compared with the comparison algorithms TDQ and FUA under different numbers of users. Detailed Implementation

[0020] The technical solutions of the present invention will be clearly and thoroughly described below with reference to the accompanying drawings of the embodiments. The described embodiments are merely some embodiments of the present invention.

[0021] The user association and unmanned aerial vehicle trajectory optimization method of a multi-unmanned aerial vehicle communication network based on multi-intelligent reflecting surface assistance comprises the following steps:

[0022] Step 1: Construct a system model, confirm a transmission channel model and a transmission rate model.

[0023] As shown in Figure 1 , the present application constructs a system model, which contains K unmanned aerial vehicles, N users and I intelligent reflecting surfaces, each of which is composed of rxc reflecting units. r represents the number of columns of the intelligent reflecting surface array, and c represents the number of rows of the intelligent reflecting surface array. In each time slot t, the unmanned aerial vehicles and the users are divided into two channel conditions according to the propagation environment, i.e. only a reflection channel and a mixed reflection and direct channel. In each time slot t, multiple unmanned aerial vehicles serve different users. The coordinates of unmanned aerial vehicle k in time slot t are q k [t]=[x k [t],y k [t],z k [t]] T , the intelligent reflecting surface i and the user n are fixed at coordinates q i =[x i ,y i ,z i ] T and q n =[x n ,y n ,z n ] T . The distance d k,i [t]=||q k [t]-q i || between unmanned aerial vehicle k and intelligent reflecting surface i in time slot t, the distance d k,n [t]=||q k [t]-q n || between unmanned aerial vehicle and user, and the distance d i,n [t]=||q i -q n || between intelligent reflecting surface and user can be obtained respectively. Considering that the moving distance of the unmanned aerial vehicle in the time slot is much smaller than the distance between the unmanned aerial vehicle and the user, the present application approximates the position of the unmanned aerial vehicle in the time slot as a constant value. In time slot t, the direct channel between unmanned aerial vehicle k and user n is:

[0024]

[0025] wherein α k,n and represent the path loss exponent and the Rice fading factor between unmanned aerial vehicle k and user n respectively; γ represents the channel gain at one meter reference; denotes the non-line-of-sight component.

[0026] The reflection link is composed of two parts, the first part is the link from UAV to smart reflector, the second part is the link from smart reflector to user. In this invention, considering that there is an obstacle between UAV and user, a virtual line-of-sight channel is established through the smart reflector. Then, the non-line-of-sight channel component from UAV k to user n in time slot t is:

[0027] L k,i,n [t] = a(H i,n [t]) H Φ i,n [t]H k,i,n [t]

[0028] where a denotes the amplitude loss caused by the smart reflector; H k,i,n [t] denotes the channel vector from UAV k to smart reflector i in time slot t; H i,n [t] denotes the channel vector from smart reflector i to user n in time slot t. Φ i,n [t] denotes the phase shift matrix between user n and smart reflector i. In time slot t, the channel vector from UAV k to smart reflector i is denoted as:

[0029]

[0030] where h k,i,n [t] denotes the line-of-sight link component from UAV k to smart reflector i. In time slot t, the channel component from smart reflector i to user n is denoted as:

[0031]

[0032] where h i,n [t] denotes the line-of-sight link component from smart reflector i to user n; denotes the non-line-of-sight link component from smart reflector i to user n. α i,n and denote the path loss exponent and the corresponding Rician factor between smart reflector i and user n, respectively. In this invention, a smart reflector segmentation strategy is designed to uniformly distribute the reflection units of the smart reflector by column to the users it serves, and the number of columns of reflection units of smart reflector i allocated to user n can be obtained as β i,n [t] denotes the association variable between smart reflector i and user n, c denotes the row of the smart reflector array, and f and denote the carrier frequency and the speed of light, respectively, and denote the row spacing and the column spacing of the reflection units, respectively, h k,i,n [t] is denoted as:

[0033]

[0034] Similarly, h i,n [t] can be expressed as:

[0035]

[0036] where θ k,i and denote the vertical and horizontal angles of arrival of UAV k to smart reflector i; and denote the vertical and horizontal angles of departure of smart reflector i to user n. In addition, the phase shift matrix Φ k,i,n [t] is expressed as:

[0037]

[0038] where, φ[t] denotes the source transmit power, φ r,c [t] denotes the reflection power of the smart reflector, denotes the array of the smart reflector. A binary variable β i,n = 1 indicates that user n is served by smart reflector i. Then, the present invention combines the direct channel between UAV k and user n with the reflected channel to obtain the hybrid channel:

[0039]

[0040] where H k,n [t] denotes the line-of-sight channel component. μ k,n [t] denotes the association variable of UAV k and user n. β i,n [t] denotes the association variable of smart reflector i and user n. Since the line-of-sight channel component dominates in the communication scenario of UAVs and changes slowly relative to the non-line-of-sight channel component, the present invention mainly considers the line-of-sight channel component. The UAV hybrid channel is expressed as:

[0041]

[0042] Q k,n [t] denotes the line-of-sight link component of the direct channel between UAV k and user n in t time slot, Q k,i,n [t] denotes the L k,i,n [t] in the line-of-sight link component.

[0043] The following introduces a transmission rate model, which is used in the present application to transmit using non-orthogonal frequency division multiple access technology. Considering that too many users will bring greater transmission interference to the entire system and increase the design difficulty at the receiver, the present application groups users, with each UAV serving two users, and users in different groups using orthogonal frequency bands, and users in the same group using the same frequency band. The present application fixes the decoding order and power allocation of the users served by each UAV, so that the transmission rate of the priority decoding user n served by the UAV k in time slot t can be represented as:

[0044]

[0045] where p n represents the transmission power allocated to user n, δ 2 represents the power of the additive white Gaussian noise. Then the transmission rate of another user m served by the UAV is:

[0046]

[0047] where p m represents the transmission power allocated to user m. Then the sum of the transmission rates of the user group of each UAV in time slot t can be represented as:

[0048] R k [t]=R k,n [t]+R k,m [t]

[0049] The optimization objective of the present application is to maximize the average system total transmission rate. Then, the optimization problem is expressed as:

[0050]

[0051] q represents the position of the UAV, τ represents the time slot, and T represents the number of time slots.

[0052] Constraint C1 defines that the value of the binary variable β i,n [t] is 0 and 1; constraint C2 indicates that each user n is served by and only by one intelligent reflecting surface; constraint C3 indicates that the number of users served by each intelligent reflecting surface cannot exceed the number of its rows; constraint C4 indicates that each UAV uniformly flies at a speed v, and the moving distance in each time slot is fixed; constraint C5 indicates that the value of the line-of-sight link judgment variable μ k,n [t] is binary; constraint C6 limits the value of the phase shift variable φ n,i [t] to be between 0 and 2π.

[0053] Step 2: Transform the average transmission rate maximization problem in step 1.

[0054] Since the problem in Step 1 is an integer non-convex optimization problem, and the variables are highly coupled. The present application develops a phase-shift strategy to obtain the hybrid channel gain between the UAVs and the user. Then the converted optimization problem is modeled as a Markov decision process.

[0055] The present application models the optimization problem as a Markov continuous decision process, specifically, the next time step is determined by the interaction of the current environment. Through training, the agent can learn the optimal strategy ψ, that is, the optimal action A[t] made according to the state S[t] of the current time step.

[0056] First, the state space is defined as The state space of all UAVs, the state space of each UAV is defined as the position coordinates of the UAV at different time slots, denoted as S k [t] = {q k [t]}. Therefore, the state of each time slot t can be represented as s[t] = {q1[t],..., q k [t],..., q K [t]}. The action space of the UAVs, where the action space mainly includes the movement direction indication of each UAV at time slot t and the association variable of the user, that is, m k [t] represents the movement direction indication of UAV k at time slot t. Therefore, the actions of all UAVs in each time slot t together form the action A[t] = (A1[t],..., A k [t],..., A K [t]). According to the optimization objective of maximizing the average transmission rate, the present application sets the reward function to the maximum average transmission rate of this time slot brought by the user association and UAV movement strategy corresponding to the action A[t], that is,

[0057] Step 3: use the improved A3C algorithm to solve the optimization problem in Step 2.

[0058] The pseudo code of the specific training process is shown in Table 1.

[0059]

[0060] First, define the noise network. Specifically, parameterize the fully connected layer of the policy network as a network with added noise. Then the noise linear layer can be represented as:

[0061]

[0062] where, and respectively, instead of the weight matrix and bias. σ, ε represent noise random variables, ω is a weight array, b is a noise bias. Each noisy linear layer has ge+e noise variables. g, e represent the size of the input and output. Since there is no exploratory action selection scheme in A3C as in the classic DRL algorithm such as deep Q network, the selected action is based on the current policy. Therefore, the entropy reward of the policy loss is introduced in the traditional A3C to hinder the update of the deterministic policy. Considering the above problems, the noise-based A3C algorithm adds noise weights in the network, which is equivalent to forming a different current policy. This method is conducive to the exploration of the model. Compared with direct exploration in the policy space, this method can save the artificial entropy loss on the policy. After each step of optimization, the new parameters in the policy network are sampled. Since A3C uses a specified number of steps to return the policy, optimization is performed once every specified number of steps. Since A3C is a policy-based algorithm, the gradient is unbiased when the noise in the network remains consistent throughout the backtracking process.

[0063] In each step of training, first select the noise parameter from the noise random variable. Then, based on the action selected by the current noise-introduced policy, the action is selected and executed to obtain the reward in the specified number of steps. If it is the last specified step, reset the T-step return estimate Otherwise, the return estimate is the agent's estimate of the value function for state S[t]. In the present invention, the thread specified step is set to the number of time slots T. In each step, the return estimate needs to be updated By Where ρ represents the reward discount factor. Then, the gradient policy update is:

[0064]

[0065] Where ζ represents the parameters of the policy network, π(·) represents the action selection policy, V(·) represents the value function, represents the noise, ζ' represents the parameters of the thread policy network, and then the value policy update is:

[0066]

[0067] Where ξ represents the parameters of the value network, and ξ' represents the parameters of the thread value network. Finally, the parameters of the policy network and the value network are updated asynchronously until the end of the set maximum time step, so that the maximum average data transmission rate in each time slot can be obtained.

[0068] Figure 2It can be seen that the convergence speed and convergence value of the proposed MUMR algorithm are optimal, followed by TDQ, and finally FUA. This is because the MUMR algorithm uses an asynchronous update strategy, which can improve the training speed and make the model converge faster. In addition, due to the use of the actor-critic structure in the MUMR algorithm, it can learn the strategy more effectively than the traditional DRL method. In the MUMR algorithm based on noise A3C proposed by the application, the addition of noise reduces the exploration cost of the policy network during updating. In addition, due to the large action space in the Markov decision process converted by the optimization problem of the application, the adaptability of the MUMR algorithm is stronger than that of the traditional DRL method. In addition, it can be observed that FUA has small convergence amplitude and fast convergence. This is because it fixes the user association and is easy to fall into local optimum.

[0069] Figure 3 It can be seen that the average transmission rate increases with the increase of the number of reflection units, and the MUMR algorithm achieves the highest average transmission rate under different numbers of reflection units. In addition, it can be observed that the performance of the FUA algorithm under different numbers of reflection units is the worst, because the MUMR and TDQ algorithms dynamically optimize the user association decision.

[0070] Figure 4 It can be seen that as the transmission power increases, the average transmission rate increases. In addition, it can be observed from the figure that the performance of the MUMR algorithm proposed by the application is relatively stable, while the performance of the TDQ algorithm fluctuates with the change of the transmission power. Specifically, when the transmission power is 25dBm, the performance of TDQ is close to MUMR, and the relative difference is larger when the size is larger.

[0071] Figure 5 It can be seen that as the number of unmanned aerial vehicles increases, the average transmission rate increases. In addition, it can be observed from the figure that the performance of the MUMR algorithm proposed by the application is more stable, followed by TDQ, and the performance of FUA is relatively poor.

[0072] Figure 6 It can be seen that when the number of unmanned aerial vehicles is constant, the average transmission rate decreases obviously with the increase of the number of users. Compared with other algorithms, the MUMR algorithm proposed by the application is more stable, because the MUMR not only optimizes the user association decision, but also considers the dynamic selection of the best channel gain for the transmission between unmanned aerial vehicles and users.

[0073] The above is the specific embodiment of the application and the technical principle used. If changes are made in accordance with the concept of the application, the resulting functional effects still fall within the scope of the application as covered by the specification and drawings.

Claims

1. A method for user association and UAV trajectory optimization based on a multi-intelligent reflector-assisted multi-UAV communication network, characterized in that, Includes the following steps: Step 1: Construct a system model, confirm the transmission channel model and transmission rate model, and construct the optimization problem with maximizing the average total system transmission rate as the optimization objective: K represents the number of drones, N represents the number of users, I represents the number of smart reflectors, and R represents the number of drones. k [t] represents the total transmission rate of each user group of drones in time slot t, q k [t] represents the coordinates of UAV k in time slot t, r represents the number of columns of the intelligent reflective array, q represents the position of the UAV, τ represents the time slot, and T represents the number of time slots; Constraint C1 defines a binary variable β i,n The value of [t] is 0 or 1; constraint C2 indicates that each user n has one and only one intelligent reflector providing services; constraint C3 indicates that the number of users served by each intelligent reflector cannot exceed the number of its rows; constraint C4 indicates that each UAV flies uniformly at speed v, and the distance traveled in each time slot is fixed; constraint C5 indicates the line-of-sight link determination variable μ. k,n The value of [t] is binary; constraint C6 restricts the phase shift variable φ. n,i The value of [t] is between 0 and 2π; The transmission channel model is constructed as follows: The system model includes K UAVs, N users, and I intelligent reflectors. Each intelligent reflector consists of r×c reflective elements, where r represents the number of columns and c represents the number of rows in the intelligent reflector array. Within each time slot t, the communication between the UAVs and users is divided into two channel scenarios based on the propagation environment: a channel with only reflection and a mixed reflection and direct channel. Within each time slot t, multiple UAVs provide services to different users. The coordinates of UAV k in time slot t are q. k [t], where the coordinates of the intelligent reflective surface i and the user n are q respectively. i and q n The distance d between the UAV k and the intelligent reflective surface i within time slot t can be obtained separately. k,i [t] = ||q k [t]-q i ||, the distance d between the drone and the user k,n [t] = ||q k [t]-q n || The distance d between the smart reflective surface and the user i,n [t] = ||q i -q n ||; Within time slot t, the direct connection channel between UAV k and user n is: Where, α k,n and denoted by and representing the path loss exponent and Ricean fading factor between UAV k and user n, respectively; γ represents the channel gain at a one-meter reference point. Indicates the non-line-of-sight component; The reflection link consists of two parts: the first part is the link from the UAV to the intelligent reflective surface, and the second part is the link from the intelligent reflective surface to the user. Considering the obstruction between the UAV and the user, a virtual line-of-sight channel path is established through the intelligent reflective surface. Therefore, in time slot t, the virtual channel line-of-sight distance from UAV k to user n is: L k,i,n [t]=a(H i,n [t]) H Φ i,n [t]H k,i,n [t] Where 'a' represents the amplitude loss caused by the smart reflector; H k,i,n [t] represents the channel vector from UAV k to smart reflector i in time slot t; H i,n [t] represents the channel vector from intelligent reflector i to user n in time slot t, Φ i,n [t] represents the phase shift matrix between user n and intelligent reflector i; within time slot t, the channel vector from UAV k to intelligent reflector i is expressed as: Among them, h k,i,n [t] represents the line-of-sight link component from UAV k to smart reflector i; within time slot t, the channel component from smart reflector i to user n is represented as: Among them, h i,n [t] represents the line-of-sight link component from smart reflector i to user n; α i,n and Let i and n represent the path loss exponent and the corresponding Rice factor, respectively. This represents the non-line-of-sight link component from smart reflector i to user n; Next, the direct channel and the reflection channel between UAV k and user n are combined to obtain the hybrid channel, which is represented as follows: Among them, H k,n [t] represents the direct channel, μ k,n [t] represents the correlation variable between drone k and user n, β i,n [t] represents the correlation variable between the intelligent reflective surface i and the user n; Considering primarily the line-of-sight channel component, the hybrid channel representation for UAVs is as follows: Q k,n [t] represents the line-of-sight link component of the direct channel between UAV k and user n within time slot t; Q k,i,n [t] represents L k,i,n The line-of-sight link component in [t]; The transmission rate model is constructed as follows: Users are grouped, with each drone serving two users. Users in different groups use orthogonal frequency bands, while users in the same group use the same frequency band. The decoding order and power allocation for each drone serving a user are fixed. Therefore, within time slot t, the transmission rate of the priority decoding user n served by drone k is expressed as: Where, p n δ represents the transmit power allocated to user n. 2 Let represent the power of additive white Gaussian noise; then the transmission rate of the other user m served by the drone is: Where, p m Let m represent the transmit power allocated to user m. Then, the total transmission rate of each user group of drones in time slot t can be expressed as: R k [t]=R k,n [t]+R k,m [t] Step 2: For the optimization problem of maximizing the average total system transmission rate in Step 1, firstly, the optimal phase shift strategy is obtained by beam alignment and the problem is transformed to derive the hybrid channel gain between the UAV and the user. Then, the transformed optimization problem is modeled as a Markov decision process. Step 3: Solve the transformed optimization problem in Step 2 using the improved A3C algorithm.

2. The user association and UAV trajectory optimization method based on a multi-intelligent reflector-assisted multi-UAV communication network according to claim 1, characterized in that: A smart reflector segmentation strategy is used to evenly distribute the reflective units of the smart reflector column-wise to the users it serves, thus obtaining the number of columns of reflective units from smart reflector i allocated to user n. β i,n [t] represents the association variable between smart reflector i and user n, c represents the row of the smart reflector array, and f and These represent the carrier frequency and the speed of light, respectively. and h represents the row spacing and column spacing of the reflective elements, respectively. k,i,n [t] is represented as: Similarly, h i,n [t] is represented as: Where, θ k,i and Indicates the vertical and horizontal angles of arrival of the drone k on the intelligent reflective surface i; and This represents the vertical and horizontal departure angles from the intelligent reflective surface i to the user n.

3. The user association and UAV trajectory optimization method based on a multi-intelligent reflector-assisted multi-UAV communication network according to claim 1, characterized in that: The optimization problem utilizes channel alignment to obtain a smart reflector phase control strategy, and then derives the hybrid channel gain between the UAV and the user based on this strategy.

4. The user association and UAV trajectory optimization method based on a multi-intelligent reflector-assisted multi-UAV communication network according to claim 1, characterized in that: The Markov decision process defines the state space. The state space includes all drones, and the state space of each drone is defined as the position coordinates of the drone in different time slots, denoted as S. k [t] = {q k [t]}, the state of each time slot t is represented as s[t]={q1[t],...,q k [t],...,q K [t]}, This represents the drone's action space, which mainly includes the movement direction indication of each drone in time slot t and the user's associated variables, i.e. m k [t] represents the direction of movement of UAV k in time slot t; therefore, the actions of all UAVs in each time slot t together constitute the action A[t] = (A1[t],...,A k [t],...,A K [t]); Based on the optimization objective of maximizing the average transmission rate, the reward function is... Set to the maximum average transmission rate of this time slot resulting from the user association and drone movement strategy corresponding to the execution action A[t].

5. The user association and UAV trajectory optimization method based on a multi-intelligent reflector-assisted multi-UAV communication network according to claim 1, characterized in that: The improved A3C algorithm includes the following steps: First, parameterizing the fully connected layer of the policy network as a network with added noise, the noisy linear layer can be represented as: in, and They respectively replaced the weight matrix and the bias. σ and ε represent noise random variables, ω is the weight array, and b is the noise bias. Each noisy linear layer has ge+e noise variables, where g and e represent the magnitude of the input and output. After each optimization step, the new parameters in the policy network are sampled. Since A3C uses a specified number of steps return strategy, optimization is performed once every specified number of steps. In each training step, a noise parameter is first selected from a noisy random variable; then, within a specified number of steps, an action is selected based on the current policy that introduces noise and executed to obtain a reward; if the current step is the specified last step, the return estimate of the T steps is reset. Otherwise, the returned estimate is the agent's estimate of the value of state S[t] as a function of the state. The thread specifies the step as the number of time slots T, and the returned estimate needs to be updated at each step. pass Where ρ represents the reward discount factor. The reward function is then expressed, and the gradient policy is updated as follows: Where ζ represents the parameters of the policy network, π(·) represents the action selection policy, and V(·) represents the value function. Let A[t] represent noise, S[t] represent the drone's actions, ζ′ represent the state of time slot t, and ζ′ represent the parameters of the thread policy network. Then, the value policy is updated as follows: Here, ξ represents the parameters of the value network; ξ′ represents the parameters of the thread value network. Finally, the parameters of the policy network and the value network are updated asynchronously until the set maximum time step ends, thereby obtaining the maximum average data transmission rate in each time slot.

6. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed, it can implement the user association and UAV trajectory optimization method based on a multi-intelligent reflector-assisted multi-UAV communication network as described in any one of claims 1-5.

Citation Information

Patent Citations

  • Method for maximizing network rate of unmanned aerial vehicle assisted by multiple intelligent reflecting surfaces in emergency scene

    CN117478256A

  • User association and trajectory optimization method based on multi-intelligent reflector unmanned aerial vehicle communication

    CN117793752A