Multi-super-large-scale intelligent reflector network rate optimization method based on machine learning
Through machine learning and optimization algorithms, the flexibility and performance problems of drone-assisted XL-IRS in near-field communication systems are solved, and the communication performance and throughput of the system are improved.
Patent Information
- Application Number
- CN202510605456.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-12
- Publication Date
- 2025-08-12
AI Technical Summary
In the prior art, there are few researches on drone-assisted ultra-large-scale intelligent reflective surfaces (XL-IRSs) in near-field communication systems, and traditional research is mostly based on far-field models, which is difficult to meet the needs of flexibility and communication performance.
The machine learning algorithm is used to combine the Lagrangian multiplication method and the element coordinate block gradient descent algorithm to optimize the base station transmission precoding matrix, XL-IRS phase shift matrix and drone trajectory, and optimize the drone flight trajectory through reinforcement learning algorithms to build a near-field communication system model to maximize the system weighting sum rate.
While ensuring transmission power and user location, the base station and XL-IRS parameters are optimized to improve the communication performance and throughput of the communication system.
Smart Images

Figure CN120474914A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of near-field wireless communication, and in particular relates to a method for optimizing the network rate of multiple ultra-large-scale intelligent reflective surfaces based on machine learning. Background Art
[0002] Intelligent Reflecting Surfaces (IRS) are an emerging wireless communication technology with powerful radio environment engineering capabilities. They achieve cost-effective wireless communication by actively manipulating the radio propagation environment. Compared to traditional relays, IRS-assisted communication eliminates expensive radio frequency (RF) chains and operates in full-duplex mode, eliminating self-interference and noise amplification. However, due to the two-way loss attenuation of IRS reflected signals, achieving significant performance gains requires a sufficiently large physical and electrical size of the IRS, leading to communication scenarios with extra-large-scale IRSs (XL-IRS). Unlike traditional far-field uniform plane wave (UPA) propagation models, XL-IRS communication scenarios use non-uniform spherical waves (NUSW) in the near field. With the continuous advancement of technology, the requirements for IRS deployment in various scenarios are becoming increasingly stringent, requiring IRSs to possess high flexibility. Researchers are focusing on integrating drones with IRSs, leveraging the high maneuverability of drones to meet these requirements. In certain scenarios, drones equipped with IRSs are required to fly within a specific area, which places demands on the drone's flight trajectory. Currently, most of the research on IRS is based on the far-field UPA propagation model, and UAV-assisted IRS communication is also under far-field conditions. There is little research on UAV-assisted XL-IRS communication under near-field conditions.
[0003] To overcome this challenge, a machine learning-based study on rate optimization for multiple, ultra-large-scale smart reflector networks was conducted. This study employed machine learning algorithms to address optimization problems where obtaining large amounts of training data is difficult. Furthermore, the study modeled the near-field communication system in detail and derived closed-form expressions for the base station transmit precoding matrix and the ultra-large-scale smart reflector phase shift matrix. Summary of the Invention
[0004] In summary, the present invention provides a method for optimizing the network rate of multiple ultra-large-scale intelligent reflector networks based on machine learning. The technical solutions adopted are as follows:
[0005] A machine learning-based multi-ultra-large-scale intelligent reflector network rate optimization method comprises an antenna base station, N ultra-large-scale intelligent reflectors, one of which is deployed on a drone, and I single-antenna ground near-field user;
[0006] The optimization method comprises the following steps:
[0007] S1. Obtain scenario layout parameters, including the number of base station antennas, the center point location, and the spacing between two adjacent antennas; the number of transmitting elements of each ultra-large-scale intelligent reflective surface, the spacing between adjacent reflective elements, and the coordinates of the center point. The coordinates of the ultra-large-scale intelligent reflective surface deployed on the drone represent the coordinates of the drone; the coordinates of each ground user, the weighting coefficient of each user, the distance of each drone flight, the total flight time of the drone, the number of time slots for the drone flight, and the range of the drone's flight area. Through research, obtain the number of ultra-large-scale intelligent reflective surfaces deployed in the scene, the number of elements of the ultra-large-scale intelligent reflective surface and their locations, the location of the base station, and the values of the ground near-field users. Then, use the near-field channel theory to model the channel.
[0008] S2, confirming the optimization problem based on the system weighted sum rate; the near-field communication user connection is an indirect line-of-sight link assisted by a very large-scale intelligent reflecting surface; the base station adjusts the reflection matrix of the ground user and the very large-scale intelligent reflecting surface through the controller of the very large-scale intelligent reflecting surface. After receiving the reflected signal of the target very large-scale intelligent reflecting surface, the base station determines the existence of the LOS link between the base station and the target very large-scale intelligent reflecting surface based on the distance, direction and received signal strength information, and then determines the channel state information of the communication system;
[0009] S21, construct a communication model when the ground user connection is an indirect line-of-sight link,
[0010] The link passes through the nodes from the base station to the ultra-large-scale intelligent reflector to the ground user. At this time, the signal received by the i-th ground user in the t-th time slot is:
[0011]
[0012] Where H i represents the channel between the base station and the i-th user, X is the base station's transmitted signal, is the additive Gaussian white noise received by the i-th ground user, G=[H b1 ,H b2 ,...H bN ], the channel matrix from the base station to the nth ultra-large-scale intelligent reflection surface to the i-th user. Assuming that the Nth ultra-large-scale intelligent reflection surface is deployed on a drone, the coordinates of the drone in the tth time slot are expressed as l(t)=[x(t),y(t),z(t)],t∈[1.T],H bn represents the channel matrix between the base station and the nth ultra-large-scale intelligent reflection surface, represents the beamforming matrix of the nth ultra-large-scale smart reflector, the subscript s∈[1, S] represents the sth element of the nth ultra-large-scale smart reflector, and all ultra-large-scale smart reflectors in the communication system have S elements; H ni represents the channel between the nth ultra-large-scale intelligent reflector and the i-th ground user. The signal transmitted by the base station is expressed as follows:
[0013]
[0014] Where w i is the transmit precoding matrix of the i-th terrestrial user, s i represents the signal of the i-th terrestrial user,
[0015] This constructs the optimization problem in this case:
[0016]
[0017]
[0018] l(t)∈Area, t∈[1, t] (3e);
[0019] S3, adopt the corresponding optimization solution for the optimization problem to obtain the optimized maximum weighted sum rate of the system:
[0020] S31, the optimization steps for S21 are: taking the maximum system weighted sum rate as the goal, constructing the optimization problem while ensuring the base station power constraint, the phase shift matrix unit mode constraint, and the UAV flyable area constraint. The formula is as follows:
[0021]
[0022] l(t)∈Area,t∈[1,t] (4e)
[0023] Where, and are the STAR-RIS amplitude and phase constraints;
[0024] The rate expression of the i-th ground user is:
[0025]
[0026] In order to reduce the decoding complexity, we first consider the linear decoding matrix using the corresponding relationship between the weighted sum rate and the weighted mean square error. The mean square error of the i-th user can be written as:
[0027]
[0028] U iDenotes the decoding coefficient of the i-th terrestrial user, and introduces auxiliary variables Q = [Q1, Q2, .... Q I ]φ0, the weighted sum rate optimization problem of the system can be expressed as:
[0029]
[0030] l(t)∈Area,t∈[1,t] (7e)
[0031] It appears that the optimization problem is a concave function with respect to the auxiliary variable Q and the decoding matrix U,
[0032] When fixed w i , l(t), θ n , Q, by calculating the first-order derivative of the new objective function, the optimal decoding coefficient is obtained, which is expressed as:
[0033]
[0034] Similarly, when w is fixed i , l(t), θ n , U is obtained by calculating the first-order derivative of the new objective function, which is expressed as:
[0035]
[0036] In S32, the transmit precoding matrix is solved by fixing the phase shift matrix, the UAV position, the decoding coefficients, and the auxiliary variables. The optimization problem can be reformulated as:
[0037]
[0038] It is easy to verify that the optimization problem is a convex optimization problem. The Lagrange multiplier method is used to find the corresponding solution. According to the Lagrange multiplier method, a function related to the transmit precoding matrix and the Lagrange multiplier can be expressed as:
[0039]
[0040] Where λ≥0, the optimal transmit precoding matrix can be expressed as:
[0041] w i (λ)=ω i (A+λΞ M )H i U i Q i (12)
[0042] The optimal transmit precoding matrix is solved using the bisection method;
[0043] S33, when solving the optimal phase shift matrix, the transmit precoding matrix, the UAV position, the decoding coefficients and the auxiliary variables are fixed to solve the optimal phase shift matrix. The optimization problem can be reformulated as:
[0044]
[0045] Where Z is represented by:
[0046]
[0047] q can be expressed as:
[0048]
[0049] We use the element-wise fast gradient descent algorithm to solve it. At each step, we treat one entry of θ as a block and optimize that block while fixing the other blocks. The constraints for each item are separate; Observe the function Relative to It is actually a quadratic function, which is of the form where μ is a real number, κ is a complex number, and Unrelated, the question can be restated as:
[0050]
[0051]
[0052] The closed form expression of the problem is solved as:
[0053]
[0054] To calculate κ, the complex gradient is calculated from two different perspectives. On the one hand, The complex gradient with respect to θ can be calculated as:
[0055]
[0056] because yes The quadratic function of right The complex derivative of can be calculated as:
[0057]
[0058] The complex gradient is unique, and the update expression for κ can be expressed as:
[0059]
[0060] S34, when solving the optimal UAV trajectory, the transmit precoding matrix, phase shift matrix, decoding coefficients and auxiliary variables are fixed to solve the optimal UAV trajectory. The optimization problem can be reformulated as:
[0061]
[0062] st l(t)∈Area, t∈[1, t] (23b)
[0063] The goal is to obtain the best movement to maximize the system weighted sum rate of the next drone position; since the optimization problem is difficult to handle, an effective algorithm is needed to solve the horizontal trajectory optimization problem; since the system weighted sum rate is related to the distance between the drone and all users and the distance to the base station, an algorithm based on reinforcement learning is proposed to solve the optimization problem; in the Q-learning model, the drone acts as an intelligent agent. The Q-learning model consists of four elements: state space S, action space A, reward function Ra and q table. In each step of the Q-learning network, the agent explores the environment from the initial state, calculates the reward of the selected action, and gradually updates the q table. The reward function Ra depends on the current state s and the selected action a. According to the optimization problem in (7a), Ra can be modeled as:
[0064]
[0065] The state update can be expressed as:
[0066]
[0067] S35, return the phase shift matrix θ obtained in S33 to step S32 to recalculate the transmit precoding matrix W, and repeat this cycle until the final optimization target is stable or the number of alternating optimizations reaches a limited number, then return the drone position l(t) obtained in S34 to S31 to recalculate the decoding coefficients and auxiliary variables, and repeat returning the phase shift matrix θ obtained in S33 to step S32 to recalculate the transmit precoding matrix W until the final optimization target is stable or the number of alternating optimizations reaches a limited number, and repeat this cycle until the final optimization target is stable or the number of alternating optimizations reaches a limited number.
[0068] Beneficial effects:
[0069] This method maximizes the system's co-weighted sum rate by optimizing the base station's transmit precoding matrix, the XL-IRS phase shift matrix, and the drone's trajectory, given the exact transmit power and user location. This improves the system's communication performance and throughput, and has high application value. BRIEF DESCRIPTION OF THE DRAWINGS
[0070] Figure 1 A system model diagram for the implementation of the present invention;
[0071] Figure 2 Flowchart of the present invention. DETAILED DESCRIPTION
[0072] In order to more clearly illustrate the technical methods used in the examples of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings.
[0073] Machine learning-based multi-super-large-scale intelligent reflector network rate optimization method:
[0074] S1. Obtain scene layout parameters through investigation, including the number and location of elements of the ultra-large-scale intelligent retro-reflective surface, the location of the base station, the number of base station antennas, the user location, the spacing between adjacent base station antennas, and the spacing between adjacent elements of the ultra-large-scale intelligent retro-reflective surface.
[0075] like Figure 1 As shown, in a near-field wireless communication network, there is a single base station, multiple ground near-field users, and multiple ultra-large-scale intelligent reflecting surfaces. The number of base station antennas is M, the number of ultra-large-scale intelligent reflecting surfaces is N, one of which is deployed on a drone and the number of reflecting elements is S, and the number of single-antenna ground near-field users is I.
[0076] Considering that in a near-field communication system, the base station altitude, drone altitude, XL-IRS deployment altitude, and ground user altitude all have a certain impact on the system's weighted sum rate. In the drone-assisted, multi-ultra-large-scale intelligent reflector-assisted near-field wireless communication network configured in the present invention, the user altitude, base station altitude, and drone altitude are all set to typical heights in practical applications. This invention optimizes the phase parameters of the ultra-large-scale intelligent reflectors serving the users, the base station's transmit precoding matrix, and the drone's trajectory to maximize the system's weighted sum rate within given constraints.
[0077] In order to facilitate the distinction, the present invention sets the ultra-large-scale intelligent reflective surface deployed on the UAV as the Nth one.
[0078] The present invention denotes the channel parameter between the base station and the nth ultra-large-scale intelligent reflector as H bn , which includes path loss and channel fading, and n∈{1, 2, ..., N}. represents the beamforming matrix of the nth ultra-large-scale smart reflector. ni represents the channel between the nth ultra-large-scale smart reflector and the i-th ground user.
[0079] S2, the optimization problem can be determined based on the weighted sum rate of ground users; the base station can adjust the base station's transmit precoding matrix through the base station controller, the ultra-large-scale intelligent reflector can control its own phase shift matrix, and the drone can adjust its own flight trajectory through its own control platform.
[0080] S21, constructing a communication model when the ground user connection is an indirect line-of-sight link, the specific steps include:
[0081] The link passes through the nodes from the base station to the ultra-large-scale intelligent reflector to the ground user. At this time, the signal received by the i-th ground user in the t-th time slot is:
[0082]
[0083] Where H i represents the channel between the base station and the i-th user, X is the base station's transmitted signal, is the additive Gaussian white noise received by the i-th ground user, G=[H b1 ,H b2 ,...H bN ], the channel matrix from the base station to the nth ultra-large-scale intelligent reflection surface to the i-th user. Assuming that the Nth ultra-large-scale intelligent reflection surface is deployed on a drone, the coordinates of the drone in the tth time slot are expressed as l(t)=[x(t),y(t),z(t)],t∈[1.T],H bn represents the channel matrix between the base station and the nth ultra-large-scale intelligent reflection surface, represents the beamforming matrix of the nth ultra-large-scale smart reflector, the subscript s∈[1, S] represents the sth element of the nth ultra-large-scale smart reflector, and all ultra-large-scale smart reflectors in the communication system have S elements; H ni represents the channel between the nth ultra-large-scale intelligent reflector and the i-th ground user. The signal transmitted by the base station is expressed as follows:
[0084]
[0085] Where w i is the transmit precoding matrix of the i-th terrestrial user, s i represents the signal of the i-th terrestrial user.
[0086] This constructs the optimization problem in this case:
[0087]
[0088]
[0089] l(t)∈Area, t∈[1, t] (3e);
[0090] The specific steps of S3 are:
[0091] S31, the optimization steps for S21 are: taking the maximum system weighted sum rate as the goal, constructing the optimization problem while ensuring the base station power constraint, the phase shift matrix unit mode constraint, and the UAV flyable area constraint; the formula is as follows:
[0092]
[0093] l(t)∈Area,t∈[1,t] (4e)
[0094] Where, and are the STAR-RIS amplitude and phase constraints;
[0095] The rate expression of the i-th ground user is:
[0096]
[0097] In order to reduce the decoding complexity, we first consider the linear decoding matrix that uses the corresponding relationship between the weighted sum rate and the weighted mean square error. Then the mean square error of the i-th user can be written. It is expressed as:
[0098]
[0099] U i Denotes the decoding coefficient of the i-th terrestrial user, and introduces auxiliary variables Q = [Q1, Q2, .... Q I ]φ0, the weighted sum rate optimization problem of the system can be expressed as:
[0100]
[0101] l(t)∈Area,t∈[1,t] (7e)
[0102] At this point, the optimization problem appears to be a concave function of the auxiliary variable Q and the decoding matrix U. When w is fixed i , l(t), θ n , Q, by calculating the first-order derivative of the new objective function, the optimal decoding coefficient is obtained, which is expressed as:
[0103]
[0104] Similarly, when w is fixed i , l(t), θ n , U can be obtained by calculating the first-order derivative of the new objective function, which is expressed as:
[0105]
[0106] S32, when the decoding coefficients and auxiliary variables are fixed, the optimal transmit precoding matrix, phase shift matrix, and drone trajectory can be solved using the alternating optimization algorithm. The transmit precoding matrix is solved by fixing the phase shift matrix, drone position, decoding coefficients, and auxiliary variables. The optimization problem can be reformulated as:
[0107]
[0108] Since it can be easily verified that the optimization problem is a convex optimization problem, the Lagrange multiplier method is used to find the corresponding solution. According to the Lagrange multiplier method, a function related to the transmit precoding matrix and the Lagrange multiplier can be constructed and expressed as:
[0109]
[0110] Where λ≥0, from its first-order optimal condition, the optimal transmit precoding matrix is expressed as:
[0111] w i (λ)=ω i (A+λΞ M )H i U i Q i (12)
[0112] For the optimal transmit precoding matrix, we can use the bisection method to solve;
[0113] S33, since the transmit precoding matrix is solved by fixing the phase shift matrix, the UAV position, the decoding coefficients and the auxiliary variables, when solving the optimal phase shift matrix, the transmit precoding matrix, the UAV position, the decoding coefficients and the auxiliary variables are fixed to solve the optimal phase shift matrix, the optimization problem is reformulated as:
[0114]
[0115] Where Z is represented by:
[0116]
[0117] q is expressed as:
[0118]
[0119] We use element-wise fast gradient descent to solve it. At each step, we treat one entry of θ as a block and optimize that block while fixing the other blocks. The constraints for each item are separate. Now that we have fixed It can be observed that the function Relative to It is actually a quadratic function, which is of the form where μ is a real number, κ is a complex number, and Unrelated, the question can be restated as:
[0120]
[0121] The closed form expression of the problem is solved as:
[0122]
[0123] To calculate κ, the complex gradient is calculated from two different perspectives;
[0124] on the one hand, The complex gradient with respect to θ is calculated as:
[0125]
[0126] In addition, due to yes The quadratic function of right The complex derivative of is calculated as:
[0127]
[0128] Since the complex gradient is unique, the update expression for κ can be expressed as:
[0129]
[0130] S34, since the optimal transmit precoding matrix is solved by fixing the phase shift matrix, the UAV position, the decoding coefficients and the auxiliary variables, when solving the optimal UAV trajectory, the transmit precoding matrix, the phase shift matrix, the decoding coefficients and the auxiliary variables are fixed to solve the optimal UAV trajectory. The optimization problem can be reformulated as:
[0131]
[0132] s.tl(t)∈Area,t∈[1,t] (23b)
[0133] The goal is to obtain the best motion to maximize the system weighted sum rate of the next drone position; since the optimization problem is difficult to handle, an effective algorithm is needed to solve the horizontal trajectory optimization problem; since the system weighted sum rate is related to the distance between the drone and all users and the distance to the base station, an algorithm based on reinforcement learning is proposed to solve the optimization problem; in the Q-learning model, the drone acts as an intelligent agent. The Q-learning model consists of four elements: state space S, action space A, reward function Ra and q table; in each step of the Q-learning network, the agent explores the environment from the initial state, calculates the reward of the selected action, and gradually updates the q table; the reward function Ra depends on the current state s and the selected action a. According to the optimization problem in (7a), Ra is modeled as:
[0134]
[0135] The status update is represented as:
[0136]
[0137] S35, first return the phase shift matrix θ obtained in S33 to step S32 to recalculate the transmit precoding matrix W, and repeat this cycle until the final optimization target is stable or the number of alternating optimizations reaches a limited number, then return the drone position l(t) obtained in S34 to S31 to recalculate the decoding coefficients and auxiliary variables, and repeatedly return the phase shift matrix θ obtained in S33 to step S32 to recalculate the transmit precoding matrix W until the final optimization target is stable or the number of alternating optimizations reaches a limited number, and repeat this cycle until the final optimization target is stable or the number of alternating optimizations reaches a limited number.
[0138] The above embodiments are merely illustrative of the present invention. Those skilled in the art may make various modifications and variations to the present invention without departing from the spirit and scope of the present invention. Therefore, the present invention is intended to include such modifications and variations as long as they fall within the scope of the claims and their equivalents.
Claims
1. A machine learning-based multi-ultra-large-scale intelligent reflector network rate optimization method, characterized by: It consists of a base station with M antennas, N ultra-large-scale intelligent reflectors, one of which is deployed on a drone, and I single-antenna ground near-field user. The optimization method comprises the following steps: S1. Obtain scenario layout parameters, including the number of base station antennas, the center point location, and the spacing between two adjacent antennas; the number of transmitting elements of each ultra-large-scale intelligent reflective surface, the spacing between adjacent reflective elements, and the coordinates of the center point. The coordinates of the ultra-large-scale intelligent reflective surface deployed on the drone represent the coordinates of the drone; the coordinates of each ground user, the weighting coefficient of each user, the distance of each drone flight, the total flight time of the drone, the number of time slots for the drone flight, and the range of the drone's flight area. Through research, obtain the number of ultra-large-scale intelligent reflective surfaces deployed in the scene, the number of elements of the ultra-large-scale intelligent reflective surface, and their locations, the location of the base station, and the values of the ground near-field users. Then, use near-field channel theory to model the channel. S2, confirming the optimization problem based on the system weighted sum rate; the near-field communication user connection is an indirect line-of-sight link assisted by a very large-scale intelligent reflecting surface; the base station adjusts the reflection matrix of the ground user and the very large-scale intelligent reflecting surface through the controller of the very large-scale intelligent reflecting surface. After receiving the reflected signal of the target very large-scale intelligent reflecting surface, the base station determines the existence of the LOS link between the base station and the target very large-scale intelligent reflecting surface based on the distance, direction and received signal strength information, and then determines the channel state information of the communication system; S21, constructing a communication model when the ground user connection is an indirect line-of-sight link, the specific steps include: The link passes through the nodes from the base station to the ultra-large-scale intelligent reflector to the ground user. At this time, the signal received by the i-th ground user in the t-th time slot is: Where H i represents the channel between the base station and the i-th user, X is the base station's transmitted signal, is the additive Gaussian white noise received by the i-th ground user, G=[H b1 , H b2 ,...H bN ], the channel matrix from the base station to the nth ultra-large-scale intelligent reflection surface to the i-th user. Assuming that the Nth ultra-large-scale intelligent reflection surface is deployed on a drone, the coordinates of the drone in the t-th time slot are expressed as l(t)) = [x(t), y(t), z(t)], t∈[1.T], H bn represents the channel matrix between the base station and the nth ultra-large-scale intelligent reflection surface, represents the beamforming matrix of the nth ultra-large-scale smart reflector, the subscript s∈[1, S] represents the sth element of the nth ultra-large-scale smart reflector, and all ultra-large-scale smart reflectors in the communication system have S elements; H ni represents the channel between the nth ultra-large-scale intelligent reflector and the i-th ground user. The signal transmitted by the base station is expressed as follows: Where w i is the transmit precoding matrix of the i-th terrestrial user, s i represents the signal of the i-th terrestrial user, This constructs the optimization problem in this case: l(t)∈Area, t∈[1, t](3e); S3, adopt the corresponding optimization scheme for the optimization problem to obtain the optimized maximum weighted sum rate of the system; S31, the optimization steps for S21 are: taking the maximum system weighted sum rate as the goal, constructing the optimization problem while ensuring the base station power constraint, the phase shift matrix unit mode constraint, and the UAV flyable area constraint; the formula is as follows: l(t)∈Area,t∈[1,t] (4e) Where, and are the STAR-RIS amplitude and phase constraints; The rate expression of the i-th ground user is: In order to reduce the decoding complexity, we first consider the linear decoding matrix using the corresponding relationship between the weighted sum rate and the weighted mean square error. Then, the mean square error of the i-th user can be written as: U i Denotes the decoding coefficient of the i-th terrestrial user, and introduces the auxiliary variable Q = [Q1, Q 2, ....Q I ]φ0, the weighted sum rate optimization problem of the system can be expressed as: l(t)∈ Area, t∈[1, t](7e) Now the optimization problem appears to be a concave function of the auxiliary variable Q and the decoding matrix U. When w is fixed i , l(t), θ n , Q, by calculating the first-order derivative of the new objective function, the optimal decoding coefficient is obtained, which is expressed as: Similarly, fixed w i , l(t), θ n , U, by calculating the first-order derivative of the new objective function, the optimal auxiliary variable is obtained, which is expressed as: In step S32, when the decoding coefficients and auxiliary variables are fixed, the optimal transmit precoding matrix, phase shift matrix, and UAV trajectory are solved using the alternating optimization algorithm. The transmit precoding matrix is solved by fixing the phase shift matrix, UAV position, decoding coefficients, and auxiliary variables. The optimization problem can be reformulated as: It is easy to verify that the optimization problem is a convex optimization problem, so the Lagrange multiplier method is used to find the corresponding solution. According to the Lagrange multiplier method, a function related to the transmit precoding matrix and the Lagrange multiplier can be constructed and expressed as: Where λ≥0, the optimal transmit precoding matrix can be expressed as: w i (λ)=ω i (A+λΞ M )H i U i Q i (12) The optimal transmit precoding matrix is solved using the bisection method; S33, by fixing the phase shift matrix, UAV position, decoding coefficients and auxiliary variables to solve the transmit precoding matrix. When solving the optimal phase shift matrix, the transmit precoding matrix, UAV position, decoding coefficients and auxiliary variables are fixed to solve the optimal phase shift matrix. The optimization problem is reformulated as: Where Z is represented by: q is expressed as: We use element-wise fast gradient descent to solve it. At each step, we treat one entry of θ as a block and optimize that block while fixing the other blocks. The constraints for each item are separate; since it has been fixed It can be observed that the function Relative to is a quadratic function of the form where μ is a real number, κ is a complex number, and Unrelated, the question can be restated as: The closed form expression of the problem is solved as: To calculate κ, the complex gradient is calculated from two different perspectives; on the one hand, The complex gradient with respect to θ is calculated as: In addition, due to yes The quadratic function of right The complex derivative of is calculated as: The complex gradient is unique, and the update expression for κ is expressed as: S34, when solving the optimal UAV trajectory, the transmit precoding matrix, phase shift matrix, decoding coefficients and auxiliary variables are fixed to solve the optimal UAV trajectory, and the optimization problem is reformulated as: s.tl(t)∈Area,t∈[1,t] (23b) The goal is to obtain the best movement to maximize the system weighted sum rate of the next drone position. Since the system weighted sum rate is related to the distance between the drone and all users and the distance to the base station, an algorithm based on reinforcement learning is proposed to solve the optimization problem. In the Q-learning model, the drone acts as an intelligent agent. The Q-learning model consists of four elements: state space S, action space A, reward function Ra and q table. In each step of the Q-learning network, the agent explores the environment from the initial state, calculates the reward of the selected action, and gradually updates the q table. The reward function Ra depends on the current state s and the selected action a. According to the optimization problem in (7a), Ra is modeled as: The status update is represented as: S35, return the phase shift matrix θ obtained in S33 to step S32 to recalculate the transmit precoding matrix W, and repeat this cycle until the final optimization target is stable or the number of alternating optimizations reaches a limited number, then return the drone position l(t) obtained in S34 to S31 to recalculate the decoding coefficients and auxiliary variables, and repeat returning the phase shift matrix θ obtained in S33 to step S32 to recalculate the transmit precoding matrix W until the final optimization target is stable or the number of alternating optimizations reaches a limited number, and repeat this cycle until the final optimization target is stable or the number of alternating optimizations reaches a limited number.