Multi-objective optimization method for high-speed dynamic communication based on unmanned aerial vehicle cooperative beamforming

By using a multi-objective optimization model for cooperative beamforming of UAV virtual antenna arrays, combined with long short-term memory networks and reinforcement learning algorithms, the weights of UAV position and excitation current are optimized, solving the energy consumption and channel interference problems of UAV cooperative beamforming in mobile user communication, and realizing high-speed dynamic communication.

CN119071815BActive Publication Date: 2025-10-24JILIN UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411022416.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-29
Publication Date
2025-10-24
Estimated Expiration
2044-07-29

AI Technical Summary

Technical Problem

Unmanned aerial vehicle (UAV) collaborative beamforming cannot meet the high-speed dynamic communication requirements when serving mobile users, and it also suffers from high energy consumption and severe channel interference.

Method used

By establishing a multi-objective optimization model for collaborative beamforming of UAV virtual antenna arrays, combined with long short-term memory networks and reinforcement learning algorithms, the UAV's moving position and excitation current weight are optimized in real time to improve the transmission rate and reduce energy consumption.

Benefits of technology

It enables high-speed dynamic communication in wireless networks, improves the transmission rate for mobile users, reduces the energy consumption of drones, and enhances the real-time response capability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119071815B_ABST
    Figure CN119071815B_ABST
Patent Text Reader

Abstract

The application discloses a high-speed dynamic communication multi-target optimization method based on unmanned aerial vehicle cooperative beam forming, which comprises the following steps: step one, establishing a calculation model of the sum of transmission rates of the virtual antenna array of the unmanned aerial vehicle in the whole communication time as a first target function, establishing a calculation model of the total energy consumption of the unmanned aerial vehicle as a second target function, and obtaining the position of the mobile user in the communication system in each time slot; step two, taking the maximum of the first target function and the minimum of the second target function as the optimization target, and determining the optimal position of the unmanned aerial vehicle and the optimal excitation current weight of the unmanned aerial vehicle; and step three, the unmanned aerial vehicle moves to the optimal position, and transmits communication data to the user according to the optimal excitation current weight. The high-speed dynamic communication multi-target optimization method based on unmanned aerial vehicle cooperative beam forming can output the optimal moving position of the unmanned aerial vehicle and the optimal excitation current weight according to the position of the mobile user in real time, and realizes high-speed dynamic communication of wireless network communication.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of wireless network communication, and particularly relates to a high-speed dynamic communication multi-objective optimization method based on unmanned aerial vehicle cooperative beamforming. BACKGROUND

[0002] In the field of wireless networks, unmanned aerial vehicles can assist ground networks as air base stations and air relays to provide reliable and efficient wireless communication services for ground users. In the scenario where unmanned aerial vehicles serve as air base stations, unmanned aerial vehicles located in a certain area need to communicate with remote mobile users. The communication between the unmanned aerial vehicles and the remote mobile users can be established through a swarm of unmanned aerial vehicles, and then a single-hop or multi-hop method is used to realize communication. However, this method may cause link failure and additional energy consumption, and at the same time, due to time-varying channel conditions and interference from other base stations, it is difficult for the remote mobile users to obtain satisfactory communication rates. In addition, the current work on unmanned aerial vehicle cooperative beamforming only considers unmanned aerial vehicles serving static ground terminals, and ignores the mobility of users, which cannot obtain satisfactory results when serving mobile users. Due to the high dynamic nature caused by the random movement of users, the real-time response capability of the communication system is challenged. SUMMARY

[0003] The purpose of the application is to provide a high-speed dynamic communication multi-objective optimization method based on unmanned aerial vehicle cooperative beamforming, which can output the optimal mobile position and optimal excitation current weight of the unmanned aerial vehicle in real time according to the position of the mobile user by establishing a multi-objective joint optimization model for improving the transmission rate of the receiving mobile user in the cooperative beamforming of the unmanned aerial vehicle virtual antenna array and reducing the energy consumption of the unmanned aerial vehicle movement, and realizing high-speed dynamic communication of wireless network communication.

[0004] The technical scheme provided by the application is as follows:

[0005] The high-speed dynamic communication multi-objective optimization method based on unmanned aerial vehicle cooperative beamforming comprises the following steps:

[0006] Step one, establishing a calculation model of the sum of the transmission rates of the unmanned aerial vehicle virtual antenna array within the entire communication time as a first objective function, and establishing a calculation model of the total energy consumption of the unmanned aerial vehicle as a second objective function; and

[0007] In each time slot, the position of the mobile user in the communication system is obtained.

[0008] Step two, taking the maximum of the first objective function and the minimum of the second objective function as the optimization target, determining the optimal position of the unmanned aerial vehicle and the optimal excitation current weight of the unmanned aerial vehicle.

[0009] Step three, the UAV moves to the optimal position, and transmits communication data to the user according to the optimal excitation current weight.

[0010] Preferably, the calculation model of the sum of the transmission rates of the UAV virtual antenna array in the entire communication time is:

[0011]

[0012] wherein, represents the horizontal movement direction of the UAV, represents the horizontal movement distance of the UAV, represents the vertical movement distance of the UAV, represents the excitation current weight of the UAV; R UM [t] represents the transmission rate of the UAV virtual antenna array in the t time slot when communicating with the mobile user; Y UM [t] represents the signal-to-interference-plus-noise ratio of the UAV virtual antenna array in the t time slot when communicating with the mobile user.

[0013] Preferably, the signal-to-interference-plus-noise ratio of the UAV virtual antenna array in the t time slot when communicating with the mobile user is calculated by the following formula:

[0014]

[0015] wherein, P U [t] represents the total transmission power of the UAV virtual antenna array in the t time slot, G UM [t] represents the antenna gain of the UAV virtual antenna array in the t time slot towards the mobile user position, g UM [t] represents the channel power gain between the UAV virtual antenna array in the t time slot and the mobile user, σ 2 represents the noise power, P B [t] represents the transmission power of the base station in the t time slot, g BM [t] represents the channel power gain between the uncorrelated base station in the t time slot and the mobile user.

[0016] Preferably, the calculation model of the total energy consumption of the UAV is:

[0017]

[0018] wherein, E i [t] represents the energy consumption of the UAV i in the t time slot, T represents the end time of flight, v i [t] represents the speed of the UAV i in the t time slot, m UAV represents the mass of the UAV, g represents the acceleration of gravity, h i [0] and v i [0] respectively represent the initial height and initial speed of the UAV i, hi [T] and v i [T] represent the height and speed of the unmanned aerial vehicle i at time T, respectively, P(v i [t]) represents the propulsion energy consumption of the unmanned aerial vehicle in a two-dimensional horizontal space within time slot t; N UAV represents the number of unmanned aerial vehicles.

[0019] Preferably, the propulsion energy consumption of the unmanned aerial vehicle in a two-dimensional horizontal space is calculated by the following formula:

[0020]

[0021] wherein P B and P I are constants, v tip , v0, d0, s, p and A represent the rotor tip speed, hover average rotor induced velocity, fuselage drag ratio, rotor solidity, air density and rotor disc area, respectively; v i represents the speed of the unmanned aerial vehicle i.

[0022] Preferably, in the step two, the optimal position of the unmanned aerial vehicle and the optimal excitation current weight of the unmanned aerial vehicle are determined based on the long short-term memory network, comprising the following steps:

[0023] Step 1, randomly initializing network parameters of a plurality of long short-term memory networks and weight vectors of the first objective function and the second objective function, obtaining a plurality of policy networks and weight vectors corresponding to the policy networks as an initialization population; performing iterative updating on the policy networks in the initialization population based on a reinforcement learning algorithm of the long short-term memory network, and all updated populations and the initialization population constitute a primary population;

[0024] Step 2, screening a plurality of policy networks from the primary population by using a performance buffer method, obtaining a first population;

[0025] Step 3, screening the first population by using a hyperplane method, obtaining a policy network corresponding to each weight vector uniquely as a second population;

[0026] Step 4, updating the policy networks in the second population based on the reinforcement learning method of the long short-term memory network, and screening non-dominated policy networks in the updating result according to the first objective function and the second objective function, and a set of the non-dominated networks constitutes a Pareto solution set;

[0027] Step 4 is repeated until a maximum number of iterations is reached, and an optimal Pareto solution set is obtained.

[0028] Step 5, taking the user's position as an input parameter, using the strategy network in the optimal Pareto solution set to output the optimal position of the unmanned aerial vehicle and the optimal excitation current weight of the unmanned aerial vehicle.

[0029] Preferably, the first objective function and the second objective function weight vector are arrays (a, b) ;

[0030] Wherein, the value conditions of a and b satisfy a [0, 1], b [0, 1], and a+b=1.

[0031] The beneficial effects of the present application are:

[0032] The high-speed dynamic communication multi-objective optimization method based on unmanned aerial vehicle cooperative beam forming provided by the present application can output the optimal moving position of the unmanned aerial vehicle and the optimal excitation current weight in real time according to the position of the mobile user by establishing a multi-objective joint optimization model for improving the transmission rate of the receiving mobile user and reducing the energy consumption of the unmanned aerial vehicle in the cooperative beam forming of the unmanned aerial vehicle virtual antenna array, and realizes high-speed dynamic communication of wireless network communication.

[0033] The present application combines reinforcement learning algorithm and evolutionary learning method to solve multi-objective optimization problem, combines long short-term memory network to strengthen the learning ability of the algorithm to high dynamic environment, and uses performance buffer method and task selection method based on hyperplane to strengthen the diversity of the optimal Pareto solution set obtained, so that the decision maker can select the suitable strategy network according to different demand preferences. BRIEF DESCRIPTION OF DRAWINGS

[0034] Figure 1 It is a schematic diagram of the cooperative beam dynamic communication scene of the unmanned aerial vehicle virtual antenna array.

[0035] Figure 2 It is a flowchart of the high-speed dynamic communication method of the cooperative beam forming of the unmanned aerial vehicle virtual antenna array.

[0036] Figure 3 It is a schematic diagram of the performance buffer pool method.

[0037] Figure 4 It is a schematic diagram of the strategy selection method based on hyperplane. DETAILED DESCRIPTION

[0038] The present application will be further described in detail below with reference to the accompanying drawings, so that those skilled in the art can implement it according to the description.

[0039] As Figure 1As shown in the wireless network communication process, due to the limitation of on-board energy and transmission power, a single UAV can only serve a limited area; These constraints make it difficult for UAVs to meet the communication needs of remote users. In addition, time-varying channels and interference from other base stations make it difficult to achieve satisfactory communication rates. In addition, current work on UAV cooperative beamforming only focuses on UAVs serving static ground terminals, without considering the mobility of users served (such as pedestrians, robots or vehicles), which poses challenges to the real-time response capability of the communication system.

[0040] As Figure 2 shown, the present application provides a high-speed dynamic communication multi-objective optimization method based on UAV cooperative beamforming, and establishes a multi-objective joint optimization model to improve the transmission rate of the receiving mobile user and reduce the energy consumption of the UAV. Then, the improved multi-objective evolutionary reinforcement learning algorithm is used to design the best position of the UAV movement and the optimal excitation current weight of the UAV according to the observed environment state at each time. The evolutionary reinforcement learning algorithm improves the diversity and effectiveness of the solution by introducing a hyperplane-based task selection method and a long short-term memory network. The specific implementation process is as follows.

[0041] I. According to the rate and energy consumption requirements of long-term dynamic communication, two optimization functions are designed, and a multi-objective optimization problem is constructed according to the multi-objective optimization theory;

[0042] First, the total objective function is designed as follows:

[0043] maxF={f1,f2} (1)

[0044]

[0045] Where,

[0046] R UM [t]=log2(1+Υ UM [t]) (4)

[0047]

[0048] g UM [t]=K0d UM [t] -α Ω UM [t] (6)

[0049] g BM [t]=K0d BM [t] -α Ω BM [t] (7)

[0050]

[0051]

[0052] wherein, denotes the horizontal moving direction of the UAV, denotes the horizontal moving distance of the UAV, denotes the vertical moving distance of the UAV, denotes the exciting current weight of the UAV.

[0053] The first objective function f1 denotes the sum of the transmission rates of the UAV virtual antenna array in the whole communication time. Wherein R UM [t] denotes the transmission rate of the UAV virtual antenna array when communicating with the mobile user at time slot t, Y UM [t] denotes the signal-to-interference-plus-noise ratio of the UAV virtual antenna array when communicating with the mobile user at time slot t, P U [t] denotes the total transmission power of the UAV virtual antenna array at time slot t, g UM [t] denotes the channel power gain between the UAV virtual antenna array and the mobile user at time slot t, K0 denotes the path loss constant of the UAV virtual antenna array and the mobile user, d UM [t] denotes the distance between the center of the UAV virtual antenna array and the mobile user at time slot t, a denotes the path loss exponent, Ω UM [t] denotes the small-scale fading at time slot t, which is simulated by a Rician distribution; g BM [t] denotes the channel power gain between the base station and the mobile user at time slot t, P B [t] denotes the transmission power of the base station at time slot t; s 2 denotes the noise power; G UM [t] denotes the antenna gain of the UAV virtual antenna array towards the position of the mobile user at time slot t, wherein, (0 UM [t], f UM [t]) denotes the direction of the mobile user, 0 UM [t] denotes the elevation angle of the mobile user at time slot t, f UM [t] denotes the azimuth angle of the mobile user at time slot t, w(0, f) denotes the amplitude of the far-field beam pattern of each element of the UAV, which is usually taken as 1, and n e [0, 1] denotes the antenna array efficiency; AF is the array factor, denotes the position of the i-th UAV at time slot t, I i [t] denotes the exciting current weight of the i-th UAV at time slot t, k c = 2p / l denotes the phase constant, and l denotes the wavelength;

[0054] The second objective function f2 denotes the total energy consumption of the UAV during the communication process. Wherein E i[t] represents the energy consumption of the UAV i in the t time slot, T represents the flight end time, v i [t] represents the speed of the UAV i in the t time slot, m UAV m represents the mass of the UAV, g represents the gravitational acceleration, and h[0] represents the initial height of the UAV; P(v i [t]) represents the propulsion energy consumption of the UAV when flying in a two-dimensional horizontal space, P B and P I respectively represent two constants, v tip , v0, d0, s, p, and A respectively represent the rotor tip speed, the hover average rotor induced velocity, the fuselage drag ratio, the rotor solidity, the air density, and the rotor disc area.

[0055] Based on the above objective function, an optimal policy network is trained using an evolutionary reinforcement learning algorithm, which can output the best position of the UAV movement and the optimal excitation current weight of the UAV according to the observed environment state at each time, so as to maximize the above two objectives.

[0056] II. The initialization of the population is implemented. n policy networks and n uniformly distributed weight vectors (each policy network corresponds to a weight vector) are randomly initialized, and the reinforcement learning algorithm based on the long short-term memory network is used to optimize these policy networks to generate the initial population.

[0057] wherein the weight vector is an array (a, b) representing the weights of the first objective function and the second objective function; the values of a and b satisfy a∈[0, 1], b∈[0, 1], and a+b=1.

[0058] The specific optimization process of each policy network is as follows:

[0059] (1) The policy network outputs the action to be performed by the UAV (i.e., the horizontal movement direction, the horizontal movement distance, the vertical movement distance, and the excitation current weight) according to the environment state information, and inputs the action to the environment.

[0060] (2) After the UAV performs the action, the environment returns a reward vector (including the transmission rate of the UAV and the user link and the energy consumption of the UAV) in real time.

[0061] (3) The reward vector is linearly weighted by the weight vector corresponding to the policy network to obtain a weighted reward value, and the reinforcement learning algorithm based on the long short-term memory network updates the policy network parameters according to the reward value.

[0062] (4) Steps (1)-(3) are repeated n iter times, and the policy network generated by each update is saved to the population to form a new population.

[0063] All updated populations and the initialized population constitute the initial population. Since the number of policy networks in the initial population is greater than the number of policy networks in the initialized population, one weight vector in the initial population will correspond to multiple policy networks.

[0064] Third, update the population using the performance buffer method to maintain the diversity and quality of the offspring population. Figure 3 As shown in Figure 2, all possible values ​​of f1 and f2 constitute the performance space, and the policy network in the population is mapped to the performance space, and the population is saved through the performance buffer. The specific process is as follows:

[0065] (1) The performance space is divided into B num buffers, each of which can store up to B size A strategic network.

[0066] (2) The objective function value corresponding to the policy network is calculated based on the two objective functions established previously. The position of the policy network in the performance space is determined by the values ​​of the two objective functions.

[0067] (3) Each policy network is stored in the nearest buffer in the performance space.

[0068] (4) If the number of policy networks in the buffer exceeds B size , only retain B with the largest distance from the origin size A strategic network.

[0069] (5) A new population (the first population) is formed by all the strategy networks stored in the performance buffer pool.

[0070] Fourth, use the hyperplane-based strategy selection method to select the strategy network to be updated in the population (the second population), such as Figure 4 As shown, in the performance space, with each policy network as the center, r hyper Establish a hyperplane for the radius, and determine the probability of the policy network corresponding to each hyperplane being selected for update based on the number of policy networks contained in the plane (that is, the denser the policy network in the hyperplane, the lower the probability of being selected, and vice versa). The specific steps are as follows:

[0071] (1) Calculate the objective function value corresponding to the policy network based on the two objective functions in step 1, and determine the position of the policy network in the performance space based on the objective function value.

[0072] (2) Calculate the number of policy networks contained in each hyperplane (given a hyperplane, if the Euclidean distance between a policy network and the center of the hyperplane is less than r hyper , then the hyperplane contains the policy network).

[0073] (3) For each weight vector obtained by initialization, a unique corresponding policy network is selected for training, and the probability of selecting the policy network is as follows:

[0074]

[0075] where N i represents the number of policy networks contained in the hyperplane corresponding to the i-th policy network, and c is a constant greater than 1.

[0076] Five, the selected policy networks are optimized using a reinforcement learning algorithm based on a long short-term memory network to generate a new offspring population, and the parameters of the policy network are updated according to the following formula:

[0077]

[0078] where J clip is the policy gradient loss function, ∈ is a hyperparameter used to control the range of truncation, r[t](θ) represents the probability ratio, A ω [t] represents the extended advantage function at time slot t, λ∈[0,1] controls the balance between bias and variance, γ represents the discount factor, V π represents the vectorized value function, J V represents the value function loss, represents the target value function.

[0079] After each update, the objective function value is calculated according to the policy networks in the current population and the policy networks in the last generation Pareto set, and the non-dominated policy networks are found, and the set of non-dominated policy networks obtained after each update is used as the new Pareto set.

[0080] The iteration termination condition is whether the maximum number of iterations is reached; if the iteration termination condition is met, the final Pareto combination is output as the optimal Pareto set, otherwise the iteration update process continues.

[0081] The user selects a policy network from the Pareto set as the final policy network according to the demand preference. For example, if the requirement for energy consumption is higher, the policy network with the lowest energy consumption is selected.

[0082] The UAV nodes synchronize the data and form a virtual antenna array, and at the beginning of each time slot, the UAV obtains the user's position information and inputs the final policy network.

[0083] The policy network outputs the optimal position and optimal excitation current weight of the UAV.

[0084] The above process is repeated at each time slot until the communication ends.

[0085] The present application provides a method of high-reliable dynamic communication based on unmanned aerial vehicle virtual antenna array cooperative beamforming, the unmanned aerial vehicle can move to the optimal position and adjust the optimal excitation current weight to form a virtual antenna array according to the real-time position of the mobile user in real time, and then form a beam with high antenna gain and high directivity, and directly communicate with the mobile user far away without multi-hop, the main lobe direction of the beam points to the direction of the communicated mobile user, thereby realizing high signal-to-interference-plus-noise ratio of the receiving user and improving the transmission rate and anti-interference ability. In addition, by designing the moving position of the unmanned aerial vehicle, a certain trade-off between the maximum communication rate and the minimum energy consumption of the unmanned aerial vehicle can be achieved. Therefore, the communication method based on unmanned aerial vehicle virtual antenna array cooperative beamforming can achieve the purpose of high-speed dynamic communication of wireless network. The present application combines reinforcement learning algorithm with evolutionary learning method to solve multi-objective optimization problem, and combines long short-term memory network to strengthen the learning ability of the algorithm to high dynamic environment. In addition, the performance buffer method and the superplane-based task selection method are used to strengthen the diversity of the obtained final Pareto solution set, and more choices are provided for the user to select the appropriate strategy network according to different situations.

[0086] Although the embodiments of the present application have been disclosed as above, they are not limited to the application listed in the specification and embodiments, and can be fully applied to various fields suitable for the present application, and additional modifications can be easily realized by those skilled in the art, and therefore the present application is not limited to specific details and the figures shown and described herein, without departing from the general concept defined by the claims and the equivalent scope.

Claims

1. A high-speed dynamic communication multi-objective optimization method based on unmanned aerial vehicle cooperative beamforming, characterized in that, The method comprises the following steps: Step 1, establishing a calculation model of the sum of the transmission rates of the virtual antenna array of the UAV in the whole communication time as a first objective function, and establishing a calculation model of the total energy consumption of the UAV as a second objective function; and In each time slot, the position of the mobile user in the communication system is obtained; Step 2, determining the optimal position of the UAV and the optimal excitation current weight of the UAV as the optimization target of the maximum first objective function and the minimum second objective function; Step 3, the UAV moves to the optimal position and transmits communication data to the user according to the optimal excitation current weight; The calculation model of the sum of the transmission rates of the virtual antenna array of the UAV in the whole communication time is as follows: In the formula, represents the horizontal moving direction of the UAV, represents the horizontal moving distance of the UAV, represents the vertical moving distance of the UAV, represents the exciting current weight of the UAV; R UM [t] represents the transmission rate of the t time slot UAV virtual antenna array when communicating with the mobile user; Y UM [t] represents the signal-to-interference-plus-noise ratio of the t time slot UAV virtual antenna array when communicating with the mobile user; The signal-to-interference-plus-noise ratio (SINR) of the virtual antenna array of the UAV when communicating with the mobile user in the t time slot is calculated by the following formula: where P U [t] denotes the total transmission power of the t-th time slot UAV virtual antenna array, G UM [t] denotes the antenna gain of the t-th time slot UAV virtual antenna array towards the mobile user location, g UM [t] denotes the channel power gain between the t-th time slot UAV virtual antenna array and the mobile user, σ 2 denotes the noise power, P B [t] denotes the transmission power of the t-th time slot base station, g BM [t] denotes the channel power gain between the t-th time slot uncorrelated base station and the mobile user; The calculation model of the total energy consumption of the UAV is as follows: In the formula, E i [t] represents the energy consumption of the UAV i in the t time slot, T represents the flight end time, v i [t] represents the speed of the UAV i in the t time slot, m UAV represents the mass of the UAV, g represents the gravitational acceleration, h i [0] and v i [0] respectively represent the initial height and initial speed of the UAV i, h i [T] and v i [T] respectively represent the height and speed of the UAV i at T time, P(v i [t]) represents the propulsion energy consumption of the UAV when flying in a two-dimensional horizontal space in the t time slot; N UAV represents the number of UAVs.

2. The method of claim 1, wherein, The propulsive energy consumption of the UAV when flying in a two-dimensional horizontal space is calculated by the following formula: Among them, P B and P I are all constants, v tip , v0, d0, s, ρ and A represent the rotor tip speed, hovering average rotor induced speed, fuselage drag ratio, rotor solidity, air density and rotor disc area respectively; v i represents the speed of drone i.

3. The high-speed dynamic communication multi-objective optimization method based on UAV cooperative beamforming according to claim 1 or 2, characterized in that, In the step 2, the optimal position of the UAV and the optimal excitation current weight of the UAV are determined based on the long short-term memory network, comprising the following steps: Step 1, randomly initializing the network parameters of a plurality of long short-term memory networks and the weight vectors of the first objective function and the second objective function, obtaining a plurality of policy networks and the weight vectors corresponding to the policy networks as an initialization population; the policy networks in the initialization population are iteratively updated based on the reinforcement learning algorithm of the long short-term memory network, and all the updated populations and the initialization population form a first-generation population; Step 2, a plurality of policy networks are selected from the first-generation population by using the performance buffer method to obtain a first population; Step 3, the first population is screened by using the hyperplane method to obtain a policy network corresponding to each weight vector uniquely as a second population; Step 4, the policy networks in the second population are updated based on the reinforcement learning method of the long short-term memory network, and the non-dominated policy networks in the update results are selected according to the first objective function and the second objective function, and the set of the non-dominated policy networks forms a Pareto solution set; Step 4 is repeated until the maximum number of iterations is reached, and an optimal Pareto solution set is obtained; Step 5, the position of the user is taken as an input parameter, and the optimal position of the UAV and the optimal excitation current weight of the UAV are output by the policy network in the optimal Pareto solution set.

4. The high-speed dynamic communication multi-objective optimization method based on UAV cooperative beamforming according to claim 3, characterized in that, The first objective function and the second objective function weight vector are an array (a, b); Wherein, the value conditions of a and b satisfy a∈[0, 1], b∈[0, 1], and a+b=1.

Citation Information

Patent Citations

  • Unmanned aerial vehicle swarm secure communication method

    CN114221694A

  • Unmanned aerial vehicle self-organizing multi-hop network data transmission method for disaster area

    CN115866575A