An anti-interference method for unmanned aerial vehicle communication network joint power-position optimization
By constructing an anti-jamming communication system for unmanned aerial vehicles (UAVs) and using deep reinforcement learning for power-position joint optimization, the problem of balancing coverage and anti-jamming strength in UAV communication networks in complex electromagnetic interference environments was solved, maximizing the number of user connections.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- GUANGDONG UNIV OF TECH
- Filing Date
- 2026-01-15
- Publication Date
- 2026-05-29
AI Technical Summary
In complex electromagnetic interference environments, it is difficult for UAV communication networks to simultaneously achieve both wide coverage and strong anti-interference capabilities. Existing optimization methods have failed to effectively address the issues of the number of user connections and connection stability.
Construct an anti-jamming communication system for unmanned aerial vehicles (UAVs) by using deep reinforcement learning for power-position joint optimization, designing a multi-dimensional state space and reward function for UAVs and users, and optimizing the UAV's transmission power and position to maximize the number of stably connected users.
It improves the interference perception and collaboration efficiency of drones, and maximizes the number of user connections in interference environments.
Smart Images

Figure CN122120804A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the technical field of unmanned aerial vehicle (UAV) communication, and more specifically, to an anti-interference method for joint power-position optimization of UAV communication networks. Background Technology
[0002] In the practical application of Unmanned Aerial Vehicle (UAV) communication networks, one of the core requirements is to maximize the number of ground user connections. This requirement not only demands that the network cover more potential users within a wider geographical area, but also that each connected user's connection possess stable and reliable communication quality to meet basic service needs such as voice calls, data transmission, and emergency command issuance. However, in actual operating environments, UAV communication networks typically face the dual challenges of complex electromagnetic interference and multi-dimensional collaborative optimization. These two challenges directly constrain the achievement of the goal of maximizing the number of user connections.
[0003] On the one hand, complex electromagnetic interference becomes a core obstacle to disrupting connection stability: Real-world communication environments contain various interference sources, including malicious interference devices, electromagnetic noise in industrial environments (such as electromagnetic radiation from factory motors and power transmission lines), and interference from the same or adjacent channels (such as signal interference from surrounding ground base stations and other UAV communication systems). These interference sources directly cause a significant decrease in the signal-to-interference-plus-noise ratio (SINR) of the received signal at the user end, falling below the communication threshold, leading to connection interruptions, soaring data transmission error rates, and access failures, severely damaging communication stability. On the other hand, there is a strong coupling relationship between the multi-dimensional control variables of UAVs, making it difficult to simultaneously balance "coverage breadth" and "interference resistance strength" through optimization of a single dimension. The core control variables for UAVs include horizontal position, transmit power, flight altitude, and beam angle. Among these, position and transmit power are key variables affecting user connectivity: horizontal position directly determines the UAV's ground coverage area; a well-planned deployment allows the UAV to cover more potential users. Transmit power affects signal propagation distance and interference resistance; higher transmit power enhances signal penetration and mitigates some interference, but also increases UAV power consumption and shortens runtime. The coupling relationship between these two variables is as follows: if only horizontal position is optimized to expand coverage, a fixed transmit power may not be sufficient to resist interference within the coverage area, leading to connection interruptions for users in that area; conversely, if only transmit power is optimized to improve interference resistance, a fixed deployment may result in edge users being unable to access the network due to excessive path loss, even with increased power, thus limiting coverage breadth. Existing research rarely focuses on optimizing user connection performance and does not consider the impact of interference on user connection status. The methods are limited to optimization of a single dimension or some dimensions, failing to fully consider the coupling effect between variables and the impact of complex interference environments. As a result, it is difficult to simultaneously meet the requirements of coverage breadth and anti-interference strength in real-world scenarios. The number of user connections and connection stability cannot reach the ideal level, making it difficult to meet the application needs in complex scenarios. Summary of the Invention
[0004] To address the issues of weak interference perception, limited optimization dimensions, and low collaboration efficiency in existing UAV communication technologies, this invention proposes a joint power-position optimization anti-interference method for UAV communication networks, which aims to maximize the number of connected users under interference conditions.
[0005] To achieve the above-mentioned technical effects, the technical solution of the present invention is as follows: In a first aspect, this application proposes an anti-interference method for joint power-position optimization in UAV communication networks, comprising the following steps: S1: Construct an anti-jamming communication system for unmanned aerial vehicles (UAVs); the anti-jamming communication system for UAVs includes a jammer, a UAV, and a user, the UAV provides communication services to the user, and the jammer interferes with the user's communication; S2: Within a certain time period, with the objective function of maximizing the number of users that the UAV anti-interference communication system can stably connect to, a power-position joint optimization model is constructed based on the user-UAV connection state constraints, UAV allocable bandwidth constraints, UAV transmit power constraints, and UAV position constraints. S3: Design a multi-dimensional state space for UAV-users, a position-power adjustment action space for UAVs, and a reward function. Based on the multi-dimensional state space for UAV-users, the position-power adjustment action space for UAVs, and the reward function, use deep reinforcement learning to solve the power-position joint optimization model to obtain the optimized UAV transmit power, UAV position, the optimal movement action of the UAV at each time step, and the maximum number of users that can be stably connected.
[0006] Preferably, the UAV anti-jamming communication model constructed in S10 includes... I Launch drones at a constant altitude H Fly over the target area, for U Each user provides communication services; each drone moves its position in a discretized grid space, and the emission energy of each drone is concentrated at the aperture angle below the drone. θ The ground coverage area of each drone corresponds to a dynamic coverage radius of [missing information]. R i The disk; the dynamic coverage radius R i Satisfying the expression:
[0007] in, and These are the minimum and maximum values of the drone's transmission power, respectively. P i Indicates the first i The launch power of the drone; The U Users are randomly distributed in the target area; all UAVs share the same spectrum, and orthogonal frequency division multiple access technology is used to achieve user spectrum access. Users of the same UAV communicating on different resource blocks (RBs) do not interfere with each other because the spectrum is orthogonal. Jammers are deployed in a fixed location. J By calculating the distance between the user and the jammer in real time, it is determined whether the user is within the range of medium-to-strong interference. The coverage radius of medium-to-strong interference satisfies the calculation expression:
[0008] in, Indicates the strong interference coverage radius in the reference. Indicates the interference power attenuation coefficient; Preferably, the power-location joint optimization model includes an objective function, expressed as:
[0009] in, Indicates user u With drones i At any moment t The connection status, A value of 0 indicates no connection. Setting it to 1 indicates a connection; Indicates drone i At any moment t The optimal movement action; Indicates drone i At any moment t The transmission power; T Indicates the total time; I Indicates the total number of drones; U This represents the total number of users.
[0010] Preferably, the power-location joint optimization model further includes constraints, which include: The connection state constraints between the user and the drone, wherein the expression satisfying the connection state constraints is: , ; The allocatable bandwidth constraint for the UAV is expressed as follows: ; The UAV transmission power constraint satisfies the following expression: ; The UAV position constraint satisfies the following expression: ; in, For users u Number of resource blocks required; Bandwidth can be allocated for drones; For resource block bandwidth; and These represent the x-coordinate and y-coordinate of the UAV's projected position, respectively.
[0011] Preferably, bandwidth can be allocated to each drone. The resource is divided into several blocks, and the user... u Number of resource blocks required Satisfying the expression:
[0012] in, For the user's minimum throughput requirement, To associate drones i To users u The signal-to-interference-plus-noise ratio; the signal-to-interference-plus-noise ratio Satisfying the expression:
[0013] in, The noise power spectral density; The power spectral density of the jammer; The transmit power spectral density of the UAV; For channel gain, For free space path loss, The center carrier frequency, For drones i With users u The straight-line distance c At the speed of light, Additional loss for line-of-sight channels.
[0014] Preferably, the UAV-user multidimensional state space is constructed based on the UAV's own location, the number of currently connected users, the UAV's real-time transmission power, the jammer's global information, and the user's jamming state vector. The expression is:
[0015] in,( () represents the location coordinates of the UAV within the grid space; Indicates the number of currently connected users; Indicates the real-time transmission power of the drone; () indicates the coordinates of the jammer's position; Indicates the jammer's transmission power; Indicates the strong interference coverage radius of the jammer; This represents the user interference state vector, where each element of the user interference state vector is identified by 1 or 0, indicating whether the corresponding user is within the strong coverage range of the jammer.
[0016] Preferably, the UAV moves in a grid space at fixed step sizes. The UAV position-power adjustment action space includes position movement actions and power adjustment actions. The position movement actions include up, down, left, right, and hovering. The power adjustment actions include increasing power, decreasing power, and maintaining power. The UAV position-power adjustment action space is constructed through joint optimization of position movement and power adjustment. Satisfies the expression: ; in, This indicates the number of combinations of position movement actions and power adjustment actions.
[0017] Preferably, the reward function is designed by guiding the optimization direction with positive rewards and constraining unreasonable behaviors with negative penalties. Satisfies the expression:
[0018] Among them, in positive rewards, , , All are positive reward parameters. Basic connection rewards, , This represents the current number of connected users. Rewards for stable connections, , The number of users who maintain connection for two consecutive steps; Reconnection reward, , This represents the number of users who reconnected after a disconnection. In negative punishment , , , All are negative penalty parameters; Penalty for connection interruption, , This represents the number of users currently experiencing an interruption. Penalty for medium to strong interference zones. , The number of users whose services were disrupted while serving drones; Penalty for power usage , Reference power; To cover overlap penalty, , For drones i With drones k The dynamic average coverage radius, and drones i With drones k The dynamic coverage radius; Fixed weighting coefficients; I Indicates the total number of drones. U This represents the total number of users.
[0019] Preferably, the deep reinforcement learning uses a deep Q-network as the core decision network and adopts a dual-network design of a main network and a target network. The main network and the target network have the same structure but different parameter update mechanisms. The deep reinforcement learning utilizes the interaction between the agent and the environment. Intelligent agents acquire information through interaction with the environment. t Time-based drones - user multidimensional state space Then, standardization processing is performed to obtain a standardized UAV-user multidimensional state space. ; Utilizing the main network for standardized UAV-user multidimensional state space Drone position-power adjustment motion space After processing, the Q value is obtained, expressed as: ; Based on the Q-value and the improved ε-greedy strategy, select t Momentary drone position-power adjustment action space By executing t Momentary drone position-power adjustment action space Continue to interact with the environment and receive rewards. The improved ε-greedy strategy satisfies the expression: ,in, Number of steps executed. This refers to the decay step size parameter; The intelligent agent generates experience tuples from its interactions with the environment. Storing them in an experience pool, and using an adaptive sampling strategy to select learning experiences; among which... d This is the end marker; The target network calculates the target Q-value based on the learned experience. The target Q-value is the sum of the immediate reward and the future discounted reward, expressed as: ; Where γ is the discount factor, and at the terminal state d=1, the target Q value only includes the immediate reward, i.e. ; Minimize the target Q-value, the Q-value, and the loss function; update the main network parameters through gradient descent; and synchronize the target network parameters with the main network through a soft update mechanism. After training the deep Q-network, the final Q-value is obtained. Based on the final Q-value, the optimized drone transmission power, drone position, optimal drone movement at each moment, and the maximum number of users that can be stably connected are obtained.
[0020] Secondly, this application proposes an anti-jamming system for joint power-position optimization of UAV communication networks, used to implement the anti-jamming method for joint power-position optimization of UAV communication networks, comprising: A communication system construction module is used to construct an anti-jamming communication system for unmanned aerial vehicles (UAVs). The anti-jamming communication system for UAVs includes a jammer, a UAV, and a user. The UAV provides communication services to the user, and the jammer interferes with the user's communication. The optimization model building module is used to construct a power-position joint optimization model within a certain time period, with the objective function of maximizing the number of users that the UAV anti-interference communication system can stably connect to. This model is based on user-UAV connection state constraints, UAV allocable bandwidth constraints, UAV transmit power constraints, and UAV position constraints. The optimization model solving module is used to design the UAV-user multidimensional state space, UAV position-power adjustment action space, and reward function. Based on the UAV-user multidimensional state space, UAV position-power adjustment action space, and reward function, deep reinforcement learning is used to solve the power-position joint optimization model to obtain the optimized UAV transmit power, UAV position, optimal movement action of UAV at each time step, and the maximum number of users that can be stably connected.
[0021] Compared with the prior art, the beneficial effects of the technical solution of the present invention are: This invention proposes an anti-jamming method for UAV communication networks based on joint power-position optimization. First, an anti-jamming communication system for UAVs is constructed, including a jammer, UAVs, and users. The UAVs provide communication services to the users, while the jammer interferes with user communication. Within a certain time period, the objective function is to maximize the number of users stably connected by the anti-jamming communication system. Based on constraints on the connection state between users and UAVs, the allocable bandwidth of the UAVs, the UAV's transmit power, and the UAV's position, a power-position joint optimization model is constructed. A multi-dimensional state space for UAVs and users, a position-power adjustment action space for UAVs, and a reward function are designed. Finally, based on the multi-dimensional state space for UAVs and users, the position-power adjustment action space for UAVs, and the reward function, deep reinforcement learning is used to solve the power-position joint optimization model, obtaining the optimized UAV transmit power, UAV position, the optimal movement action of the UAV at each time step, and the maximized number of stably connected users. This invention improves the interference perception capability of UAVs, increases the optimization dimension, enhances collaborative efficiency, and can maximize the number of connected users under interference conditions. Attached Figure Description
[0022] Figure 1 A flowchart illustrating the anti-interference method for joint power-position optimization of UAV communication networks proposed in this embodiment of the invention; Figure 2 This diagram illustrates the UAV anti-jamming communication system model proposed in this embodiment of the invention. Figure 3 This diagram illustrates the deep Q-network framework proposed in this embodiment of the invention. Figure 4 This represents the convergence graph of the number of randomly distributed user connections proposed in this embodiment of the invention. Figure 5 This represents the convergence graph of the number of user sparsely distributed connections proposed in this embodiment of the invention; Figure 6 This represents a convergence graph of the number of user hotspot distribution connections proposed in this embodiment of the invention. Figure 7 This diagram illustrates the comparison between increased jammer transmission power and the number of connected users as presented in this embodiment of the invention. Figure 8 This diagram illustrates the comparison between the signal-to-interference-plus-noise ratio (SINR) communication threshold and the number of connected users as proposed in this embodiment of the invention. Figure 9 This diagram illustrates the composition of the anti-jamming system for joint power-position optimization of the UAV communication network proposed in this embodiment of the invention. Detailed Implementation
[0023] The accompanying drawings are for illustrative purposes only and should not be construed as limiting the scope of this patent. To better illustrate this embodiment, some parts of the accompanying drawings may be omitted, enlarged, or reduced, and do not represent the actual dimensions; It is understandable to those skilled in the art that some well-known details may be omitted from the accompanying drawings.
[0024] The technical solution of the present invention will be further described below with reference to the accompanying drawings and embodiments; The positional relationships depicted in the accompanying drawings are for illustrative purposes only and should not be construed as limiting this patent.
[0025] Example 1 This embodiment provides an anti-interference method for joint power-position optimization in UAV communication networks. The flowchart of this method is shown below. Figure 1 This includes the following steps: S1: Construct an anti-jamming communication system for unmanned aerial vehicles (UAVs); the anti-jamming communication system for UAVs includes a jammer, a UAV, and a user, the UAV provides communication services to the user, and the jammer interferes with the user's communication; S2: Within a certain time period, with the objective function of maximizing the number of users that the UAV anti-interference communication system can stably connect to, a power-position joint optimization model is constructed based on the user-UAV connection state constraints, UAV allocable bandwidth constraints, UAV transmit power constraints, and UAV position constraints. S3: Design a multi-dimensional state space for UAV-users, a position-power adjustment action space for UAVs, and a reward function. Based on the multi-dimensional state space for UAV-users, the position-power adjustment action space for UAVs, and the reward function, use deep reinforcement learning to solve the power-position joint optimization model to obtain the optimized UAV transmit power, UAV position, the optimal movement action of the UAV at each time step, and the maximum number of users that can be stably connected.
[0026] In this embodiment, a UAV anti-jamming communication system is first constructed. This system includes a jammer, a UAV, and users. The UAV provides communication services to the users, and the jammer interferes with user communication. Within a certain time period, with the objective function being to maximize the number of users that the UAV anti-jamming communication system can stably connect to, a power-position joint optimization model is constructed based on user-UAV connection state constraints, UAV allocable bandwidth constraints, UAV transmit power constraints, and UAV position constraints. A multi-dimensional UAV-user state space, a UAV position-power adjustment action space, and a reward function are designed. Finally, based on the UAV-user multi-dimensional state space, the UAV position-power adjustment action space, and the reward function, deep reinforcement learning is used to solve the power-position joint optimization model, obtaining the optimized UAV transmit power, UAV position, the optimal movement action of the UAV at each moment, and the maximized number of users that can stably connect.
[0027] Example 2 In this embodiment, as Figure 2 The diagram shown illustrates a UAV anti-jamming communication system model. In the UAV anti-jamming communication model constructed in S10, it includes... I Launch drones at a constant altitude H Fly over the target area, for U Each user provides communication services; each drone moves its position in a discretized grid space, and the emission energy of each drone is concentrated at the aperture angle below the drone. θ The ground coverage area of each drone corresponds to a dynamic coverage radius of [missing information]. R i The disk; the dynamic coverage radius R i Satisfying the expression:
[0028] in, and These are the minimum and maximum values of the drone's transmission power, respectively. P i Indicates the first i The launch power of the drone; The UUsers are randomly distributed in the target area; all UAVs share the same spectrum, and orthogonal frequency division multiple access technology is used to achieve user spectrum access. Users of the same UAV communicating on different resource blocks (RBs) do not interfere with each other because the spectrum is orthogonal. Jammers are deployed in a fixed location. J By calculating the distance between the user and the jammer in real time, it is determined whether the user is within the range of medium-to-strong interference. The coverage radius of medium-to-strong interference satisfies the calculation expression:
[0029] in, Indicates the strong interference coverage radius in the reference. Indicates the interference power attenuation coefficient; Specifically, the drones fly without a central control unit, and each drone is able to autonomously decide on its movement strategy in a coordinated manner based on its local observations, status, and information exchanged with other drones. This represents the baseline medium-strong interference coverage radius when the jammer's transmit power is 20dBm, in meters, with a default value of 250m.
[0030] In this embodiment, the power-location joint optimization model includes an objective function, expressed as:
[0031] in, Indicates user u With drones i At any moment t The connection status, A value of 0 indicates no connection. Setting it to 1 indicates a connection; Indicates drone i At any moment t The optimal movement action; Indicates drone i At any moment t The transmission power; T Indicates the total time; I Indicates the total number of drones; U This represents the total number of users.
[0032] The power-location joint optimization model also includes constraints, which include: The connection state constraints between the user and the drone, wherein the expression satisfying the connection state constraints is: , ; The allocatable bandwidth constraint for the UAV is expressed as follows: ; The UAV transmission power constraint satisfies the following expression: ; The UAV position constraint satisfies the following expression: ; in, For users u Number of resource blocks required; Bandwidth can be allocated for drones; For resource block bandwidth; and These represent the x-coordinate and y-coordinate of the UAV's projected position, respectively.
[0033] Bandwidth can be allocated to each drone The resource is divided into several blocks, and the user... u Number of resource blocks required Satisfying the expression:
[0034] in, For the user's minimum throughput requirement, To associate drones i To users u The signal-to-interference-plus-noise ratio; the signal-to-interference-plus-noise ratio Satisfying the expression:
[0035] in, The noise power spectral density; The power spectral density of the jammer; The transmit power spectral density of the UAV; For channel gain, For free space path loss, The center carrier frequency, For drones i With users u The straight-line distance c At the speed of light, Additional loss for line-of-sight channels.
[0036] Specifically, based on user association requirements, a resource block (RB) allocation mechanism is introduced, with each user having a minimum throughput requirement. The number of resource blocks (RBs) to be allocated Determined by both throughput constraints and signal-to-interference-plus-noise ratio (SINR), a two-stage user association strategy is adopted: In the first stage, users associate themselves with the network of networks within their coverage area. And the signal-to-interference-plus-noise ratio (SINR) is greater than or equal to the communication threshold. SThe drone sends a connection request, and the drone allocates resources in ascending order of user resource block (RB) demand. Under the same RB demand, it prioritizes connecting to the user with the highest signal-to-interference-plus-noise ratio (SINR). In the second phase, users who fail to connect extend their coverage area (distance ≤ α). And scaling the communication threshold (signal-to-interference-plus-noise ratio SINR ≥ β) S The drone requested a connection again, and the drone completed the replenishment allocation based on the remaining resources. , .
[0037] In this embodiment, the UAV-user multidimensional state space is constructed based on the UAV's own location, the number of currently connected users, the UAV's real-time transmission power, the jammer's global information, and the user's jamming state vector. The expression is:
[0038] in,( () represents the location coordinates of the UAV within the grid space; Indicates the number of currently connected users; Indicates the real-time transmission power of the drone; () indicates the coordinates of the jammer's position; Indicates the jammer's transmission power; Indicates the strong interference coverage radius of the jammer; This represents the user interference state vector, where each element of the user interference state vector is identified by 1 or 0, indicating whether the corresponding user is within the strong coverage range of the jammer.
[0039] Specifically, the UAV's position coordinates accurately reflect its deployment location within the mission area, i.e., the grid space; the current number of connected users directly reflects the communication service effect; the UAV's real-time transmission power directly affects coverage and anti-interference capabilities; the inclusion of global jammer information provides a basis for UAVs to avoid interference; and the user interference state vector, with a length equal to the total number of users within the UAV anti-interference communication system, uses 1 or 0 to indicate whether the corresponding user is within the strong coverage range of the jammer, achieving accurate perception of the user's communication environment. Therefore, the UAV-user multi-dimensional state space includes eight basic states and... U User interference status.
[0040] In this embodiment, the UAV moves in a grid space at fixed step sizes. The UAV position-power adjustment action space includes position movement actions and power adjustment actions. The position movement actions include up, down, left, right, and hovering. The power adjustment actions include increasing power, decreasing power, and maintaining power. The UAV position-power adjustment action space is constructed through joint optimization of position movement and power adjustment. Satisfies the expression: ; in, This indicates the number of combinations of position movement actions and power adjustment actions.
[0041] Specifically, hovering ensures stable service at the current location; power adjustment meets the UAV's transmit power constraints. In this embodiment, the drone's position-power adjustment motion space The expression is: It includes 15 discrete actions, and achieves anti-interference decision-making through joint optimization of position movement and power adjustment.
[0042] In this embodiment, the reward function is designed by guiding optimization direction with positive rewards and constraining unreasonable behavior with negative penalties. Satisfies the expression:
[0043] Among them, in positive rewards, , , All are positive reward parameters. Basic connection rewards, , This represents the current number of connected users. Rewards for stable connections, , The number of users who maintain connection for two consecutive steps; Reconnection reward , This represents the number of users who reconnected after a disconnection. In negative punishment , , , All are negative penalty parameters; Penalty for connection interruption, , This represents the number of users currently experiencing an interruption. Penalty for medium to strong interference zones. , The number of users whose services were disrupted while serving drones; Penalty for power usage , Reference power; To cover overlap penalty, , For drones i With drones k The dynamic average coverage radius, and drones i With drones k The dynamic coverage radius; Fixed weighting coefficients; I Indicates the total number of drones. U This represents the total number of users.
[0044] Specifically, By normalizing the current number of connected users, the reward magnitude is ensured to be consistent across different scale scenarios, directly incentivizing the improvement of connection capabilities. Encourage drones to maintain continuous connectivity. Incentivize drones to reconnect to disconnected users. To constrain disconnection caused by interference or improper strategies. Guide the drone to avoid areas with moderate to strong interference. Balancing power gain with energy consumption costs This means that the closer the drones are to each other and the more overlapping their coverage, the greater the penalty.
[0045] The deep reinforcement learning uses a deep Q-network as the core decision network and adopts a dual-network design of a main network and a target network. The main network and the target network have the same structure but different parameter update mechanisms. The deep reinforcement learning utilizes the interaction between the agent and the environment. Constrain unstable behavior of connections; Guide the drone to avoid the interference area; Balancing power gain with energy consumption cost; Distributed deployment is encouraged to reduce mutual interference. The positive reward parameters and negative penalty parameters are determined through experimental tuning to ensure that the reward function accurately guides the drones to find the optimal balance between anti-interference, connectivity maintenance, low energy consumption, and collaborative deployment.
[0046] In this embodiment, the deep reinforcement learning uses a deep Q-network as the core decision network and adopts a dual-network design of a main network and a target network. The main network and the target network have the same structure but different parameter update mechanisms. The deep reinforcement learning utilizes the interaction between the agent and the environment. Intelligent agents acquire information through interaction with the environment. t Time-based drones - user multidimensional state space Then, standardization processing is performed to obtain a standardized UAV-user multidimensional state space. ; Utilizing the main network for standardized UAV-user multidimensional state space Drone position-power adjustment motion space After processing, the Q value is obtained, expressed as: ; Based on the Q-value and the improved ε-greedy strategy, select tMomentary drone position-power adjustment action space By executing t Momentary drone position-power adjustment action space Continue to interact with the environment and receive rewards. The improved ε-greedy strategy satisfies the expression: ,in, Number of steps executed. This refers to the decay step size parameter; The intelligent agent generates experience tuples from its interactions with the environment. Storing them in an experience pool, and using an adaptive sampling strategy to select learning experiences; among which... d This is a termination marker; The target network calculates the target Q-value based on the learned experience. The target Q-value is the sum of the immediate reward and the future discounted reward, expressed as: ; Where γ is the discount factor, and at the terminal state d=1, the target Q value only includes the immediate reward, i.e. ; Minimize the target Q-value, the Q-value, and the loss function; update the main network parameters through gradient descent; and synchronize the target network parameters with the main network through a soft update mechanism. After training the deep Q-network, the final Q-value is obtained. Based on the final Q-value, the optimized drone transmission power, drone position, optimal drone movement at each moment, and the maximum number of users that can be stably connected are obtained.
[0047] For a detailed diagram of the deep Q-network framework, please refer to [link / reference needed]. Figure 3 The deep reinforcement learning framework used in this embodiment is based on Independent Q-Learning (IQL). This architecture balances exploration and utilization, improves sample efficiency, and adapts to decision-making needs in complex environments through a hierarchical design.
[0048] like Figure 3 As shown, during the perception phase, the agent acquires the multi-dimensional state space of the drone and the user through environmental interaction. After standardization, the data is input into the main network, which then uses the standardized UAV-user multidimensional state space. and the current drone position-power adjustment action space The Q-value is calculated. During the decision-making phase, an improved ε-greedy strategy is employed to achieve a dynamic balance between exploration and exploitation. The exploration rate ε decays exponentially, ensuring that the exploration rate smoothly converges to its minimum value. During the exploration phase, a hybrid strategy was adopted to adjust the position-power space of the UAV. Selection: A completely random action is executed with probability p, and a second-best action is selected with probability 1-p. The optimal action is then randomly selected from the 2nd to 4th best actions after the Q-value of the current state is calculated and sorted by the main network. This selection is based on the UAV's position-power adjustment action space. If the exploration rate ε is insufficient, the algorithm reverts to randomness. The hybrid strategy employed avoids overexploring ineffective actions while retaining the ability to discover potential optimal strategies. When the exploration rate ε is below a threshold, the agent selects the action with the highest Q-value output from the main network, thus utilizing known optimal strategies.
[0049] In this embodiment, the deep Q-network framework also includes an experience replay mechanism. The agent uses the experience tuples generated from each interaction. Stored in the experience pool, For the next state, d This serves as a termination marker. The Deep Q-Network framework also introduces an adaptive sampling strategy to select learning experiences to adapt to dynamic changes in the environment. The adaptive sampling strategy is as follows: a uniform random sampling mode is used in the early stages of training; when performance stagnation is detected, i.e., the average performance is lower than 95% of the best performance for 30 consecutive rounds, it automatically switches to a priority sampling mode. In the priority sampling mode, 70% of the samples come from the last 1 / 4 of recent experiences in the buffer, and 30% come from historical experiences. By increasing the weight of recent samples, the learning ability to adapt to the latest dynamics of the environment is enhanced.
[0050] In this embodiment, a multi-layer fully connected neural network is used as the Q-function approximator. A dual-network design of a main network and a target network improves training stability. The main network and the target network have identical structures but different parameter update mechanisms. Both the main network and the target network consist of an input layer, three hidden layers, and an output layer. The input layer receives the multi-dimensional state space of the UAV-user interface. The hidden layer space is mapped through a linear transformation. All three hidden layers use the ReLU activation function to introduce nonlinearity to fit the complex state-action value relationship. The output layer directly outputs the Q value corresponding to each discrete action, providing a value basis for action decision-making.
[0051] The main network calculates the Q-value through forward propagation. The Q-value expression is: The target network calculates the target Q-value based on the Bellman optimality equation. Specifically, for non-terminating states, i.e., when d=0, the target Q-value is the sum of the immediate reward and the future discounted reward, satisfying the expression... Where γ is a discount factor used to balance immediate rewards and future rewards; for the terminal state, i.e., when d=1, the objective Q value only includes immediate rewards, satisfying the expression The target Q-value is calculated by the target network and gradient propagation is truncated using detach() to ensure stability.
[0052] The core of the network update mechanism is to update parameters by minimizing the difference between the target Q-value and the current Q-value. The loss function used is Huber loss, expressed as:
[0053] Where N is the batch size. This function approximates the Mean Squared Error (MSE) when the error is small to refine parameter optimization; when the error is large, it approximates the Mean Absolute Error (MAE) to suppress the interference of outliers on training and improve robustness. The main network parameters are updated via gradient descent, specifically: first, the gradient from the previous round is cleared; then, the difference between the Q-value and the target Q-value is minimized using the loss function, and the gradient is calculated via backpropagation. A dual constraint strategy is employed to prevent gradient explosion: gradient values are clipped to [-1, -1], and the gradient norm is constrained to within 1.0. Finally, the optimizer updates the main network parameters. The target network parameters are synchronized with the main network through a soft update mechanism, completely replicating the main network parameters during initialization. This allows the target network to both track the learning progress of the main network and avoid drastic fluctuations in the target value to maintain training stability, ultimately achieving efficient fitting of complex state-action value relationships, enabling the agent to quickly adapt to environmental dynamics and learn globally optimal anti-interference strategies.
[0054] Example 3 In this embodiment, the environment is implemented in Python using OpenAI's Gym toolkit, and the above method is implemented using the PyTorch library. Five drones are set to fly within a target area of 1000 m × 1000 m, with a fixed flight altitude of 350 m and an aperture angle of 60°. The initial position of the drones is [500, 500] m, and the initial transmit power is 27 dBm. 100 users requiring connection are randomly distributed within the flight area. A ground-based jammer with a transmit power of 20 dBm is fixed at [500, 500] m. The algorithm iterates 1000 times during the training phase, with 100 time slots for drone flight in each iteration. The simulation parameters are shown in Table 1.
[0055] Table 1
[0056] Set other parameters according to Table 1.
[0057] To verify the superiority of the present invention, two sets of comparative schemes were set up. FP scheme: the UAV's transmit power was fixed, and only the position action was optimized; FA scheme: the UAV's initial position was fixed, and only the power action was optimized.
[0058] like Figure 4 , Figure 5 , Figure 6 As shown, the horizontal axis represents the number of iterations, the vertical axis represents the average number of connected users, and IQL represents the scheme proposed in this invention. Figure 4 , Figure 5 and Figure 6 The convergence curves of user connection counts for the IQL scheme and the comparison scheme are presented in three typical scenarios: random user distribution, sparse distribution, and hotspot distribution, with a signal-to-interference-plus-noise ratio (SNR) threshold of 5 dBm. The results show that the IQL scheme exhibits a significant convergence advantage and performance superiority in all scenarios. These results fully verify that the "location-power" joint optimization mechanism can quickly adapt to different user distributions and interference environments. It expands coverage through dynamic location adjustment and enhances anti-interference capabilities through power optimization, effectively overcoming the limitations of single-dimensional optimization schemes and achieving a synergistic improvement in coverage and connection stability.
[0059] like Figure 7 As shown, the horizontal axis represents the UAV's transmission power, the vertical axis represents the number of connected users, and IQL represents the solution proposed in this invention. Figure 7 The study demonstrates the changing trends in the number of stable connected users for three schemes as the jammer's transmit power gradually increases from 20 dBm to 40 dBm. With increasing jamming power, the number of connected users decreases for all schemes, but the IQL scheme shows the smallest decrease and the best robustness. This result indicates that the IQL scheme's "interference awareness + joint optimization" mechanism effectively combats changes in jamming intensity: when jamming intensifies, the UAV maximizes the preservation of effective connections by employing strategies such as reducing user association priority in the jamming area, increasing transmit power, and moving to the edge of the jamming zone. In contrast, single-dimensional optimization schemes lack flexible anti-jamming adjustment methods, leading to a sharp deterioration in connection performance when jamming intensifies.
[0060] like Figure 8 As shown, the horizontal axis represents the signal-to-interference-plus-noise ratio (SINR) threshold, the vertical axis represents the number of connected users, and IQL represents the scheme proposed in this invention. Figure 8 The study demonstrates the change in the number of stable connected users for three schemes as the SINR communication threshold increases from 5dB to 15dB. A higher SINR threshold indicates stricter communication quality requirements and demands higher interference resistance from the schemes. As the SINR threshold increases, the number of connections decreases, with the IQL scheme showing the smallest decrease. This result indicates that the IQL scheme, by dynamically adjusting power to improve user SINR and optimizing deployment locations to select areas with better signal quality, better adapts to different communication quality requirements. In contrast, the comparative schemes, lacking collaborative optimization capabilities, struggle to meet user connection needs under strict communication thresholds.
[0061] Example 4 This embodiment provides an anti-jamming system for a UAV communication network with joint power-position optimization. See [link to documentation]. Figure 9 The system is used to implement the anti-interference method of joint power-position optimization for the UAV communication network, including: A communication system construction module is used to construct an anti-jamming communication system for unmanned aerial vehicles (UAVs). The anti-jamming communication system for UAVs includes a jammer, a UAV, and a user. The UAV provides communication services to the user, and the jammer interferes with the user's communication. The optimization model building module is used to construct a power-position joint optimization model within a certain time period, with the objective function of maximizing the number of users that the UAV anti-interference communication system can stably connect to. This model is based on user-UAV connection state constraints, UAV allocable bandwidth constraints, UAV transmit power constraints, and UAV position constraints. The optimization model solving module is used to design the UAV-user multidimensional state space, UAV position-power adjustment action space, and reward function. Based on the UAV-user multidimensional state space, UAV position-power adjustment action space, and reward function, deep reinforcement learning is used to solve the power-position joint optimization model to obtain the optimized UAV transmit power, UAV position, optimal movement action of UAV at each time step, and the maximum number of users that can be stably connected.
[0062] The same or similar labels correspond to the same or similar parts; The terms used to describe positional relationships in the accompanying drawings are for illustrative purposes only and should not be construed as limiting the invention. Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention, and are not intended to limit the implementation of the present invention. Those skilled in the art can make other variations or modifications based on the above description. It is neither necessary nor possible to exhaustively describe all embodiments here. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the scope of protection of the claims of the present invention.
Claims
1. An anti-interference method for joint power-position optimization in unmanned aerial vehicle (UAV) communication networks, characterized in that, Includes the following steps: S1: Construct an anti-jamming communication system for unmanned aerial vehicles (UAVs); the anti-jamming communication system for UAVs includes a jammer, a UAV, and a user, the UAV provides communication services to the user, and the jammer interferes with the user's communication; S2: Within a certain time period, with the objective function of maximizing the number of users that the UAV anti-interference communication system can stably connect to, a power-position joint optimization model is constructed based on the user-UAV connection state constraints, UAV allocable bandwidth constraints, UAV transmit power constraints, and UAV position constraints. S3: Design a multi-dimensional state space for UAV-users, a position-power adjustment action space for UAVs, and a reward function. Based on the multi-dimensional state space for UAV-users, the position-power adjustment action space for UAVs, and the reward function, use deep reinforcement learning to solve the power-position joint optimization model to obtain the optimized UAV transmit power, UAV position, the optimal movement action of the UAV at each time step, and the maximum number of users that can be stably connected.
2. The anti-interference method for joint power-position optimization of UAV communication networks according to claim 1, characterized in that, The anti-jamming communication model for UAVs constructed in S10 includes I Launch drones at a constant altitude H Fly over the target area, for U Each user provides communication services; each drone moves its position in a discretized grid space, and the emission energy of each drone is concentrated at the aperture angle below the drone. θ The ground coverage area of each drone corresponds to a dynamic coverage radius of [missing information]. R i The disk; the dynamic coverage radius R i Satisfying the expression: in, and These are the minimum and maximum values of the drone's transmission power, respectively. P i Indicates the first i The launch power of the drone; The U Users are randomly distributed in the target area; all UAVs share the same spectrum, and orthogonal frequency division multiple access technology is used to achieve user spectrum access. Users of the same UAV communicating on different resource blocks (RBs) do not interfere with each other because the spectrum is orthogonal. Jammers are deployed in a fixed location. J By calculating the distance between the user and the jammer in real time, it is determined whether the user is within the range of medium-to-strong interference. The coverage radius of medium-to-strong interference satisfies the calculation expression: in, Indicates the strong interference coverage radius in the reference. This represents the interference power attenuation coefficient.
3. The anti-interference method for joint power-position optimization of UAV communication networks according to claim 1, characterized in that, The power-location joint optimization model includes an objective function, expressed as follows: in, Indicates user u With drones i At any moment t The connection status, A value of 0 indicates no connection. Setting it to 1 indicates a connection; Indicates drone i At any moment t The optimal movement action; Indicates drone i At any moment t The transmission power; T Indicates the total time; I Indicates the total number of drones; U This represents the total number of users.
4. The anti-interference method for joint power-position optimization of UAV communication networks according to claim 3, characterized in that, The power-location joint optimization model also includes constraints, which include: The connection state constraints between the user and the drone, wherein the expression satisfying the connection state constraints is: , ; The allocatable bandwidth constraint for the UAV is expressed as follows: ; The UAV transmission power constraint satisfies the following expression: ; The UAV position constraint satisfies the following expression: ; in, For users u Number of resource blocks required; Bandwidth can be allocated for drones; For resource block bandwidth; and These represent the x-coordinate and y-coordinate of the UAV's projected position, respectively.
5. The anti-interference method for joint power-position optimization of UAV communication networks according to claim 4, characterized in that, Bandwidth can be allocated to each drone The resource is divided into several blocks, and the user... u Number of resource blocks required Satisfying the expression: in, For the user's minimum throughput requirement, To associate drones i To users u The signal-to-interference-plus-noise ratio; the signal-to-interference-plus-noise ratio Satisfying the expression: in, The noise power spectral density; The power spectral density of the jammer; The transmit power spectral density of the UAV; For channel gain, For free space path loss, The center carrier frequency, For drones i With users u The straight-line distance c At the speed of light, Additional loss for line-of-sight channels.
6. The anti-interference method for joint power-position optimization of UAV communication networks according to claim 2, characterized in that, Based on the UAV's own location, the number of currently connected users, the UAV's real-time transmission power, the jammer's global information, and the user's jamming state vector, the UAV-user multidimensional state space is constructed. The expression is: in,( () represents the location coordinates of the UAV within the grid space; Indicates the number of currently connected users; Indicates the real-time transmission power of the drone; () indicates the coordinates of the jammer's position; Indicates the jammer's transmission power; Indicates the strong interference coverage radius of the jammer; This represents the user interference state vector, where each element of the user interference state vector is identified by 1 or 0, indicating whether the corresponding user is within the strong coverage range of the jammer.
7. The anti-interference method for joint power-position optimization of UAV communication networks according to claim 1, characterized in that, The UAV moves in a grid space at fixed step sizes. The UAV position-power adjustment action space includes position movement actions and power adjustment actions. The position movement actions include up, down, left, right, and hovering. The power adjustment actions include increasing power, decreasing power, and maintaining power. The UAV position-power adjustment action space is constructed through joint optimization of position movement and power adjustment. Satisfies the expression: ; in, This indicates the number of combinations of position movement actions and power adjustment actions.
8. The anti-interference method for joint power-position optimization of UAV communication networks according to claim 2, characterized in that, The reward function is designed by guiding optimization with positive rewards and constraining unreasonable behavior with negative penalties. Satisfies the expression: Among them, in positive rewards, , , All are positive reward parameters. Basic connection rewards, , This represents the current number of connected users. Rewards for stable connections, , The number of users who maintain connection for two consecutive steps; Reconnection reward, , This represents the number of users who reconnected after a disconnection. In negative punishment , , , All are negative penalty parameters; Penalty for connection interruption, , This represents the number of users currently experiencing an interruption. Penalty for medium to strong interference zones. , The number of users whose services were disrupted while serving drones; Penalty for power usage , Reference power; To cover overlap penalty, , For drones i With drones k The dynamic average coverage radius, and drones i With drones k The dynamic coverage radius; Fixed weighting coefficients; I Indicates the total number of drones. U This represents the total number of users.
9. The anti-interference method for joint power-position optimization of UAV communication networks according to claim 1, characterized in that, The deep reinforcement learning uses a deep Q-network as the core decision network and adopts a dual-network design of a main network and a target network. The main network and the target network have the same structure but different parameter update mechanisms. The deep reinforcement learning utilizes the interaction between the agent and the environment. Intelligent agents acquire information through interaction with the environment. t Time-based drones - user multidimensional state space Then, standardization processing is performed to obtain a standardized UAV-user multidimensional state space. ; Utilizing the main network for standardized UAV-user multidimensional state space Drone position-power adjustment motion space After processing, the Q value is obtained, expressed as: ; Based on the Q-value and the improved ε-greedy strategy, select t Momentary drone position-power adjustment action space By executing t Momentary drone position-power adjustment action space Continue to interact with the environment and receive rewards. The improved ε-greedy strategy satisfies the expression: ,in, Number of steps executed. This refers to the decay step size parameter; The intelligent agent generates experience tuples from its interactions with the environment. Storing them in an experience pool, and using an adaptive sampling strategy to select learning experiences; among which... d This is the end marker; The target network calculates the target Q-value based on the learned experience. The target Q-value is the sum of the immediate reward and the future discounted reward, expressed as: ; Where γ is the discount factor, and at the terminal state d=1, the target Q value only includes the immediate reward, i.e. ; Minimize the target Q-value, the Q-value, and the loss function; update the main network parameters through gradient descent; and synchronize the target network parameters with the main network through a soft update mechanism. After training the deep Q-network, the final Q-value is obtained. Based on the final Q-value, the optimized drone transmission power, drone position, optimal drone movement at each moment, and the maximum number of users that can be stably connected are obtained.
10. An anti-interference system for joint power-position optimization of a UAV communication network, used to implement the anti-interference method for joint power-position optimization of a UAV communication network as described in any one of claims 1-9, characterized in that, include: A communication system construction module is used to construct an anti-jamming communication system for unmanned aerial vehicles (UAVs). The anti-jamming communication system for UAVs includes a jammer, a UAV, and a user. The UAV provides communication services to the user, and the jammer interferes with the user's communication. The optimization model building module is used to construct a power-position joint optimization model within a certain time period, with the objective function of maximizing the number of users that the UAV anti-interference communication system can stably connect to. This model is based on user-UAV connection state constraints, UAV allocable bandwidth constraints, UAV transmit power constraints, and UAV position constraints. The optimization model solving module is used to design the UAV-user multidimensional state space, UAV position-power adjustment action space, and reward function. Based on the UAV-user multidimensional state space, UAV position-power adjustment action space, and reward function, deep reinforcement learning is used to solve the power-position joint optimization model to obtain the optimized UAV transmit power, UAV position, optimal movement action of UAV at each time step, and the maximum number of users that can be stably connected.