Unmanned aerial vehicle relay communication method based on multi-branch attention deep reinforcement learning
By employing a multi-branch attention deep reinforcement learning method, an integrated optimization model was constructed, and a feature extraction structure and attention mechanism were designed. This solved the problems of high-dimensional state space and dynamic behavior constraints in UAV relay communication systems, thereby improving the system's coverage performance and throughput.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-16
- Publication Date
- 2026-04-10
AI Technical Summary
Unmanned aerial vehicle (UAV) relay communication systems present complex decision-making problems involving high-dimensional state spaces, dynamic behavioral constraints, and multi-objective trade-offs. Traditional methods struggle to adapt to dynamically changing scenarios, have high computational complexity, and are difficult to obtain feasible solutions within a limited timeframe.
We employ a multi-branch attention deep reinforcement learning method to construct an ensemble optimization model, design a multi-branch feature extraction structure and attention mechanism, and combine it with a priority experience replay strategy to optimize the decision-making process of the UAV relay communication system.
It significantly enhances the service coverage performance and total system throughput of the UAV relay communication system, improves decision accuracy and convergence stability, and enables dynamic focusing on nodes with unfulfilled communication needs and priority learning of high-value samples.
Smart Images

Figure CN121841431A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of communication technology, specifically to a UAV relay communication method based on multi-branch attention deep reinforcement learning. Background Technology
[0002] With the deep integration of wireless communication technology and IoT applications, the demand for scenario-based communication, such as communication coverage in remote areas and emergency disaster relief communication support, is becoming increasingly prominent. These scenarios typically face challenges such as lack of infrastructure, wide user distribution, and dynamically changing demands, placing higher requirements on the flexibility, real-time performance, and coverage capabilities of relay services. Compared to traditional fixed ground communication facilities such as base stations and fiber optic cables, drones, as relay nodes, offer advantages such as low deployment costs, high mobility, and rapid adaptation to complex terrain. They are widely regarded as a key component for achieving integrated air-ground coverage in next-generation wireless communication systems such as 5G-Advanced / 6G. Against this backdrop, optimizing drone relay communication technology is of significant practical importance for improving system performance and service capabilities.
[0003] The high dimensionality, dynamic nature, and non-convexity of the aforementioned problems lead to an exponential increase in computational complexity with system size, making it difficult to obtain a feasible solution within a finite timeframe. Secondly, the dynamic changes in user distribution and channel states in typical communication environments require algorithms to have real-time adjustment capabilities, while traditional methods, relying on fixed model parameters and complete environmental information, struggle to adapt to dynamically changing scenarios. Therefore, Deep Reinforcement Learning (DRL) algorithms, with their advantages of model freedom and online learning mechanisms, have become an effective approach to address these challenges. DRL does not rely on precise system modeling but autonomously perceives dynamic patterns such as user demand distribution and channel fading through interaction with the environment, and gradually learns the optimal strategy based on a reward feedback mechanism, aiming to maximize long-term cumulative rewards. Therefore, DRL demonstrates significant advantages and applicability in handling complex decision-making problems in UAV relay communication systems, including high-dimensional state spaces, dynamic behavioral constraints, and multi-objective trade-offs. Summary of the Invention
[0004] The technical problem to be solved by this invention is to provide a UAV relay communication method based on multi-branch attention deep reinforcement learning, which can handle complex decision-making problems such as high-dimensional state space, dynamic behavioral constraints and multi-objective trade-offs in UAV relay communication systems.
[0005] To solve the above-mentioned technical problems, the technical solution adopted by the present invention is as follows: A multi-branch attention-based deep reinforcement learning method for drone relay communication includes the following steps: Step 1: Construct an integrated optimization model for the UAV relay communication system, specifically including the UAV flight and energy consumption model, communication model, and integrated optimization model; Step 2: Model the optimization problem in Step 1 as a Markov decision process to adapt to the input requirements of the subsequent deep reinforcement learning (DRL) algorithm. This includes state design, action design, and reward function design. Step 3: Design a multi-branch feature extraction structure to independently encode information such as UAV location, user location, and communication requirements, thereby achieving feature decoupling from the source and effectively improving the extraction capability and optimization efficiency of state features. Step 4: Propose an attention mechanism and a priority experience replay strategy to achieve dynamic focusing on nodes with incomplete communication needs and priority learning of high-value samples. Step 5: Propose a multi-branch attention deep reinforcement learning algorithm to optimize and solve the model established in Step 1.
[0006] The specific process of establishing the UAV flight and energy consumption model in step 1 above is as follows: Step 1.1: State update and decision execution are performed using discrete time slots; analysis shows that the UAV's motion follows a Markov decision process, which occurs in time slots. t The +1 state depends only on the time slot. t State actions; define time slots t The location of the drone at that time was Its action space is , indicating in time slot t Inner along x and y The displacement in the direction; therefore, the UAV position update design is as follows: (1) In addition, the power consumption of the UAV during propulsion comes from three parts, specifically including rotor blade profile power, induced power, and parasitic power; therefore, the total power consumption of the UAV can be modeled as follows: (2) In the formula, t For time slots, P 0、 P s These represent the blade profile power and induced power when the drone is hovering, respectively. v 1 represents the flight speed of the drone. U r For rotor tip speed, v 0 represents the rotor design parameter. d 0、 s , ρ a , A and are the fuselage drag coefficient, rotor solidity, air density, and rotor disk area, respectively.
[0007] The specific process of establishing the communication model in step 1 above is as follows: Step 1.2: Use the Ricean fading model to characterize the channel characteristics, and the channel coefficient between the base station and the UAV. It can be represented as: (3) In the formula, α B-U This is the path loss index. d B-U The distance between the base station and the drone. ρ For reference distance loss coefficient; K 1 is the Rice factor of the BU link, which represents the power ratio of the line-of-sight component to the non-line-of-sight component; and These are the Los component and the NLoS transmission component in the channel, respectively. The elements of are independent and identically distributed random variables that follow a complex Gaussian distribution with a mean of 0 and a variance of 1. Determined by the array response vectors of the base station and the drone: (4) In the formula, and These are the array response vectors for the base station and the drone, respectively. for The conjugate transpose of . ϕ and θ These represent the departure angle and arrival angle of the signal, respectively. Similarly, the UAV-user link channel It can be modeled as: (5) In the formula, α U-M This is the path loss index. d U-M The distance between the drone and the user. ρ For reference distance loss coefficient; K 2 is the Rice factor of the UM link; The drone uses a gain amplification relay protocol, and its gain G The design ensures that the UAV's transmit power meets the following constraints: (6) In the formula, P u and P b These represent the transmission power of the drone and the base station, respectively. For receiving noise power for drones; In each time slot, the UAV provides relay services to users within its communication coverage area using orthogonal frequency division multiple access (OFDM) technology. The total system bandwidth B is divided into multiple orthogonal sub-channels. Due to spectrum resource constraints, the number of users the UAV can serve in each time slot is determined by the number of available orthogonal sub-channels. n Users in each time slot k Received signal-to-noise ratio It can be represented as: (7) In the formula, The noise power received by user k; in For the sub-bandwidth, according to Shannon's formula, the first... k A user in time slot n Transmission rate at time for: ;(8).
[0008] The specific process of establishing the integrated optimization model in step 1 above is as follows: Step 1.3, the optimization objective can be formally expressed as: (9) (10) (11) (12) (13) Regarding the above optimization objectives and constraints, the specific meanings of each symbol and formula can be further clarified: M Represents the total number of user nodes. N The total number of time slots in the scheduling cycle; Indicates time slot t Internal users i Communication rate, It is the set of UAV position offsets in each time slot, and the objective function (9) is obtained through the characteristic function. (Take 1 if the requirement is met, otherwise take 0) Maximize the fulfillment of preset communication requirements. The number of users; the objective function (10) corresponds to maximizing the total communication rate of the system; in constraint (11), and Let be the position coordinates of the UAV in time slot t. and They are respectively in X , Y The positional boundary of the dimension; constraint (12) , It is the maximum displacement limit of the UAV within a single time slot; constraint (13) ensures the distance between the UAV and the user through the distance formula. i The distance does not exceed the effective communication radius. R This is to ensure service accessibility.
[0009] The specific process of state design in step 2 above is as follows: Step 2.1, the state design is as follows: (14) In the formula, x U [ t ]and y U [ t ] respectively represent the drone in t The horizontal and vertical axes of the time slot; I for M Service status of each user node; Ψ The set of locations for all user nodes; this design integrates drone status with... M The complete information of each user node is integrated to form a 3-dimensional... M The +2 state space provides a sufficient environmental perception basis for decision-making.
[0010] The specific process of motion design in step 2 above is as follows: Step 2.2, the specific action space is represented as follows: (15) In the formula, x t and y t They represent t The continuous displacement of the time-slotted UAV in the horizontal and vertical directions; by directly outputting displacement commands, the intelligent agent can achieve smooth trajectory planning and realize dynamic coverage of high-priority users.
[0011] The specific process of designing the reward function in step 2 above is as follows: Step 2.3: The reward function, by quantifying and integrating key system performance indicators, provides fine-grained optimization guidance for policy learning. Its expression is as follows: (16) ; In the formula, The coverage reward represents the number of users within the drone's coverage area in the current time slot, directly incentivizing the expansion of service coverage. The throughput reward is calculated based on the total data transmission volume of the system within the current time slot, which directly promotes the improvement of spectrum resource utilization efficiency and data transmission capability. The number of users served in that time slot, in order to accelerate the overall service process. , , These are the corresponding weighting coefficients; It is a signal variable; it equals 1 when all users have been served, and 0 otherwise. K com It is a reward given after serving all users. For crossing the boundary, Penalty for moving to an empty space. Penalty for delays in progress, These are the weight coefficients corresponding to each penalty item, constraining invalid actions to ensure the effectiveness of the strategy.
[0012] The specific process of step 3 above is as follows: A multi-branch feature extraction structure is proposed. By adopting a parallel branch architecture, the original state space is deconstructed into three feature subspaces, each establishing an independent feature processing path. Each branch adopts a unified network structure, wherein... The branch index is represented as follows: (17) Equation (17) is the implementation carrier of the initial transformation for feature extraction, in which the linear transformation step is composed of the weight matrix. With bias vector Collaborative completion— Responsible for processing the raw input Mapping to a high-dimensional feature space expands the representational dimension of features to enhance their representational power; This is used to compensate for feature shifts after linear transformation, making the transformed feature distribution more suitable for the processing needs of subsequent networks. The subsequent layer normalization operation, by re-centering and scaling the linear transformation result, effectively eliminates the adverse interference of differences in physical dimensions and distributions between the original inputs on the training process, providing more stable feature inputs for subsequent network layers. Based on this, a nonlinear transformation needs to be introduced into the normalized features to further enhance their nonlinear representation capability. (18) Equation (18) provides the network with the necessary nonlinear mapping capability, enabling the model to learn and fit complex feature relationships; at the same time, by setting the output of some neurons to zero, the sparsity representation of features is achieved, which helps to filter noise information and enhance the discriminativeness of features; finally, the nonlinear features are refined and their dimensions are compressed to output the final encoding of this branch: (19) In the above processing, the layer normalization operation is specifically defined as follows: (20) In the formula, and These are the input vectors. x The mean and standard deviation, through Achieve feature standardization; Scaling factor For offset; through learnable parameters and Affine transformations are performed to enable the network to adaptively adjust the normalized feature distribution according to its own needs, while maintaining its expressive power.
[0013] The specific process of step 4 above is as follows: A scaling dot product-based attention mechanism is designed to achieve accurate identification and adaptive resource allocation of key service targets by establishing a dynamic interaction model between UAVs and user nodes. The mechanism uses UAV state features extracted by the front-end multi-branch feature module as the query vector, which deeply encodes the UAV's real-time location information and service capability status. Simultaneously, the integrated user node features are used as keys and values, where each node feature contains its precise geographical location and real-time communication requirement status. To achieve effective interaction modeling, the system uses three sets of independent learnable weight matrices to map the query vector and key-value features to the same metric space, as detailed below: ;(twenty one) ;(twenty two) ;(twenty three) This mapping operation transforms the query features representing drone status and the key-value features representing user group status into a unified vector space; through independent linear transformations. W Q , W K and W V This enables features from different modalities to be compared for similarity in the same metric space, laying the foundation for subsequent correlation calculations; Furthermore, the attention weights are calculated using scaled dot product attention as follows: ;(twenty four) In the formula, d kThe dimension representing the key vector is used to adjust the order of magnitude of the dot product result, preventing excessively large dot product values due to high vector dimensions, which could cause the softmax function to enter the gradient saturation region and affect the model's learning stability. The essence of this calculation process is to establish a correlation measurement mechanism between the drone's state and the features of each user node. By performing the dot product operation between the query vector and the key vector, a correlation score is generated for each user node, accurately reflecting the node's importance at the current decision moment. The softmax function converts these scores into a normalized probability distribution, ensuring that the sum of the attention weights of all user nodes is 1, forming a reasonable weight allocation. Finally, the algorithm uses the calculated attention weights... A value vector V We perform weighted fusion to generate a refined representation of user group characteristics: (25) The weighted summation operation enables adaptive aggregation of information, generating feature representations. H att In this process, the feature information of users identified as key users is enhanced, while the feature information of non-key users is suppressed. This dynamic weight allocation mechanism provides the policy network with feature mappings that have priority hints, guiding the UAV to focus on high-value service targets during the decision-making process.
[0014] The specific process of step 5 above is as follows: The methods in steps 3 and 4 are integrated into the Deep Deterministic Policy Gradient (DDPG) algorithm to form a multi-branch attention deep reinforcement learning algorithm, which specifically includes: Step 5.1, Experience Replay and Sampling Mechanism; Samples generated by the interaction between the agent and the environment. The samples are stored in an experience replay pool that is prioritized. Based on the temporal difference error (TIF), the algorithm assigns a priority to each sample and performs non-uniform sampling accordingly. Samples with high TIF contain more unlearned environmental dynamics and are therefore given a higher sampling probability. To overcome the bias that priority sampling may introduce, the algorithm combines importance sampling weights to correct gradient updates. This priority experience replay mechanism effectively improves the utilization efficiency of training data by guiding the agent to focus on samples with high learning value, thereby accelerating the policy convergence process as a whole. Step 5.2, Target value calculation and network update; The calculation of the target Q value is based on the Bellman equation expansion: (26) In the formula, Represents an instant reward. This is a discount factor, with a value ranging from 0 to 1, used to balance the importance of immediate rewards and future cumulative rewards; d t This is the round end marker; and These are the outputs of the target value network and the target policy network, respectively. Their synergistic effect provides a stable benchmark reference for the calculation of the target value. The update of the value network aims to minimize the mean squared error, and its loss function is defined as: (27) This loss function introduces importance sampling weights. w i This effectively corrects for distribution biases that may be caused by priority sampling. Parameters are updated using gradient descent. Value networks can gradually improve their prediction accuracy, making the Q-value estimate continuously approach the true cumulative return expectation, thus providing a reliable evaluation basis for strategy optimization. Furthermore, the policy network update mechanism is based on the deterministic policy gradient theorem, and its update direction is guided by the gradient information of the value network: (28) In the formula F For sample batches, this gradient calculation process demonstrates the tight coupling between the value network and the policy network: Value Network It provides directions for improving the action space, indicating which actions will yield higher long-term returns; policy network These directional guidelines are then translated into actual adjustments to the network parameters. Through this chain-like gradient propagation, the policy network can continuously optimize its decision-making capabilities, generate superior drone displacement actions, and ultimately maximize long-term cumulative returns. Step 5.3: Target network soft update; The target network parameters are tracked using a soft update method to track the main network. (29) In the formula, This is the soft update coefficient.
[0015] This invention presents a multi-branch attention-based deep reinforcement learning method for UAV relay communication. First, it constructs an integrated optimization model for the UAV relay communication system, laying the foundation for improved overall system performance. Second, it models the optimization problem as a Markov decision process, adapting it to the input requirements of deep reinforcement learning algorithms. Then, it designs a multi-branch feature extraction structure, independently encoding information such as UAV location, user location, and communication needs, achieving feature decoupling from the source and effectively enhancing the extraction capability and optimization efficiency of state features. Furthermore, it proposes a strategy that integrates an attention mechanism and priority experience replay, enabling dynamic focusing on nodes with incomplete communication needs and prioritizing the learning of high-value samples, synergistically improving the algorithm's decision accuracy and convergence stability. Finally, it organically integrates the multi-branch feature extraction structure and the attention-priority replay mechanism into the Deep Deterministic Policy Gradient (DDPG) algorithm, forming a multi-branch attention-based deep reinforcement learning algorithm, which ultimately significantly enhances the service coverage performance and total throughput of the UAV relay communication system. This invention not only provides new ideas for UAV relay communication network optimization research but also expands the depth and breadth of the DDPG algorithm's application in this field. Attached Figure Description
[0016] The present invention will be further described below with reference to the accompanying drawings and embodiments: Figure 1 This is a schematic diagram of the model of the present invention based on UAV relay communication; Figure 2 This is a schematic diagram of the MBA-DDPG framework for the multi-branch attention depth deterministic strategy gradient algorithm in this invention; Figure 3 This is a schematic diagram of the reward function when the number of users is 30 in an embodiment of the present invention; Figure 4 This is a schematic diagram of the reward function when the number of users is 50 in an embodiment of the present invention; Figure 5 This is a schematic diagram of the reward function when the number of users is 70 in an embodiment of the present invention; Figure 6 This refers to the average throughput of the algorithm under different user scales in this embodiment of the invention. Figure 7 This is a drone trajectory diagram when the number of users is 30 in this embodiment of the invention; Figure 8 This is a drone trajectory diagram when the number of users is 50 in this embodiment of the invention; Figure 9 This is a drone trajectory diagram when the number of users is 90 in this embodiment of the invention. Detailed Implementation
[0017] The technical solution of the present invention will be described in detail below with reference to the accompanying drawings and embodiments.
[0018] A multi-branch attention-based deep reinforcement learning method for drone relay communication includes the following steps: Step 1: Construct an integrated optimization model for the UAV relay communication system, specifically including the UAV flight and energy consumption model, communication model, and integrated optimization model; Step 2: Model the optimization problem in Step 1 as a Markov decision process to adapt to the input requirements of the subsequent deep reinforcement learning (DRL) algorithm. This includes state design, action design, and reward function design. Step 3: Design a multi-branch feature extraction structure to independently encode information such as UAV location, user location, and communication requirements, thereby achieving feature decoupling from the source and effectively improving the extraction capability and optimization efficiency of state features. Step 4: Propose an attention mechanism and a priority experience replay strategy to achieve dynamic focusing on nodes with incomplete communication needs and priority learning of high-value samples. Step 5: Propose a multi-branch attention deep reinforcement learning algorithm to optimize and solve the model established in Step 1.
[0019] The specific process of establishing the UAV flight and energy consumption model in step 1 above is as follows: Step 1.1: In relay communication tasks, the mobility and endurance of the UAV are key factors determining the system's service efficiency. Therefore, it is necessary to establish an accurate UAV flight and energy consumption model to provide a decision-making environment that conforms to actual physical laws for the subsequent DRL algorithm. Considering the discrete decision-making characteristics of the DRL algorithm, discrete time slots are used for state updates and decision execution. Analysis shows that the UAV's motion follows a Markov decision process, and its time slots... t The +1 state depends only on the time slot. t State actions; define time slots t The location of the drone at that time was Its action space is , indicating in time slot t Inner along x and y The displacement in the direction; therefore, the UAV position update design is as follows: (1) In addition, the power consumption of the UAV during propulsion mainly comes from three parts, specifically including rotor blade profile power, induced power, and parasitic power; therefore, the total power consumption of the UAV can be modeled as follows: (2) In the formula, t For time slots, P 0、 P sThese represent the blade profile power and induced power when the drone is hovering, respectively. v 1 represents the flight speed of the drone. U r For rotor tip speed, v 0 represents the rotor design parameter. d 0、 s , ρ a , A and are the fuselage drag coefficient, rotor solidity, air density, and rotor disk area, respectively.
[0020] The specific process of establishing the communication model in step 1 above is as follows: Step 1.2: The reliability of the communication link directly determines the service capability of the UAV. Considering that a stable line-of-sight propagation main path usually exists between the UAV and the ground node when flying at low altitudes, accompanied by multipath scattering effects, its channel characteristics are highly consistent with the Ricean fading model. Therefore, the Ricean fading model is used to characterize the channel characteristics, and the channel coefficient between the base station and the UAV... It can be represented as: (3) In the formula, α B-U This is the path loss index. d B-U The distance between the base station and the drone. ρ For reference distance loss coefficient; K 1 is the Rice factor of the BU link, which represents the power ratio of the line-of-sight component to the non-line-of-sight component; and These are the Los component and the NLoS transmission component in the channel, respectively. The elements of are independent and identically distributed random variables that follow a complex Gaussian distribution with a mean of 0 and a variance of 1. Determined by the array response vectors of the base station and the drone: (4) In the formula, and These are the array response vectors for the base station and the drone, respectively. for The conjugate transpose of . ϕ and θ These represent the departure angle and arrival angle of the signal, respectively. Similarly, the UAV-user link channel It can be modeled as: (5) In the formula, α U-M This is the path loss index.d U-M The distance between the drone and the user. ρ For reference distance loss coefficient; K 2 is the Rice factor of the UM link; The drone uses a gain amplification relay protocol, and its gain G The design ensures that the UAV's transmit power meets the following constraints: (6) In the formula, P u and P b These represent the transmission power of the drone and the base station, respectively. For receiving noise power for drones; In each time slot, the UAV provides relay services to users within its communication coverage area using orthogonal frequency division multiple access (OFDM) technology. The total system bandwidth B is divided into multiple orthogonal sub-channels. Due to spectrum resource constraints, the number of users the UAV can serve in each time slot is determined by the number of available orthogonal sub-channels. n Users in each time slot k Received signal-to-noise ratio It can be represented as: (7) In the formula, The noise power received by user k; in For the sub-bandwidth, according to Shannon's formula, the first... k A user in time slot n Transmission rate at time for: ;(8).
[0021] The specific process of establishing the integrated optimization model in step 1 above is as follows: Step 1.3: In the UAV relay communication system, the core optimization objective is to maximize the number of user nodes served and the total system throughput while satisfying system constraints by dynamically adjusting the UAV's movement trajectory and resource allocation. This optimization objective can be formally expressed as: (9) (10) (11) (12) (13) Regarding the above optimization objectives and constraints, the specific meanings of each symbol and formula can be further clarified: M Represents the total number of user nodes.N The total number of time slots in the scheduling cycle; Indicates time slot t Internal users i Communication rate, It is the set of UAV position offsets in each time slot, and the objective function (9) is obtained through the characteristic function. (Take 1 if the requirement is met, otherwise take 0) Maximize the fulfillment of preset communication requirements. The number of users; the objective function (10) corresponds to maximizing the total communication rate of the system; in constraint (11), and Let be the position coordinates of the UAV in time slot t. and They are respectively in X , Y The positional boundary of the dimension; constraint (12) , It is the maximum displacement limit of the UAV within a single time slot; constraint (13) ensures the distance between the UAV and the user through the distance formula. i The distance does not exceed the effective communication radius. R This is to ensure service accessibility.
[0022] The specific process of state design in step 2 above is as follows: Step 2.1: To achieve effective mobility decision-making, the UAV needs to comprehensively perceive the environmental state. The state space needs to include key information such as the UAV's own position and the distribution of user nodes to support the coordinated optimization of path planning and resource scheduling. Therefore, the state is designed as follows: (14) In the formula, x U [ t ]and y U [ t ] respectively represent the drone in t The horizontal and vertical axes of the time slot; I for M Service status of each user node; Ψ The set of locations for all user nodes; this design integrates drone status with... M The complete information of each user node is integrated to form a 3-dimensional... M The +2 state space provides a sufficient environmental perception basis for decision-making.
[0023] The specific process of motion design in step 2 above is as follows: Step 2.2: To achieve precise trajectory control in continuous space, the UAV's actions are directly defined as position offsets. This design avoids the performance loss caused by discretization while conforming to the actual motion constraints of the UAV. The specific action space is represented as follows: (15) In the formula, x t and y t They represent t The continuous displacement of the time-slotted UAV in the horizontal and vertical directions; by directly outputting displacement commands, the intelligent agent can achieve smooth trajectory planning and realize dynamic coverage of high-priority users.
[0024] The specific process of designing the reward function in step 2 above is as follows: Step 2.3: To guide UAVs in achieving multi-objective collaborative optimization of service coverage, transmission efficiency, and service progress in a dynamic communication environment, a reward function with clear guidance needs to be designed. The reward function constructed in this invention provides fine-grained optimization guidance for policy learning by quantifying and fusing key system performance indicators. Its expression is as follows: (16) ; In the formula, The coverage reward represents the number of users within the drone's coverage area in the current time slot, directly incentivizing the expansion of service coverage. The throughput reward is calculated based on the total data transmission volume of the system within the current time slot, which directly promotes the improvement of spectrum resource utilization efficiency and data transmission capability. The number of users served in that time slot, in order to accelerate the overall service process. , , These are the corresponding weighting coefficients; It is a signal variable; it equals 1 when all users have been served, and 0 otherwise. K com It is a reward given after serving all users. For crossing the boundary, Penalty for moving to an empty space. Penalty for delays in progress, These are the weight coefficients corresponding to each penalty item, constraining invalid actions to ensure the effectiveness of the strategy.
[0025] This reward function, through a mechanism combining positive and negative feedback, not only provides a clear and quantifiable optimization direction for UAV policy learning, but also establishes an effective trade-off mechanism among multiple objectives, thereby systematically guiding the agent to gradually approach the globally optimal strategy that balances coverage, efficiency, and progress.
[0026] The specific process of step 3 above is as follows: The state space of an unmanned aerial vehicle (UAV) relay communication system contains three heterogeneous types of information with distinct physical meanings and data structures: UAV position, user state, and user location. Traditional DRL methods directly concatenate these features using a single network, leading to severe feature coupling and information interference, significantly reducing the discriminative power of the state representation. To address this, a multi-branch feature extraction structure is proposed. By employing a parallel branch architecture, the original state space is deconstructed into three feature subspaces, each establishing an independent feature processing path. Each branch adopts a unified network structure, where… The branch index is represented as follows: (17) Equation (17) is the implementation carrier of the initial transformation for feature extraction, in which the linear transformation step is composed of the weight matrix. With bias vector Collaborative completion— Responsible for processing the raw input Mapping to a high-dimensional feature space expands the representational dimension of features to enhance their representational power; This is used to compensate for feature shifts after linear transformation, making the transformed feature distribution more suitable for the processing needs of subsequent networks. The subsequent layer normalization operation, by re-centering and scaling the linear transformation result, effectively eliminates the adverse interference of differences in physical dimensions and distributions between the original inputs on the training process, providing more stable feature inputs for subsequent network layers. Based on this, a nonlinear transformation needs to be introduced into the normalized features to further enhance their nonlinear representation capability. (18) Equation (18) primarily provides the network with the necessary nonlinear mapping capability, enabling the model to learn and fit complex feature relationships; simultaneously, by setting the output of some neurons to zero, sparse representation of features is achieved, which helps filter noise information and enhance the discriminative power of features; finally, the nonlinear features are refined and their dimensions compressed to output the final encoding of this branch: (19) In the above processing, the layer normalization operation is specifically defined as follows: (20) In the formula, and These are the input vectors. x The mean and standard deviation, through Achieve feature standardization; Scaling factor For offset; through learnable parameters and Affine transformations are performed to enable the network to adaptively adjust the normalized feature distribution according to its own needs, while maintaining its expressive power.
[0027] The specific process of step 4 above is as follows: Building upon the successful decoupling of heterogeneous features in the multi-branch structure, the system still needs to address another key issue: how to dynamically identify and prioritize services for users most critical to improving overall system performance from a widely distributed pool of user nodes, under constraints of limited communication resources and UAV endurance. To address this, this invention designs an attention mechanism based on scaled dot product. By establishing a dynamic interaction model between UAVs and user nodes, it achieves accurate identification and adaptive resource allocation for key service targets. The mechanism uses UAV state features extracted by the front-end multi-branch feature module as the query vector, which deeply encodes the UAV's real-time location information and service capability status. Simultaneously, the integrated user node features are used as keys and values, where each node feature contains its precise geographical location and real-time communication demand status. To achieve effective interaction modeling, the system uses three independent learnable weight matrices to map the query vector and key-value features to the same metric space, as detailed below: ;(twenty one) ;(twenty two) ;(twenty three) This mapping operation transforms the query features representing drone status and the key-value features representing user group status into a unified vector space; through independent linear transformations. W Q , W K and W V This enables features from different modalities to be compared for similarity in the same metric space, laying the foundation for subsequent correlation calculations; Furthermore, the attention weights are calculated using scaled dot product attention as follows: ;(twenty four) In the formula, d kThe dimension representing the key vector is used to adjust the order of magnitude of the dot product result, preventing excessively large dot product values due to high vector dimensions, which could cause the softmax function to enter the gradient saturation region and affect the model's learning stability. The essence of this calculation process is to establish a correlation measurement mechanism between the drone's state and the features of each user node. By performing the dot product operation between the query vector and the key vector, a correlation score is generated for each user node, accurately reflecting the node's importance at the current decision moment. The softmax function converts these scores into a normalized probability distribution, ensuring that the sum of the attention weights of all user nodes is 1, forming a reasonable weight allocation. Finally, the algorithm uses the calculated attention weights... A value vector V We perform weighted fusion to generate a refined representation of user group characteristics: (25) The weighted summation operation enables adaptive aggregation of information, generating feature representations. H att In this process, the feature information of users identified as key users is enhanced, while the feature information of non-key users is suppressed. This dynamic weight allocation mechanism provides the policy network with feature mappings that have priority hints, guiding the UAV to focus on high-value service targets during the decision-making process.
[0028] The specific process of step 5 above is as follows: The methods in steps 3 and 4 are integrated into the Deep Deterministic Policy Gradient (DDPG) algorithm to form a multi-branch attention deep reinforcement learning algorithm, which specifically includes: Step 5.1, Experience Replay and Sampling Mechanism; Samples generated by the interaction between the agent and the environment. (in The samples (marked as the end of the round) are stored in an experience replay pool ordered by priority. Unlike traditional random sampling mechanisms, this algorithm assigns priority to samples based on their Temporal-Difference Error (TD) and performs non-uniform sampling accordingly. Samples with high TD errors typically contain more unlearned environmental dynamics and are therefore given higher sampling probabilities. To overcome the bias that priority sampling may introduce, the algorithm combines importance sampling weights to correct gradient updates. This priority experience replay mechanism effectively improves the utilization efficiency of training data by guiding the agent to focus on samples with high learning value, thereby accelerating the policy convergence process as a whole. Step 5.2, Target value calculation and network update; The calculation of the target Q value is based on the Bellman equation expansion: (26) In the formula, Represents an instant reward. This is a discount factor, with a value ranging from 0 to 1, used to balance the importance of immediate rewards and future cumulative rewards; d t This is the round end marker; and These are the outputs of the target value network and the target policy network, respectively. Their synergistic effect provides a stable benchmark reference for the calculation of the target value. The update of the value network aims to minimize the mean squared error, and its loss function is defined as: (27) This loss function introduces importance sampling weights. w i This effectively corrects for distribution biases that may be caused by priority sampling. Parameters are updated using gradient descent. Value networks can gradually improve their prediction accuracy, making the Q-value estimate continuously approach the true cumulative return expectation, thus providing a reliable evaluation basis for strategy optimization. Furthermore, the policy network update mechanism is based on the deterministic policy gradient theorem, and its update direction is guided by the gradient information of the value network: (28) In the formula F For sample batches, this gradient calculation process demonstrates the tight coupling between the value network and the policy network: Value Network It provides directions for improving the action space, indicating which actions will yield higher long-term returns; policy network These directional guidelines are then translated into actual adjustments to the network parameters. Through this chain-like gradient propagation, the policy network can continuously optimize its decision-making capabilities, generate superior drone displacement actions, and ultimately maximize long-term cumulative returns. Step 5.3: Target network soft update; The target network parameters are tracked using a soft update method to track the main network. (29) In the formula, The soft update coefficients are used; this gradual parameter update strategy effectively suppresses drastic fluctuations in the target value, providing an important guarantee for the stable convergence of the algorithm in high-dimensional state space and continuous action domain.
[0029] Example 1: A multi-branch attention-based deep reinforcement learning method for drone relay communication includes the following steps: Step 1: Construct an integrated optimization model for the UAV relay communication system; Step 2: Model the optimization problem in Step 1 as a Markov decision process; Step 3: Design a multi-branch feature extraction structure to independently encode information such as UAV location, user location, and communication requirements; Step 4 proposes an attention mechanism and a priority experience replay strategy to achieve dynamic focusing on nodes with incomplete communication needs and priority learning of high-value samples. Step 5: Propose a multi-branch attention deep reinforcement learning algorithm to optimize and solve the model established in Step 1.
[0030] Specifically, it includes: Step 1 includes the following steps: Step 1.1: Establish a UAV flight and energy consumption model; Step 1.2, establish the communication model; Step 1.3: Establish an integrated optimization model.
[0031] Step 2 includes the following steps: Step 2.1, State Design; Step 2.2, Motion Design; Step 2.3, Design of reward function.
[0032] Step 3 includes the following steps: Independent encoding of information such as drone location, user location, and communication requirements Step 4 includes the following steps: Dynamically focusing on nodes with incomplete communication needs, and prioritizing the learning of high-value samples. Step 5 includes the following steps: Step 5.1, Experience Replay and Sampling Mechanism; Step 5.2, Target value calculation and network update; Step 5.3, target network soft update.
[0033] MBA-DDPG algorithm: 1. Initialize the MBA-DDPG core components: Building a policy network It integrates a multi-branch feature extraction structure and an attention mechanism to output continuous displacement actions of UAVs; and constructs a value network. A multi-branch feature encoding method is used to evaluate the value of state-action pairs; copying , The parameters are used to initialize the target policy network. and target value network Initialize the priority experience replay pool Used to store interaction samples .
[0034] 2. Begin the round cycle, total number of rounds: :
[0035] 3. Reset environment status: Reset information such as drone location, user requirements, and location. 4. Enter the time slot loop; the maximum number of time slots per round is... : The round has not ended (the user has not been fully served or the time slot has not been exhausted). 5. Observe the environmental conditions And separate heterogeneous features: obtain the location of the UAV User status User location .
[0036] 6. Multi-branch feature encoding and attention enhancement: Drone location branch: Encoding, extracting low-dimensional continuous feature representations User state branch: and Encode, using attention mechanisms to reinforce key user features, output and ; Fusion characteristics: splicing , , To obtain a unified state express.
[0037] 7. Policy network selects actions: ( (Gaussian noise, used for action space exploration) 8. Perform actions and interact with the environment, calculating immediate rewards. Observe the next state Repeat steps 6-7 to perform feature separation, encoding, and fusion.
[0038] 9. Storing Samples: Store in the priority experience replay pool .
[0039] 10. From Random sampling Sample Used to update the value network With policy network ; 11. Soft update target network: ',
[0040] 12. End condition judgment: If the end condition of the round is met, exit the time slot loop.
[0041] 13. End condition judgment: If the total round end condition is met, exit the round loop.
[0042] Simulation verification and analysis: To verify the performance of the MBA-DDPG algorithm in UAV relay communication scenarios, this chapter establishes a complete simulation experimental environment. The experiment is conducted within a 100m × 100m two-dimensional planar area, with the UAV's initial position coordinates at (0,0) and the ground base station deployed at (40m, 100m). The system operation cycle is divided into 70 equal-length time slots, with a maximum displacement limit of 10m for the UAV in a single time slot. Each experimental case is run independently for 1000 training rounds to obtain stable statistical results. The experimental platform is built based on Python 3.9 and PyTorch 2.8.0 frameworks, using an Intel i7-13700 processor and an NVIDIA RTX 5060 GPU. Table 1 details the key simulation parameters of the UAV relay communication system, including core settings such as flight area configuration, communication model parameters, and energy consumption coefficients. This ensures the rationality and comparability of the experimental design. The system parameters are shown in the table below: Table 1 System Parameters
[0043] Results Analysis: Comparative experiments were conducted between the proposed method MBA-DDPG and mainstream methods DDPG, SAC, and TD3. The results show that MBA-DDPG achieves source decoupling of heterogeneous state information through a multi-branch feature extraction structure, enabling the algorithm to accurately process features of different modalities such as UAV location, user distribution, and communication needs. Simultaneously, the introduction of an attention mechanism gives the algorithm the ability to dynamically focus on key nodes, achieving optimal scheduling decisions under limited resource constraints. This synergistic mechanism of accurate perception and dynamic focusing allows MBA-DDPG to establish high-quality state representations in the early stages of training, manifested as a rapid rise in the reward curve and a significant improvement in average throughput.
[0044] Optimized reward function such as Figure 3-5 As shown in the figure, the horizontal axis represents the training rounds, and the vertical axis represents the reward function value.
[0045] Optimized average throughput curve as follows Figure 6 As shown, the horizontal axis represents the user base, and the vertical axis represents the average throughput.
[0046] As shown in the figure, the performance advantages exhibited by MBA-DDPG are fully validated in comparison with other algorithms. While the TD3 algorithm alleviates the overestimation problem to some extent through its double-Q learning mechanism, its basic network structure exhibits significant limitations when handling high-dimensional heterogeneous state information—insufficient representational ability caused by feature coupling leads to continuous fluctuations in the policy learning process, making it difficult to break through suboptimal convergence levels. Although the SAC algorithm's maximum entropy-based exploration framework theoretically helps enhance policy diversity, its single-network structure's feature extraction efficiency bottleneck is exposed in the high-dimensional state space and sparse reward environment encountered in this study, resulting in the algorithm's inability to effectively identify key state features and ultimately significantly limiting its performance. The performance of the basic DDPG algorithm, on the other hand, highlights the fundamental defects of traditional structures: the combined effect of feature coupling and value estimation bias leads to slow policy convergence and limited performance. This performance difference fully demonstrates the necessity of the MBA-DDPG algorithm architecture design.
[0047] In contrast, MBA-DDPG achieves source decoupling of heterogeneous state information through a multi-branch feature extraction structure, enabling the algorithm to accurately process features from different modalities such as UAV location, user distribution, and communication needs. Simultaneously, the introduction of an attention mechanism gives the algorithm the ability to dynamically focus on key nodes, achieving optimal scheduling decisions under limited resource constraints. This synergistic mechanism of "precise perception" and "dynamic focusing" allows MBA-DDPG to establish high-quality state representations (manifested as a rapid rise in the reward curve) in the early stages of training, and translates them into system-level performance. Figure 6 The significantly improved average throughput performance is shown. The performance limitations of each comparative algorithm in different dimensions collectively demonstrate the effectiveness and advancement of the MBA-DDPG algorithm architecture in solving high-dimensional heterogeneous state space problems.
[0048] from Figure 7-9 The drone trajectories shown demonstrate that the MBA-DDPG algorithm can generate flight paths that highly match user spatial distribution and service needs across different user scales, directly reflecting the effectiveness of its internal algorithm mechanism. The multi-branch feature extraction structure independently encodes heterogeneous information such as drone location and user distribution, providing the algorithm with accurate environmental perception capabilities, enabling it to construct macroscopic paths covering densely populated user areas. Simultaneously, the attention mechanism, through dynamic weight allocation of user node features, achieves priority scheduling and decision guidance for key nodes, thereby allowing for fine-grained local adjustments based on the macroscopic path.
[0049] Therefore, in a 30-user scenario, the targeted detours of the trajectory stem from continuous attention to key nodes; in a 50-user scenario, the multi-regional linkage of the trajectory demonstrates the algorithm's ability to plan the overall spatial distribution; and in a high-density scenario with 70 users, the efficiency and coherence of the trajectory prove that the architecture can effectively cope with the challenges brought about by the increase in the dimensionality of the state space. This trajectory, generated by the combined effect of accurate environmental perception and dynamic decision guidance, achieves comprehensive coverage in physical space and high efficiency in service timing, fully verifying the algorithm's strong adaptability and effectiveness in different scenarios.
[0050] This invention provides a UAV relay communication method based on multi-branch attention deep reinforcement learning. Many methods and approaches exist for implementing this technical solution; the above description is merely a preferred embodiment. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of this invention, and these improvements and modifications should also be considered within the scope of protection of this invention. All components not explicitly stated in this embodiment can be implemented using existing technologies.
Claims
1. A UAV relay communication method based on multi-branch attention deep reinforcement learning, characterized in that, Includes the following steps: Step 1: Construct an integrated optimization model for the UAV relay communication system, specifically including the UAV flight and energy consumption model, communication model, and integrated optimization model; Step 2: Model the optimization problem in Step 1 as a Markov decision process to adapt to the input requirements of the subsequent deep reinforcement learning (DRL) algorithm. This includes state design, action design, and reward function design. Step 3: Design a multi-branch feature extraction structure to independently encode information such as UAV location, user location, and communication requirements, thereby achieving feature decoupling from the source and effectively improving the extraction capability and optimization efficiency of state features. Step 4: Propose an attention mechanism and a priority experience replay strategy to achieve dynamic focusing on nodes with incomplete communication needs and priority learning of high-value samples. Step 5: Propose a multi-branch attention deep reinforcement learning algorithm to optimize and solve the model established in Step 1.
2. The UAV relay communication method based on multi-branch attention deep reinforcement learning according to claim 1, characterized in that, The specific process of establishing the UAV flight and energy consumption model in step 1 is as follows: Step 1.1: State update and decision execution are performed using discrete time slots; analysis shows that the UAV's motion follows a Markov decision process, which occurs in time slots. t The +1 state depends only on the time slot. t State actions; define time slots t The location of the drone at that time was Its action space is , indicating in time slot t Inner along x and y The displacement in the direction; therefore, the UAV position update design is as follows: ;(1) In addition, the power consumption of the UAV during propulsion comes from three parts, specifically including rotor blade profile power, induced power, and parasitic power; therefore, the total power consumption of the UAV can be modeled as follows: ; (2) In the formula, t For time slots, P 0、 P s These represent the blade profile power and induced power when the drone is hovering, respectively. v 1 represents the flight speed of the drone. U r For rotor tip speed, v 0 represents the rotor design parameter. d 0、 s , ρ a , A and are the fuselage drag coefficient, rotor solidity, air density, and rotor disk area, respectively.
3. The UAV relay communication method based on multi-branch attention deep reinforcement learning according to claim 1, characterized in that, The specific process of establishing the communication model in step 1 is as follows: Step 1.2: Use the Ricean fading model to characterize the channel characteristics, and the channel coefficient between the base station and the UAV. It can be represented as: ; (3) In the formula, α B-U This is the path loss index. d B-U The distance between the base station and the drone. ρ For reference distance loss coefficient; K 1 is the Rice factor of the BU link, which represents the power ratio of the line-of-sight component to the non-line-of-sight component; and These are the Los component and the NLoS transmission component in the channel, respectively. The elements of are independent and identically distributed random variables that follow a complex Gaussian distribution with a mean of 0 and a variance of 1. Determined by the array response vectors of the base station and the drone: ; (4) In the formula, and These are the array response vectors for the base station and the drone, respectively. for The conjugate transpose of . ϕ and θ These represent the departure angle and arrival angle of the signal, respectively. Similarly, the UAV-user link channel It can be modeled as: ; (5) In the formula, α U-M This is the path loss index. d U-M The distance between the drone and the user. ρ For reference distance loss coefficient; K 2 is the Rice factor of the UM link; The drone uses a gain amplification relay protocol, and its gain G The design ensures that the UAV's transmit power meets the following constraints: ; (6) In the formula, P u and P b These represent the transmission power of the drone and the base station, respectively. For receiving noise power for drones; In each time slot, the UAV provides relay services to users within its communication coverage area using orthogonal frequency division multiple access (OFDM) technology. The total system bandwidth B is divided into multiple orthogonal sub-channels. Due to spectrum resource constraints, the number of users the UAV can serve in each time slot is determined by the number of available orthogonal sub-channels. n Users in each time slot k Received signal-to-noise ratio It can be represented as: ; (7) In the formula, The noise power received by user k; in For the sub-bandwidth, according to Shannon's formula, the first... k A user in time slot n Transmission rate at time for: ;(8)。 4. The UAV relay communication method based on multi-branch attention deep reinforcement learning according to claim 1, characterized in that, The specific process of establishing the integrated optimization model in step 1 is as follows: Step 1.3, the optimization objective can be formally expressed as: ;(9) ;(10) ;(11) ;(12) ;(13) Regarding the above optimization objectives and constraints, the specific meanings of each symbol and formula can be further clarified: M Represents the total number of user nodes. N The total number of time slots in the scheduling cycle; Indicates time slot t Internal users i Communication rate, It is the set of UAV position offsets in each time slot, and the objective function (9) is obtained through the characteristic function. Maximize the fulfillment of preset communication requirements The number of users; the objective function (10) corresponds to maximizing the total communication rate of the system; in constraint (11), and Let be the position coordinates of the UAV in time slot t. and They are respectively in X , Y The positional boundary of the dimension; constraint (12) , It is the maximum displacement limit of the UAV within a single time slot; constraint (13) ensures the distance between the UAV and the user through the distance formula. i The distance does not exceed the effective communication radius. R This is to ensure service accessibility.
5. The UAV relay communication method based on multi-branch attention deep reinforcement learning according to claim 1, characterized in that, The specific process of state design in step 2 is as follows: Step 2.1, the state design is as follows: ; (14) In the formula, x U [ t ]and y U [ t ] respectively represent the drone in t The horizontal and vertical axes of the time slot; I for M Service status of each user node; Ψ The set of locations for all user nodes; This design integrates the drone's status with... M The complete information of each user node is integrated to form a 3-dimensional... M The +2 state space provides a sufficient environmental perception basis for decision-making.
6. The UAV relay communication method based on multi-branch attention deep reinforcement learning according to claim 1, characterized in that, The specific process of motion design in step 2 is as follows: Step 2.2, the specific action space is represented as follows: ; (15) In the formula, x t and y t They represent t The continuous displacement of the time-slotted UAV in the horizontal and vertical directions; by directly outputting displacement commands, the intelligent agent can achieve smooth trajectory planning and realize dynamic coverage of high-priority users.
7. The UAV relay communication method based on multi-branch attention deep reinforcement learning according to claim 1, characterized in that, The specific process of designing the reward function in step 2 is as follows: Step 2.3: The reward function, by quantifying and integrating key system performance indicators, provides fine-grained optimization guidance for policy learning. Its expression is as follows: ;(16) ; In the formula, The coverage reward represents the number of users within the drone's coverage area in the current time slot, directly incentivizing the expansion of service coverage. The throughput reward is calculated based on the total data transmission volume of the system within the current time slot, which directly promotes the improvement of spectrum resource utilization efficiency and data transmission capability. The number of users served in that time slot, in order to accelerate the overall service process. , , These are the corresponding weighting coefficients; It is a signal variable; it equals 1 when all users have been served, and 0 otherwise. K com It is a reward given after serving all users; For crossing the boundary, Penalty for moving to an empty space. Penalty for delays in progress, These are the weight coefficients corresponding to each penalty item, constraining invalid actions to ensure the effectiveness of the strategy.
8. The UAV relay communication method based on multi-branch attention deep reinforcement learning according to claim 1, characterized in that, The specific process of step 3 is as follows: A multi-branch feature extraction structure is proposed. By employing a parallel branching architecture, the original state space is deconstructed into three feature subspaces, each establishing an independent feature processing path. Each branch adopts a unified network structure, wherein... The branch index is represented as follows: ;(17) Equation (17) is the implementation carrier of the initial transformation for feature extraction, in which the linear transformation step is composed of the weight matrix. With bias vector Collaborative completion— Responsible for processing the raw input Mapping to a high-dimensional feature space expands the representational dimension of features to enhance their representational power; This is used to compensate for feature shifts after linear transformation, making the transformed feature distribution more suitable for the processing needs of subsequent networks. The subsequent layer normalization operation, by re-centering and scaling the linear transformation result, effectively eliminates the adverse interference of differences in physical dimensions and distributions between the original inputs on the training process, providing more stable feature inputs for subsequent network layers. Based on this, a nonlinear transformation needs to be introduced into the normalized features to further enhance their nonlinear representation capability. ; (18) Equation (18) provides the network with the necessary nonlinear mapping capability, enabling the model to learn and fit complex feature relationships; at the same time, by setting the output of some neurons to zero, the sparsity representation of features is achieved, which helps to filter noise information and enhance the discriminativeness of features; finally, the nonlinear features are refined and their dimensions are compressed to output the final encoding of this branch: ; (19) In the above processing, the layer normalization operation is specifically defined as follows: ; (20) In the formula, and These are the input vectors. x The mean and standard deviation, through Achieve feature standardization; Scaling factor For offset; through learnable parameters and Affine transformations are performed to enable the network to adaptively adjust the normalized feature distribution according to its own needs, while maintaining its expressive power.
9. The UAV relay communication method based on multi-branch attention deep reinforcement learning according to claim 1, characterized in that, The specific process of step 4 is as follows: A scaling dot product-based attention mechanism is designed to achieve accurate identification and adaptive resource allocation of key service targets by establishing a dynamic interaction model between UAVs and user nodes. The mechanism uses UAV state features extracted by the front-end multi-branch feature module as the query vector, which deeply encodes the UAV's real-time location information and service capability status. Simultaneously, the integrated user node features are used as keys and values, where each node feature contains its precise geographical location and real-time communication requirement status. To achieve effective interaction modeling, the system uses three sets of independent learnable weight matrices to map the query vector and key-value features to the same metric space, as detailed below: ;(21) ;(22) ;(23) This mapping operation transforms the query features representing drone status and the key-value features representing user group status into a unified vector space; through independent linear transformations. W Q , W K and W V This enables features from different modalities to be compared for similarity in the same metric space, laying the foundation for subsequent correlation calculations; Furthermore, the attention weights are calculated using scaled dot product attention as follows: ; (24) In the formula, d k The dimension representing the key vector is used to adjust the order of magnitude of the dot product result, preventing the dot product value from becoming too large due to a high vector dimension, which would cause the softmax function to enter the gradient saturation region and affect the learning stability of the model. The essence of this calculation process is to establish a correlation measurement mechanism between the UAV state and the features of each user node. By performing the dot product operation between the query vector and the key vector, a correlation score is generated for each user node, which accurately reflects the importance of the node at the current decision moment. The softmax function transforms these scores into a normalized probability distribution, ensuring that the sum of the attention weights of all user nodes is 1, thus forming a reasonable weight allocation. Finally, the algorithm uses the calculated attention weights. A value vector V We perform weighted fusion to generate a refined representation of user group characteristics: ; (25) The weighted summation operation enables adaptive aggregation of information, generating feature representations. H att In this process, the feature information of users identified as key users is enhanced, while the feature information of non-key users is suppressed. This dynamic weight allocation mechanism provides the policy network with feature mappings that have priority hints, guiding the UAV to focus on high-value service targets during the decision-making process.
10. The UAV relay communication method based on multi-branch attention deep reinforcement learning according to claim 1, characterized in that, The specific process of step 5 is as follows: The methods in steps 3 and 4 are integrated into the Deep Deterministic Policy Gradient (DDPG) algorithm to form a multi-branch attention deep reinforcement learning algorithm, which specifically includes: Step 5.1, Experience Replay and Sampling Mechanism; Samples generated by the interaction between the agent and the environment. Samples are stored in an experience replay pool that is prioritized. Based on the temporal-difference error (TD), samples are assigned priorities and non-uniform sampling is performed accordingly. Samples with high TD errors contain more unlearned environmental dynamics and are therefore given higher sampling probabilities. To overcome the bias that priority sampling may introduce, the algorithm combines importance sampling weights to correct gradient updates. The priority experience replay mechanism effectively improves the utilization efficiency of training data by guiding the agent to focus on samples with high learning value, thereby accelerating the policy convergence process as a whole. Step 5.2, Target value calculation and network update; The calculation of the target Q value is based on the Bellman equation expansion: ;(26) In the formula, Represents an instant reward. This is a discount factor, with a value ranging from 0 to 1, used to balance the importance of immediate rewards and future cumulative rewards; d t This is the round end marker; and These are the outputs of the target value network and the target policy network, respectively. Their synergistic effect provides a stable benchmark reference for the calculation of the target value. The update of the value network aims to minimize the mean squared error, and its loss function is defined as: ;(27) This loss function introduces importance sampling weights. w i This effectively corrects for distribution biases that may be caused by priority sampling; parameters are updated using gradient descent. Value networks can gradually improve their prediction accuracy, making the Q-value estimate continuously approach the true cumulative return expectation, thus providing a reliable evaluation basis for strategy optimization. Furthermore, the policy network update mechanism is based on the deterministic policy gradient theorem, and its update direction is guided by the gradient information of the value network: ;(28) In the formula F For sample batches, this gradient calculation process demonstrates the tight coupling between the value network and the policy network: Value Network It provides directions for improving the action space, indicating which actions will yield higher long-term returns; policy network These directional guidelines are then translated into actual adjustments to the network parameters. Through this chain-like gradient transmission, the policy network can continuously optimize its decision-making capabilities, generate superior drone displacement actions, and ultimately maximize long-term cumulative returns. Step 5.3: Target network soft update; The target network parameters are tracked using a soft update method to track the main network. ;(29) In the formula, This is the soft update coefficient.