A Non-Orthogonal Multiple Access Method and System for Uplink of UAV Swarm Based on Reinforcement Learning
By employing non-orthogonal multiple access technology and multi-agent deep deterministic policy gradient reinforcement learning algorithm, the problems of low resource utilization efficiency and synchronization difficulties in UAV swarm communication are solved, thereby improving spectrum utilization and enhancing communication stability.
Patent Information
- Application Number
- CN202510084909.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-12-02
- Estimated Expiration
- 2045-01-20
AI Technical Summary
In existing drone swarm communication access technologies, orthogonal multiple access technology results in low resource utilization efficiency, difficulty in coping with data traffic fluctuations and synchronization difficulties, and affects communication reliability and efficiency.
By employing non-orthogonal multiple access technology combined with multi-agent deep deterministic policy gradient reinforcement learning algorithm, multiple UAVs can transmit signals in parallel on the same spectrum. The access decision is optimized through autonomous learning of the agents, thereby improving spectrum utilization and flexibly allocating transmission power.
It improves the spectrum utilization of UAV swarm communication, reduces interference and access latency, enhances communication stability and reliability, and optimizes overall communication performance.
Smart Images

Figure CN120018295B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of unmanned aerial vehicle (UAV) swarm communication technology, and relates to UAV swarm communication access technology, specifically to a non-orthogonal multiple access method and system for UAV swarm uplink based on reinforcement learning. Background Technology
[0002] Currently, orthogonal multiple access (OMA) technology dominates the field of UAV swarm communication access technology. Its principle lies in orthogonally partitioning wireless resources in time, frequency, or code domain to ensure that multiple users can share communication resources without interference. However, existing UAV swarm communication access technologies, primarily based on OMA, have significant shortcomings. On the one hand, resource utilization efficiency is low, especially in scenarios with frequent changes in UAV missions and drastic fluctuations in data traffic. Fixed-allocation time slots, frequency bands, or code channels are prone to idleness or overload. For example, in response to emergencies, the data volume of UAVs responsible for key monitoring areas surges, and conventionally allocated resources are insufficient to meet their transmission needs, leading to information delays. Meanwhile, relatively idle UAV resources in other areas cannot be flexibly allocated to support these busy UAVs. On the other hand, the system lacks flexibility. OMA technology heavily relies on precise synchronization mechanisms. As the scale of the UAV swarm expands or faces interference from complex electromagnetic environments, maintaining precise synchronization becomes extremely difficult, easily leading to synchronization problems such as time slot misalignment and frequency deviation, which in turn cause communication interruptions or a sharp increase in the bit error rate, posing a serious threat to the overall collaborative performance of the UAV swarm.
[0003] Therefore, there is an urgent need in this field to develop a non-orthogonal multiple access technology for UAV swarm uplink based on reinforcement learning, aiming to break through the limitations of traditional orthogonal technology, so as to achieve high efficiency and flexibility in UAV swarm communication access, and lay a solid foundation for the further expansion of UAV swarm applications. Summary of the Invention
[0004] To address the aforementioned problems in existing technologies, this invention provides a non-orthogonal multiple access (NOMA) method and system for the uplink of a drone swarm based on reinforcement learning. Firstly, this invention employs NOMA technology, breaking the spectrum utilization limitations of traditional orthogonal multiple access (OMA) and allowing multiple drones to transmit signals in parallel on the same spectrum, abandoning the strict orthogonal spectrum partitioning. Simultaneously, in steps 4-8, a multi-agent deep deterministic policy gradient (MADDPG) reinforcement learning algorithm is introduced to simulate the communication process of the drone swarm, enabling each drone to autonomously learn and optimize its access decisions based on the environment. This invention optimizes resource allocation and transmission performance in the drone swarm uplink communication system, improves spectrum utilization, and enables flexible on-demand allocation of drone transmit power and access order. These optimizations effectively reduce interference, access latency, and collision probability, thereby significantly improving communication efficiency and reliability and ensuring data transmission stability.
[0005] The present invention adopts the following technical solution:
[0006] A non-orthogonal multiple access method for uplink of UAV swarm based on reinforcement learning is described in the following steps:
[0007] Step 1: Determine the communication architecture for non-orthogonal multiple access uplink of the drone swarm;
[0008] Step 2: Based on the communication architecture, quantify the UAV parameters and complete the mathematical modeling of non-orthogonal multiple access in the uplink of the UAV swarm.
[0009] Step 3: Based on the mathematical model of non-orthogonal multiple access in the uplink of the UAV swarm, build a simulated communication scenario for non-orthogonal multiple access in the uplink of the UAV swarm;
[0010] Step 4: Define the state space and action space of each drone agent, and design the policy Actor network and value Critic network for each agent;
[0011] Step 5: Initialize the current state, actions, and parameters of the Actor and Critic networks for each agent;
[0012] Step 6: In the current time slot, calculate the signal-to-interference-plus-noise ratio (SINR) for each agent. If the SINR is greater than the minimum SINR threshold, the access is considered successful; otherwise, the access is considered unsuccessful.
[0013] Step 7: Review the experience and update the strategy and value network;
[0014] Step 8: Return to step 5 and continue until the maximum number of iterations is reached.
[0015] Preferably, in step 1, the communication architecture of the non-orthogonal multiple access uplink of the UAV cluster consists of one receiving UAV and N transmitting UAVs, wherein each transmitting UAV has a different channel state and allocated transmission power.
[0016] Preferably, in step 2, based on the overall communication architecture, key parameters such as UAV channel status and transmission power are quantified to complete the mathematical modeling of non-orthogonal multiple access in the uplink of the UAV cluster.
[0017] The channel is modeled as follows:
[0018]
[0019] set up Let h be the channel coefficient between the sending drone i and the receiving drone, where h i The parameter is λ i The small-scale Rayleigh fading channel coefficients, and |h i | 2 This is the corresponding channel power gain, which follows an exponential distribution |h i | 2 ~Exp(λ i L i It is a large-scale fading path loss factor, which typically depends on the communication distance between the transmitting and receiving UAVs and the surrounding environment.
[0020] set up This is the set of drone indexes that successfully accessed the receiving drone during time slot t. This represents the set of drones that successfully connected within the past t time slots. Therefore, the signal received by the drone in the current time slot k is:
[0021]
[0022] Where, p i is the transmit power of UAV i, and n is the noise of the channel.
[0023] The signal-to-interference-plus-noise ratio (SIR) of UAV i is calculated as shown in the following expression:
[0024]
[0025] in, This represents the channel gain of drone i.
[0026] To determine whether drone i has successfully connected in the current time slot, the following conditions must also be met:
[0027] εi >δ (4)
[0028] Where δ is the minimum signal-to-interference-plus-noise ratio threshold.
[0029] Preferably, in step 3, a simulated communication scenario for non-orthogonal multiple access (NOMA) of the UAV swarm uplink is built by combining the mathematical model of NOMA uplink NOMA, and the following simulation parameters are defined: number of UAVs N, UAV channel state The allocated UAV transmit power P = {P1, P2, ..., P N The maximum number of iterations T in the simulation scenario max The maximum number of access time slots t per iteration max .
[0030] Preferably, in step 4, the state space and action space of each drone agent are first defined. The state s of each agent... i This represents the interaction between the agent and the environment. For agent i in a drone swarm, its state space s represents the interaction between the agent and the environment. i for:
[0031]
[0032] Among them, o i This represents the current observation information of agent i.
[0033] Action space a i These are the actions that an agent can take at each moment. For each agent i, the actions that can be taken are:
[0034]
[0035] Therefore, the action space A of each agent i ={0,1} is a binary decision.
[0036] Then, the Actor network and Critic network for each agent i are designed. The Actor network is used to select the agent's action in a given state, and the policy for each agent i is... Based on its current state s i Generate action a i The function.
[0037] Policy function π i With action a i The relationship can be represented as:
[0038]
[0039] Where, π i For the policy of agent i, Its network parameters.
[0040] Set the objective function J(π) of the expected cumulative reward for the i-th agent in the Actor network. i By maximizing J(π) i To adjust the strategy π i :
[0041]
[0042] in, For the expected cumulative reward, r i t Let γ be the instantaneous reward that agent i receives in time slot t, and let γ be the discount factor.
[0043] A Critic network is set up to evaluate agent i in a given state s and all agent action sequences (a1,...,a2). N The expected cumulative reward under (i.e., the policy estimate function)
[0044]
[0045] in, Here are the parameters for the Critic network: s0 and a0 are the state to be given and the action sequence of all agents, and s and a are the state given to agent i and the action sequence of all agents.
[0046] Preferably, in step 5, at the beginning of each iteration, the current state, action, Actor network, Critic network, and target network parameters of each agent i are initialized. The target network is a copy of the Actor and Critic networks, used to improve training stability. Soft updates are used to reduce the variance of gradient updates and avoid unstable behavior. All drones are sorted from best to worst channel status, and their transmit power is assigned sequentially from highest to lowest. Simultaneously, the current time slot number t and the set of indices of the connected drones are initialized. And an experience replay buffer, wherein the experience replay buffer is used to store the agent's experience of interacting with the environment, including the current state, the action in the current state, the reward obtained, and the next state (s). i ,a i ,r i ,s i '). r i ,s i 'These represent the reward obtained by drone agent i and its next state, respectively.
[0047] Preferably, in step 6, in the current time slot t, the actions of each agent are based on the Actor network policy. Count the number of drone agents selected for access, and calculate the signal-to-interference-plus-noise ratio ε for each agent i according to the access order. i If the minimum signal-to-interference-plus-noise ratio threshold ε is satisfied i If the value is greater than δ, the connection is considered successful, and an immediate reward r is given. i t And include it in the set of drone indexes that successfully accessed during time slot t. No further connection is required in this iteration; otherwise, it will be considered a connection failure and an immediate penalty of -r will be applied. i t Then, stop all drones from accessing the current time slot. Then update the status of all drones. i This includes channel status, current transmit power, and current observation information.
[0048] Preferably, in step 7, experience replay and policy network updates are performed, including the current state, action, reward, and next state (s) of each agent i. i ,a i ,r i ,s i Stored in the experience replay buffer. and from A batch of experience samples is randomly selected from the data, and these samples are used to update the Actor and Critic networks of the agent.
[0049] First, the Critic network is updated by minimizing the loss function, which is calculated based on the Temporal Difference (TD) method, i.e.:
[0050]
[0051] The target value y is calculated using the following formula:
[0052]
[0053] Among them, Q i' Represents the target Critic network, π i' Represents the target Actor network. For the target Critic network parameters, The parameters of the target Actor network.
[0054] Then, the Actor policy network is updated using the policy gradient method. The policy gradient update formula is:
[0055]
[0056] in, To calculate the objective function J(π) i )right gradient, To calculate strategy π i right gradient, To compute the policy estimate function Q i For action a i The gradient.
[0057] Finally, the parameters of the target network are adjusted through soft updates, with the update rules as follows:
[0058]
[0059] Where τ is the step size of the soft update.
[0060] Preferably, in step 8, steps 5-7 are repeated, and in each time slot, the process of UAV access, experience storage, and network update is repeated. This continues until the maximum number of access time slots t in each iteration is met. max or drone index set When the number of drones equals N, initialize and retrain in the next iteration until the maximum number of iterations T is reached. max The training is now over.
[0061] This invention also discloses a non-orthogonal multiple access system for the uplink of a UAV swarm based on reinforcement learning, used to execute the above-described method, comprising the following modules:
[0062] Communication architecture determination module: Determines the communication architecture for non-orthogonal multiple access uplink of the UAV swarm;
[0063] Mathematical modeling module: Based on the communication architecture, the module quantifies the UAV parameters and completes the mathematical modeling of non-orthogonal multiple access in the uplink of the UAV swarm.
[0064] Simulated communication scenario building module: Based on the mathematical model of non-orthogonal multiple access in the uplink of UAV swarm, a simulated communication scenario of non-orthogonal multiple access in the uplink of UAV swarm is built.
[0065] Define the space and network design module: Define the state space and action space of each UAV agent, and design the policy Actor network and value Critic network for each agent;
[0066] Initialization module: Initializes the current state, actions, and parameters of the Actor and Critic networks for each agent;
[0067] Judgment module: In the current time slot, calculate the signal-to-interference-plus-noise ratio (SINR) of each agent. If the SINR is greater than the minimum SINR threshold, the access is considered successful; otherwise, the access is considered unsuccessful.
[0068] Update module: Performs experience playback and updates strategies and value networks;
[0069] Iteration module: In each time slot, the process of drone access, experience storage and network update is repeated until the maximum number of iterations is reached.
[0070] The significant technical advantages of the non-orthogonal multiple access method and system for uplink of UAV swarm based on reinforcement learning in this invention are as follows:
[0071] (1) This invention uses non-orthogonal multiple access technology to replace the orthogonal allocation of spectrum resources in the traditional communication mode, allowing multiple UAVs to transmit signals concurrently on the same spectrum resource. This method effectively utilizes spectrum resources, avoids resource idleness and waste caused by fixed allocation, greatly improves spectrum efficiency, and provides solid support for the massive data transmission of UAV swarms.
[0072] (2) Deep reinforcement learning endows UAV swarms with highly intelligent autonomous learning and precise environmental adaptation capabilities. Through neural network training, a multi-agent deep reinforcement learning framework and policy optimization mechanism that includes both competition and cooperation mechanisms are established to simulate the access process of UAV swarms, realizing autonomous training and intelligent decision-making of UAVs. This invention can intelligently adjust the access strategy according to the real-time dynamic environment of the swarm, reducing the probability of collisions caused by mutual interference, improving the stability and reliability of UAV swarm communication, ensuring smooth communication links, and thus optimizing the overall communication performance of the swarm. Attached Figure Description
[0073] Figure 1 The communication architecture diagram involved in the non-orthogonal multiple access method for uplink of UAV cluster provided in a preferred embodiment of the present invention is shown below.
[0074] Figure 2 A flowchart illustrating the steps of a non-orthogonal multiple access method for uplink of a drone swarm based on reinforcement learning, provided in a preferred embodiment of the present invention.
[0075] Figure 3 A flowchart illustrating the steps of traditional orthogonal multiple access (OMA) in the uplink of a drone swarm based on reinforcement learning for simulation comparison.
[0076] Figure 4 This is a reward training result diagram in one embodiment of the present invention;
[0077] Figure 5 The image shows the reward training results for the simulation comparison scheme.
[0078] Figure 6 This is a training result diagram of access latency in one embodiment of the present invention;
[0079] Figure 7 The training results are shown in the simulation comparison diagram for access latency.
[0080] Figure 8 This is a system block diagram of non-orthogonal multiple access for the uplink of a drone cluster, provided as a preferred embodiment of the present invention. Detailed Implementation
[0081] The preferred embodiments of the present invention will now be described in detail with reference to the accompanying drawings.
[0082] This embodiment provides a non-orthogonal multiple access method for the uplink of a drone swarm based on reinforcement learning, comprising the following steps:
[0083] Step 1: Determine the communication architecture for the uplink non-orthogonal multiple access of the drone swarm, such as... Figure 1 As shown. From Figure 1 As can be seen, the communication architecture of the non-orthogonal multiple access (NOMA) uplink of a drone swarm mainly consists of one receiving drone and N transmitting drones. Each transmitting drone has a different channel state (including small-scale fading and large-scale fading) and allocated transmit power. NOMA technology allows signal superposition between drones and distinguishes drones through different power allocations and serial interference cancellation techniques, thereby effectively reducing collisions between drones.
[0084] Step 2: Based on the overall communication architecture, quantify key parameters such as UAV channel status and transmission power to complete the mathematical modeling of non-orthogonal multiple access in the uplink of the UAV swarm.
[0085] The channel modeling in this embodiment is as follows:
[0086]
[0087] set up Let h be the channel coefficient between the sending drone i and the receiving drone, where h i The parameter is λ i The small-scale Rayleigh fading channel coefficients, and |h i | 2 This is the corresponding channel power gain, which follows an exponential distribution |h i | 2 ~Exp(λ i L i It is a large-scale fading path loss factor, which typically depends on the communication distance between the transmitting and receiving UAVs and the surrounding environment.
[0088] set up This is the set of drone indexes that successfully accessed the receiving drone during time slot t. This represents the set of drones that successfully connected within the past t time slots. Therefore, the signal received by the drone in the current time slot k is:
[0089]
[0090] Where, p i is the transmit power of UAV i, and n is the noise of the channel.
[0091] According to non-orthogonal multiple access (NOA) technology, drones with good channel conditions will be prioritized for access, while they will be subject to interference from drones with poor channel conditions. Therefore, the signal-to-interference-plus-noise ratio (SIR) of drone i is calculated as follows:
[0092]
[0093] in, represents the channel gain of drone i, and i' represents a drone that has not yet been connected and whose channel state is worse than that of drone i.
[0094] To determine whether drone i has successfully connected in the current time slot, the following conditions must also be met:
[0095] ε i >δ (4)
[0096] Where δ is the minimum signal-to-interference-plus-noise ratio threshold.
[0097] Step 3: As Figure 2 As shown, based on the mathematical model of non-orthogonal multiple access (NOMA) in the uplink of a drone swarm, a simulated communication scenario for NOMA in the uplink is constructed, and the following simulation parameters are defined: number of drones N, drone channel state. The allocated UAV transmit power P = {P1, P2, ..., P N The maximum number of iterations T in the simulation scenario max The maximum number of access time slots t per iteration max .
[0098] Step 4: First, define the state space and action space for each drone agent. The state s of each agent... i This represents the interaction between the agent and its environment. For agent i in a drone swarm, this information is crucial for the agent to make subsequent access decisions. Its state space s i for:
[0099]
[0100] Among them, o i This represents the current observation information of agent i, which mainly includes its own and the access status of other observed agents.
[0101] Action space a i These are the actions that an agent can take at each moment. For each agent i, the actions that can be taken are:
[0102]
[0103] Therefore, the action space A of each agent i Using a binary decision mechanism ({0,1}), the agent only needs to make a decision on whether to connect or not in each time slot. This binary decision-making method reduces computational complexity while maintaining sufficient flexibility to adapt to different communication environments, directly impacting the overall communication efficiency and collision probability of the drone swarm.
[0104] Secondly, a policy Actor network and a value Critic network are designed for each agent. The Actor-Critic network combines policy gradient and value function methods, enabling simultaneous optimization of both the policy and value function. The Actor network selects the agent's action in a given state, and the policy of each agent i... Based on its current state s i Generate action a i The Actor network is responsible for evaluating the quality of actions and providing feedback to guide the optimization of the Actor network. This structure allows the agent to learn effective access strategies more quickly, improving the overall performance of the system.
[0105] The policy function π of the Actor network i It can be represented as:
[0106]
[0107] Where, π i For the policy of agent i, Its network parameters.
[0108] Set the objective function J(π) of the expected cumulative reward for the i-th agent in the Actor network. i By maximizing J(π) i To adjust the strategy π i :
[0109]
[0110] in, For the expected cumulative reward, r i t Let γ be the instantaneous reward that agent i receives in time slot t, and let γ be the discount factor.
[0111] A Critic network is set up to evaluate agent i in a given state s and all agent action sequences (a1,...,a2). N The expected cumulative reward under (i.e., the policy estimate function)
[0112]
[0113] in, Here are the parameters for the Critic network: s0 and a0 are the state to be given and the action sequence of all agents, and s and a are the state given to agent i and the action sequence of all agents.
[0114] Step 5: At the beginning of each iteration, initialize the current state, action, Actor network, Critic network, and target network parameters for each agent i. The target network is a copy of the Actor and Critic networks, used to improve training stability. Soft updates are used to reduce the variance of gradient updates and avoid unstable behavior. All drones are sorted from best to worst channel status, and their transmit power is assigned sequentially from highest to lowest. The current time slot number t and the set of indices of the connected drones are initialized. And an experience replay buffer, wherein the experience replay buffer is used to store the agent's experience of interacting with the environment, including the current state, the action in the current state, the reward obtained, and the next state (s). i ,a i ,r i ,s i The initialization process serves as the starting point for the agent's autonomous learning. In each subsequent time slot, the agent selects an action based on the current state, observes the reward, and updates its network parameters to optimize future decisions. The experience replay mechanism accelerates the learning process by storing and reusing past experiences, thereby improving sample efficiency.
[0115] Step 6: In the current time slot t, perform actions according to the Actor network policy of each agent. Count the number of drone agents selected for access, and calculate the signal-to-interference-plus-noise ratio ε for each agent i according to the access order. i If the minimum signal-to-interference-plus-noise ratio threshold ε is satisfied i If the value is greater than δ, the connection is considered successful, and an immediate reward r is given. i t And include it in the set of drone indexes that successfully accessed during time slot t. No further connection is required in this iteration; otherwise, it will be considered a connection failure and an immediate penalty of -r will be applied. i t This interrupts the access process for all drones in the current time slot. Then, it updates the status s of all drones.i This includes channel status, current transmit power, and the access status of all observed UAVs. Meanwhile, all UAVs not currently connected will follow the new action strategy developed after training. Reconnect in a subsequent time slot.
[0116] Step 7: Perform experience replay and policy network update, recording the current state, action, reward, and next state (s) of each agent i. i ,a i ,r i ,s i Stored in the experience replay buffer. and from A batch of experience samples is randomly selected from the data, and these samples are used to update the Actor and Critic networks of the agent.
[0117] First, the Critic network is updated by minimizing the loss function, which is typically calculated using the Temporal Difference (TD) method. This method can more accurately evaluate the long-term returns of the current policy. The calculation method is shown in the following expression:
[0118]
[0119] The target value y is calculated using the following formula:
[0120]
[0121] Among them, Q i' Represents the target Critic network, π i' Represents the target Actor network. For the target Critic network parameters, The parameters of the target Actor network.
[0122] Then, the Actor policy network is updated using the policy gradient method, which uses the policy estimates fed back by the Critic network. The strategy is evaluated, and the parameters of the Actor network are optimized.
[0123] The policy gradient update formula is:
[0124]
[0125] in, To calculate the objective function J(π) i )right gradient, To calculate strategy π i right gradient, To compute the policy estimate function Q i For action a i The gradient.
[0126] Finally, the parameters of the target network are adjusted through soft updates, and the parameters of the current network are gradually transferred to the target network in a smooth manner to reduce instability and fluctuations during training.
[0127] Its update method is as follows:
[0128]
[0129] Where τ is the step size of the soft update, used to smooth the parameter updates of the target network.
[0130] Step 8: Repeat steps 5-7, performing the drone access, experience storage, and network update process in each time slot. This continues until the maximum number of access time slots t in each iteration is met. max or drone index set When the number of drones equals N, initialize and retrain in the next iteration until the maximum number of iterations T is reached. max Through iterative training, the drone agent can gradually learn effective access strategies that maximize the overall system performance. After training, access decision rules for each drone can be extracted from the Actor network. These rules dynamically adjust the drone's access behavior based on real-time channel conditions and allocated transmit power. This intelligent access decision rule significantly improves the communication efficiency and spectrum utilization of drone swarms, providing a novel solution for uplink communication in drone swarms.
[0131] The following simulation demonstrates the non-orthogonal multiple access method for UAV swarm uplink based on reinforcement learning proposed in this invention. A traditional orthogonal multiple access technology for UAV swarm uplink based on reinforcement learning is introduced as a simulation comparison scheme for the embodiments of this invention.
[0132] like Figure 3 This paper presents a flowchart illustrating the main steps of a traditional orthogonal multiple access (OMA) method for uplink in UAV swarms based on reinforcement learning. This scheme also employs the Multi-Agent Deep Deterministic Policy Gradient (MADDPG) algorithm from reinforcement learning to train all UAV agents. Figure 2The present invention, based on reinforcement learning, differs from traditional orthogonal multiple access (OMA) schemes in that, before access, instead of dynamically allocating transmit power based on the different channel states of each UAV, it randomly assigns a fixed transmit power value to each UAV during initialization. Furthermore, in each time slot t, this scheme allows only one UAV to access the network. When the number of UAVs exceeds one, a collision occurs, interrupting the access process in that time slot and imposing an immediate penalty of -r on all UAVs. i t Therefore, currently, access is only considered successful when there is only one drone and the minimum signal-to-interference-plus-noise ratio threshold is met, and that drone is given an immediate reward r. i t The Actor and Critic network parameters for each drone are then updated until training is complete.
[0133] According to the embodiments and simulation comparison schemes of the present invention, the various simulation parameters of the schemes are first defined, and the specific parameter information is shown in Table 1. Based on this, the reward training results and access latency training results of the two schemes are compared and simulated to verify the performance and advantages of the present invention.
[0134] Table 1. Various simulation parameters of the embodiments of the present invention and the simulation comparison scheme.
[0135]
[0136] The simulation training results are analyzed as follows:
[0137] First, the reward training results diagrams of the embodiments of the present invention and the simulation comparison scheme are analyzed, such as... Figure 4 and 5 As shown.
[0138] exist Figure 4 In the training process, the reward curve initially decreased and then rapidly increased, stabilizing after 1000 iterations, indicating a relatively stable overall training process. Particularly in the initial training phase, although the reward showed a slight decrease, the fluctuation was small and the algorithm's training efficiency was high. This demonstrates that the non-orthogonal multiple access method of this invention effectively allocates power, reducing the probability of collisions between UAVs, thereby minimizing instability during the trial-and-error process of the reinforcement learning algorithm. The agent can then rapidly optimize its decisions and successfully converge during training. Figure 5 In the process, the reward curve first decreases and then slowly rises, with large fluctuations, only gradually stabilizing after the 900th iteration. This is because traditional orthogonal multiple access methods cannot support concurrent access by multiple drones in the same time slot, causing the algorithm to require a large number of trials and errors during the training phase, resulting in large reward fluctuations, relatively unstable training process, and low algorithm training efficiency.
[0139] from Figure 4-5 The comparison shows that the non-orthogonal multiple access method of this invention exhibits a relatively stable convergence trend during training, demonstrating its significant advantages in handling multiple UAV access. Through an efficient power allocation mechanism, it avoids training instability caused by frequent access trials by UAV agents, thereby improving overall training efficiency and stability.
[0140] Then, the access latency training results of the embodiments of the present invention and the simulation comparison scheme are analyzed in detail, such as... Figure 6 and 7 As shown.
[0141] exist Figure 6 In the simulation, the access latency curve of this invention exhibits a stable and rapidly decreasing trend, especially after 1000 iterations, where the curve gradually stabilizes, and the access latency of the five drones remains relatively stable at 3.5 time slots. This indicates that this invention can dynamically adjust the drone access strategy according to different communication environments. Throughout the training process, the fluctuation in access latency is small. By reasonably allocating transmission power and optimizing the drone access sequence, collisions can be avoided when multiple drones access simultaneously, thereby significantly reducing the overall access latency of the communication system. Figure 7 In the first 900 iterations, the curve exhibited significant fluctuations, initially dropping sharply, then rising suddenly, then dropping again, and finally stabilizing after 900 iterations. The access latency for the five drones remained relatively stable within the five time slots. This is because the orthogonal multiple access scheme only allows one drone to access per time slot, and the channel state and transmit power of each drone are non-uniform, leading to frequent collisions during training and causing fluctuations in access latency. The algorithm gradually converged only in the later stages of training, ultimately achieving the ideal state of one drone accessing within one time slot.
[0142] from Figure 6-7 The comparison shows that the non-orthogonal multiple access method of this invention significantly reduces the access latency required by UAVs, improving performance by approximately 30%, and also significantly outperforms traditional orthogonal multiple access schemes in terms of training stability. This fully demonstrates the superiority of non-orthogonal multiple access in UAV swarm communication, especially in reducing access latency, lowering the probability of UAV collisions, and improving resource utilization efficiency.
[0143] like Figure 8 As shown, this embodiment provides a non-orthogonal multiple access system for the uplink of a drone swarm based on reinforcement learning, used to execute the above method, including the following modules:
[0144] Communication architecture determination module: Determines the communication architecture for non-orthogonal multiple access uplink of the UAV swarm;
[0145] Mathematical modeling module: Based on the communication architecture, the module quantifies the UAV parameters and completes the mathematical modeling of non-orthogonal multiple access in the uplink of the UAV swarm.
[0146] Simulated communication scenario building module: Based on the mathematical model of non-orthogonal multiple access in the uplink of UAV swarm, a simulated communication scenario of non-orthogonal multiple access in the uplink of UAV swarm is built.
[0147] Define the space and network design module: Define the state space and action space of each UAV agent, and design the policy Actor network and value Critic network for each agent;
[0148] Initialization module: Initializes the current state, actions, and parameters of the Actor and Critic networks for each agent;
[0149] Judgment module: In the current time slot, calculate the signal-to-interference-plus-noise ratio (SINR) of each agent. If the SINR is greater than the minimum SINR threshold, the access is considered successful; otherwise, the access is considered unsuccessful.
[0150] Update module: Performs experience playback and updates strategies and value networks;
[0151] Iteration module: In each time slot, the process of drone access, experience storage and network update is repeated until the maximum number of iterations is reached.
[0152] Other aspects of this embodiment can be found in the above method embodiments.
[0153] The preferred embodiments and principles of the present invention have been described in detail above. For those skilled in the art, there may be changes in the specific implementation based on the ideas provided by the present invention, and these changes should also be considered within the scope of protection of the present invention.
Claims
1. A non-orthogonal multiple access method for uplink of UAV swarm based on reinforcement learning, characterized by: The specific steps are as follows: Step 1: Determine the communication architecture of the uplink non-orthogonal multiple access for the drone swarm, where multiple drones transmit signals in parallel on the same spectrum; Step 2: Based on the communication architecture, quantify the UAV parameters and complete the mathematical modeling of non-orthogonal multiple access in the uplink of the UAV cluster; Step 3: Based on the mathematical model of non-orthogonal multiple access in the uplink of UAV swarm, build a simulated communication scenario for non-orthogonal multiple access in the uplink of UAV swarm; Step 4: Define the state space and action space of each drone agent, and design the policy Actor network and value Critic network for each agent; Step 5: Initialize the current state, actions, and parameters of the Actor network and Critic network for each agent; in this step, at the beginning of each iteration, the current state, actions, Actor network, Critic network, and target network parameters of each agent i are initialized. All drones are sorted from best to worst channel status, and their transmit power is assigned from highest to lowest in that order. Simultaneously initialize the current time slot number t and the set of indexes of connected drones. And an experience replay buffer, wherein the experience replay buffer is used to store the agent's experience of interacting with the environment, including the current state, the action in the current state, the reward obtained, and the next state. ; Step 6: In the current time slot, calculate the signal-to-interference-plus-noise ratio (SIR) for each agent. If the SIR is greater than the minimum SIR threshold, the access is considered successful; otherwise, the access is considered unsuccessful. In this step, in the current time slot t, based on the actions of each agent's Actor network... Count the number of drone agents selected for access, and calculate the signal-to-interference-plus-noise ratio εi for each agent i according to the access order. If the condition is met... If the connection is successful, an immediate reward will be given. And include it in the set of drone indexes that successfully accessed during time slot t. If no further connection is required in this iteration, then the connection will be considered a failure and an immediate penalty will be imposed. Then, stop all drones from accessing the current time slot; then update the status of all drones. This includes channel status, current transmit power, and current observation information; among which, For the policy of agent i, For network parameters, It is the minimum signal-to-interference-plus-noise ratio threshold; Step 7: Review the experience and update the strategy and value network; Step 8: Return to step 5 and continue until the maximum number of iterations is reached.
2. The non-orthogonal multiple access method for uplink of UAV swarm based on reinforcement learning as described in claim 1, characterized in that, In step 1, the communication architecture of the non-orthogonal multiple access uplink of the UAV cluster consists of one receiving UAV and N transmitting UAVs, where each transmitting UAV has a different channel state and an allocated transmit power.
3. The non-orthogonal multiple access method for uplink of UAV swarm based on reinforcement learning as described in claim 2, characterized in that, In step 2, the parameters include channel state and transmit power; The mathematical modeling is specifically as follows: set up Let h be the channel coefficient between the sending drone i and the receiving drone, where h i The parameter is λ i Small-scale Rayleigh fading channel coefficients, This is the corresponding channel power gain, which follows an exponential distribution. L i It is the large-scale fading path loss factor; set up This is the set of drone indexes that successfully accessed the receiving drone during time slot t. This represents the set of drones that successfully connected within the past t time slots; therefore, the signal received by the drone in the current time slot k is: Where, p i Here, n is the transmit power of UAV i, and n is the noise of the channel. The signal-to-interference-plus-noise ratio (SIR) of UAV i is calculated as follows: in, The channel gain represents the value of drone i. If the following condition is met, then it is determined that drone i has successfully accessed the network in the current time slot; 。 4. The non-orthogonal multiple access method for uplink of UAV swarm based on reinforcement learning as described in claim 3, characterized in that, In step 3, the following simulation parameters are defined: number of UAVs N, UAV channel state. Allocated UAV transmission power The maximum number of iterations T in the simulation scenario max The maximum number of access time slots t per iteration max .
5. The non-orthogonal multiple access method for uplink of UAV swarm based on reinforcement learning as described in claim 4, characterized in that, In step 4, the state s of each drone agent i This represents the interaction between the agent and the environment. For drone agent i in a drone swarm, its state s i for: Among them, o i This represents the current observation information of agent i; Action space These are the actions that a drone agent can take at any given moment. For each drone agent i, the actions that can be taken are: Therefore, the action space of each agent It is a binary decision; Design an Actor network and a Critic network for each agent i. The Actor network is used to select the agent's action in a given state, and the policy for each agent i is specified. Based on the current state s i Generate Actions The function; Set the objective function for the expected cumulative reward of the i-th agent in the Actor network. By maximizing To adjust the strategy : in, For the expected cumulative rewards, The instantaneous reward obtained by agent i in time slot t. Discount factor; A Critic network is set up to evaluate agent i in a given state s and all agent action sequences. Expected cumulative reward under the given conditions, policy estimate function as follows: in, For Critic network parameters, These represent the state to be given and the sequence of actions for all agents, respectively. Given the current state of agent i and the sequence of actions for all agents.
6. The non-orthogonal multiple access method for uplink of UAV swarm based on reinforcement learning as described in claim 5, characterized in that, In step 7, the current state, action, reward, and next state of each agent i are recorded. Store the experience samples in the experience replay buffer D; randomly extract experience samples from D and update the Actor network and Critic network of the agent using the experience samples; The Critic network is updated by minimizing a loss function, which is calculated based on the temporal difference method. The target value y is calculated using the following formula: in, Represents the target Critic network. Represents the target Actor network. For the target Critic network parameters, For the target Actor network parameters; The Actor network is updated using the policy gradient method, and the policy gradient update formula is: in, To calculate the objective function gradient, For computational strategy gradient, To compute the policy estimate function Action The gradient; The parameters of the target network are adjusted through soft updates, with the following update rules: in, This is the step size for soft updates.
7. The non-orthogonal multiple access method for uplink of UAV swarm based on reinforcement learning as described in claim 6, characterized in that, In step 8, when the maximum number of access time slots t in each iteration is satisfied... max or drone index set When the number of drones equals N, repeat steps 5-7 until the maximum number of iterations T is reached. max Training is over.
8. A non-orthogonal multiple access system for uplink of unmanned aerial vehicle (UAV) swarm based on reinforcement learning, used to perform the method as described in any one of claims 1-7, characterized in that, Includes the following modules: Communication architecture determination module: Determines the communication architecture of the uplink non-orthogonal multiple access of the UAV cluster, where multiple UAVs transmit signals in parallel on the same spectrum; Mathematical modeling module: Based on the communication architecture, the parameters of the UAV are quantified to complete the mathematical modeling of the uplink non-orthogonal multiple access of the UAV swarm; Simulated communication scenario building module: Based on the mathematical model of non-orthogonal multiple access in the uplink of UAV swarm, a simulated communication scenario of non-orthogonal multiple access in the uplink of UAV swarm is built. Define the space and network design module: Define the state space and action space of each UAV agent, and design the policy Actor network and value Critic network for each agent; Initialization module: Initializes the current state, actions, and parameters of the Actor and Critic networks for each agent; Judgment module: In the current time slot, calculate the signal-to-interference-plus-noise ratio (SINR) of each agent. If the SINR is greater than the minimum SINR threshold, the access is considered successful; otherwise, the access is considered unsuccessful. Update module: Performs experience playback and updates strategies and value networks; Iteration module: In each time slot, the process of drone access, experience storage and network update is repeated until the maximum number of iterations is reached.
Citation Information
Patent Citations
Distributed channel competition method based on multi-agent reinforcement learning
CN114375066A
Air-ground non-orthogonal multiple access uplink transmission method based on intelligent reflecting surface
CN114422056A