Unmanned aerial vehicle cluster uplink non-orthogonal multiple access method and system based on reinforcement learning
By adopting non-orthogonal multiple access technology based on reinforcement learning in drone cluster communication, the problems of low resource utilization efficiency and lack of system flexibility in orthogonal multiple access technology are solved, and efficient and flexible drone cluster communication is achieved.
Patent Information
- Application Number
- CN202510084909.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-01-20
AI Technical Summary
Among the existing drone cluster communication access technology, orthogonal multiple access technology has problems such as low resource utilization efficiency, lack of system flexibility and serious synchronization problems, which is difficult to meet the needs of frequent changes in drone tasks and fluctuations in data traffic.
Using non-orthogonal multi-access access technology based on reinforcement learning, through the multi-agent deep deterministic strategy gradient algorithm, multiple drones are allowed to transmit signals in parallel on the same spectrum, realizing flexible resource allocation and optimized access decisions for the uplink of the drone cluster.
It improves spectrum utilization, reduces interference and access delay, reduces collision probability, and significantly improves the efficiency and reliability of drone cluster communication.
Smart Images

Figure CN120018295A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of unmanned aerial vehicle cluster communication, relates to the communication access technology of unmanned aerial vehicle clusters, and specifically relates to a non-orthogonal multiple access method and system for an unmanned aerial vehicle cluster uplink based on reinforcement learning. Background Art
[0002] At present, in the field of UAV cluster communication access technology, orthogonal multiple access technology occupies a dominant position. Its principle is to implement orthogonal division of wireless resources in time, frequency or code domain to ensure that multiple users can share communication resources without interfering with each other. However, the existing UAV cluster communication access technology dominated by orthogonal multiple access technology has significant shortcomings. On the one hand, the resource utilization efficiency is low. Especially in the situation where the UAV mission changes frequently and the data traffic fluctuates sharply, the fixed allocated time slots, frequency bands or code channels are prone to idleness or overload. For example, when responding to emergencies, the amount of data from the UAVs responsible for the key monitoring areas surges, and the conventionally allocated resources are difficult to meet their transmission needs, resulting in information lags. At the same time, the relatively idle UAV resources in other areas cannot be flexibly allocated to support these busy UAVs. On the other hand, the system lacks flexibility. Orthogonal multiple access technology is highly dependent on precise synchronization mechanisms. As the scale of the UAV cluster expands or faces interference from complex electromagnetic environments, it becomes extremely difficult to maintain precise synchronization, which easily leads to synchronization problems such as time slot misalignment and frequency deviation, which in turn causes communication interruption or a sharp increase in bit error rate, posing a serious threat to the overall collaborative efficiency of the UAV cluster.
[0003] Therefore, there is an urgent need in this field to develop a non-orthogonal multiple access technology solution for the uplink of drone clusters based on reinforcement learning, aiming to break through the limitations of traditional orthogonal technology to achieve efficient and flexible communication access for drone clusters, and lay a solid foundation for the further expansion of drone cluster applications. Summary of the invention
[0004] In order to solve the above problems existing in the prior art, the present invention provides a non-orthogonal multiple access method and system for the uplink of a drone cluster based on reinforcement learning. The present invention first adopts the non-orthogonal multiple access (NOMA) technology to break the spectrum utilization restrictions brought by the traditional orthogonal multiple access (OMA), allowing multiple drones to transmit signals in parallel on the same spectrum, and abandoning the strict orthogonal division of the spectrum. At the same time, the multi-agent deep deterministic policy gradient (Multi-Agent Deep Deterministic Policy Gradient, MADDPG) reinforcement learning algorithm is introduced in steps 4-8 to simulate the communication process of the drone cluster, so that each drone can autonomously learn and optimize the access decision according to the environment. The technical solution of the present invention optimizes the resource allocation and transmission performance of the drone cluster uplink communication system, improves the spectrum utilization, and realizes the flexible on-demand allocation of drone transmission power and access order. Through these optimizations, interference can be effectively reduced, access delay and collision probability can be reduced, thereby significantly improving communication efficiency and reliability, and ensuring the stability of data transmission.
[0005] The present invention adopts the following technical scheme:
[0006] The non-orthogonal multiple access method for UAV cluster uplink based on reinforcement learning is as follows:
[0007] Step 1: Determine the communication architecture of non-orthogonal multiple access uplink of drone cluster;
[0008] Step 2: Based on the communication architecture, quantify the UAV parameters and complete the mathematical modeling of the UAV cluster uplink non-orthogonal multiple access;
[0009] Step 3: Combined with the mathematical model of non-orthogonal multiple access of the UAV cluster uplink, a simulation communication scenario of non-orthogonal multiple access of the UAV cluster uplink is constructed;
[0010] Step 4: Define the state space and action space of each UAV agent, and design the strategy actor network and value critic network of each agent;
[0011] Step 5: Initialize the current state, action, and parameters of the Actor network and Critic network of each agent;
[0012] Step 6: In the current time slot, the signal interference and noise ratio of each agent is calculated. If the signal interference and noise ratio is greater than the minimum signal interference and noise ratio threshold, it is considered as access success, otherwise, it is considered as access failure;
[0013] Step 7: Replay experience and update strategy and value network;
[0014] Step 8: Return to step 5 until the maximum number of iterations is reached.
[0015] Preferably, in step 1, the non-orthogonal multiple access communication architecture of the drone cluster uplink consists of one receiving drone and N sending drones, wherein each sending drone has a different channel state and allocated transmission power.
[0016] Preferably, in step 2, based on the overall communication architecture, key parameters such as the drone channel state and transmission power are quantified to complete the mathematical modeling of the drone cluster uplink non-orthogonal multiple access.
[0017] The channel model is:
[0018]
[0019] set up is the channel coefficient between the sending UAV i and the receiving UAV, where h i The parameter is λ i The small-scale Rayleigh fading channel coefficient, and |h i | 2 is the corresponding channel power gain, which follows an exponential distribution h i | 2 ~Exp(λ i ). i is the large-scale fading path loss factor, which usually depends on the communication distance between the transmitting UAV and the receiving UAV and the surrounding environment.
[0020] set up is the index set of drones that successfully access the receiving drone during time slot t, represents the set of drones that have successfully accessed in the past t time slots. Therefore, the signal received by the receiving drone at the current time slot k is:
[0021]
[0022] Among them, p i is the transmission power of UAV i, and n is the noise of the channel.
[0023] Calculate the signal-to-interference-to-noise ratio of drone i, as shown in the following expression:
[0024]
[0025] in, Represents the channel gain of drone i.
[0026] To determine whether drone i is successfully connected in the current time slot, the following conditions must also be met:
[0027] εi >δ (4)
[0028] Where δ is the minimum signal to interference noise ratio threshold.
[0029] Preferably, in step 3, a mathematical model of non-orthogonal multiple access of the uplink of the drone cluster is combined to build a non-orthogonal multiple access simulation communication scenario of the uplink of the drone cluster, and the following simulation parameters are clarified: the number of drones N, the channel state of the drone The allocated UAV transmission power P = {P1, P2, ..., P N}, the maximum number of iterations in the simulation scenario T max , the maximum number of access time slots t in each iteration.
[0030] max
[0031] Preferably, in step 4, the state space and action space of each UAV agent are first defined. The state s of each agent i Represents the interaction between the agent and the environment. For agent i in the drone cluster, its state space s i for:
[0032]
[0033] Among them, i Represents the current observation information of agent i.
[0034] Action Space a i is the behavior that the agent can take at each moment. For each agent i, the actions that can be taken are:
[0035]
[0036] Therefore, the action space A of each agent i ={0,1} is a binary decision.
[0037] Then design the Actor network and Critic network of each agent i. The Actor network is used to select the action of the agent in a given state. The strategy of each agent i is based on its current state s i Generate action a i function.
[0038] Policy function π i With action a i The relationship can be expressed as:
[0039]
[0040] Among them, π i is the strategy of agent i, Its network parameters.
[0041] Set the objective function J(π) of the Actor network for the expected cumulative reward of the i-th agent i ), by maximizing J(π i ) to adjust the strategy π i :
[0042]
[0043] in, is the expected cumulative reward, r i t is the immediate reward obtained by agent i in time slot t, and γ is the discount factor.
[0044] Set the Critic network to evaluate agent i in a given state s and all agent action sequences (a1,...,a N ), which is the expected cumulative reward under the policy estimation function
[0045]
[0046] in, are the parameters of the Critic network, s0 and a0 are the states to be given and the action sequences of all agents, s and a are the states given by the current agent i and the action sequences of all agents.
[0047] Preferably, in step 5, at the beginning of each iteration, the current state, action, actor network, critic network and target network parameters of each agent i are initialized. The target network is a copy of the Actor and Critic networks, which is used to improve the stability of training, reduce the variance of gradient updates through soft updates, and avoid unstable behavior. Sort all drones by channel status from good to bad, and assign transmit power from large to small to drones in this order. At the same time, initialize the current time slot number t and the index set of connected drones and experience replay buffer, where the experience replay buffer is used to store the experience of the agent's interaction with the environment, including the current state, the action in the current state, the reward obtained, the next state (s i ,a i ,r i ,s i '). i ,s i ' are the reward and next state obtained by drone agent i respectively.
[0048] Preferably, in step 6, at the current time slot t, the actions of the Actor network strategy of each agent Count the number of drone agents that choose to access, and calculate the signal-to-interference-noise ratio ε of each agent i according to the access order i , if the minimum signal-to-interference-noise ratio threshold ε is met i >δ, it is considered as a successful access and an immediate reward r is given i t , and counted into the index set of drones that successfully accessed during time slot t No more access is required in this iteration; otherwise, it is considered as access failure and an immediate penalty of -r is given. i t , and stop all drones from accessing the current time slot. Then update the status s of all drones i ', including channel status, current transmit power and current observation information.
[0049] Preferably, in step 7, experience replay and policy network update are performed to update the current state, action, reward, next state (s i ,a i ,r i ,s i ') Stored to the experience replay buffer and from A batch of experience samples are randomly selected from the network, and the Actor and Critic networks of the intelligent agent are updated through these samples.
[0050] First, the Critic network is updated by minimizing the loss function, which is calculated based on the temporal difference (TD) method, namely:
[0051]
[0052] The target value y is calculated by the following formula:
[0053]
[0054] Among them, Q i' represents the target Critic network, π i' Represents the target Actor network, is the target Critic network parameter, The target Actor network parameters.
[0055] Then the Actor policy network is updated by the policy gradient method. The policy gradient update formula is:
[0056]
[0057] in, To calculate the objective function J(π i )right The gradient of To calculate the strategy π i right The gradient of To calculate the policy estimate function Q i For action a i gradient.
[0058] Finally, the parameters of the target network are adjusted through soft update, and the update rule is:
[0059]
[0060] Among them, τ is the step size of soft update.
[0061] Preferably, step 8 repeats steps 5-7, and repeats the process of drone access, experience storage and network update in each time slot. max or drone index collection When it is equal to the number of drones N, initialize and retrain in the next iteration until the maximum number of iterations T is reached. max The training is over.
[0062] The present invention also discloses a non-orthogonal multiple access system for uplink of a drone cluster based on reinforcement learning, which is used to execute the above method and includes the following modules:
[0063] Communication architecture determination module: determines the communication architecture of non-orthogonal multiple access of the UAV cluster uplink;
[0064] Mathematical modeling module: Based on the communication architecture, the UAV parameters are quantified to complete the mathematical modeling of the UAV cluster uplink non-orthogonal multiple access;
[0065] Simulated communication scenario building module: Combined with the mathematical model of non-orthogonal multiple access of the uplink of the drone cluster, a simulated communication scenario of non-orthogonal multiple access of the uplink of the drone cluster is built;
[0066] Define space and network design module: define the state space and action space of each UAV agent, and design the strategy actor network and value critic network of each agent;
[0067] Initialization module: initialize the current state, action, and parameters of the Actor network and Critic network of each agent;
[0068] Judgment module: In the current time slot, the signal interference and noise ratio of each intelligent agent is calculated. If the signal interference and noise ratio is greater than the minimum signal interference and noise ratio threshold, it is considered as access success, otherwise, it is considered as access failure;
[0069] Update module: to replay experiences and update strategies and value networks;
[0070] Iteration module: In each time slot, the process of drone access, experience storage and network update is repeated until the maximum number of iterations is reached.
[0071] The significant technical effects of the non-orthogonal multiple access method and system for uplink of drone cluster based on reinforcement learning of the present invention are as follows:
[0072] (1) The present invention adopts non-orthogonal multiple access technology to replace the orthogonal allocation of spectrum resources in the traditional communication mode, allowing multiple drones to transmit signals concurrently on the same spectrum resource. This method effectively utilizes spectrum resources, avoids the idleness and waste of resources caused by the fixed allocation method, greatly improves spectrum efficiency, and provides solid support for the transmission of massive data in drone clusters.
[0073] (2) Deep reinforcement learning gives drone clusters highly intelligent autonomous learning and precise environmental adaptability. Through neural network training, a multi-agent deep reinforcement learning framework and strategy optimization mechanism that includes both competition and cooperation mechanisms was established to simulate the access process of drone clusters and realize autonomous training and intelligent decision-making of drones. The present invention can intelligently adjust the access strategy according to the real-time dynamic environment of the cluster, reduce the probability of collision due to mutual interference, improve the stability and reliability of drone cluster communication, ensure smooth communication links, and thus optimize the overall communication performance of the cluster. BRIEF DESCRIPTION OF THE DRAWINGS
[0074] Figure 1 A communication architecture diagram of a non-orthogonal multiple access method for uplink of a drone cluster provided in a preferred embodiment of the present invention;
[0075] Figure 2 A flowchart of the steps of a non-orthogonal multiple access method for uplink of a drone cluster based on reinforcement learning provided in a preferred embodiment of the present invention;
[0076] Figure 3 Flowchart of the steps for simulating the comparison scheme of traditional orthogonal multiple access for uplink of UAV cluster based on reinforcement learning;
[0077] Figure 4 A reward training result diagram in one embodiment of the present invention;
[0078] Figure 5 This is a graph of reward training results for the simulation comparison scheme;
[0079] Figure 6 A diagram showing the training result of access delay in one embodiment of the present invention;
[0080] Figure 7 The training result diagram of the access delay of the simulation comparison scheme is shown;
[0081] Figure 8 A system block diagram of non-orthogonal multiple access for uplink of a drone cluster provided in a preferred embodiment of the present invention. DETAILED DESCRIPTION
[0082] The preferred embodiments of the present invention are described in detail below with reference to the accompanying drawings.
[0083] This embodiment provides a non-orthogonal multiple access method for uplink of a drone cluster based on reinforcement learning, which is performed in the following steps:
[0084] Step 1: Determine the communication architecture of the UAV cluster uplink non-orthogonal multiple access Figure 1 As shown. Figure 1 As can be seen in Figure 1, the communication architecture of non-orthogonal multiple access for the uplink of the drone cluster mainly consists of one receiving drone and N sending drones, where each sending drone has a different channel state (including small-scale fading and large-scale fading) and allocated transmit power. Non-orthogonal multiple access technology allows signal superposition between drones and distinguishes drones through different power allocation and serial interference cancellation technology, thereby effectively reducing collisions between drones.
[0085] Step 2: Based on the overall communication architecture, key parameters such as the UAV channel status and transmission power are quantified to complete the mathematical modeling of the UAV cluster uplink non-orthogonal multiple access.
[0086] The channel modeling of this embodiment is:
[0087]
[0088] set up is the channel coefficient between the sending UAV i and the receiving UAV, where h i The parameter is λ i The small-scale Rayleigh fading channel coefficient, and |h i | 2 is the corresponding channel power gain, which follows an exponential distribution h i | 2 ~Exp(λ i ). i is the large-scale fading path loss factor, which usually depends on the communication distance between the transmitting UAV and the receiving UAV and the surrounding environment.
[0089] set up is the index set of drones that successfully access the receiving drone during time slot t, represents the set of drones that have successfully accessed in the past t time slots. Therefore, the signal received by the receiving drone at the current time slot k is:
[0090]
[0091] Among them, p i is the transmit power of UAV i, and n is the noise of the channel.
[0092] According to the non-orthogonal multiple access technology, drones with good channel conditions will be accessed first, and will be interfered by the signals of drones with poor channel conditions. Therefore, the signal to noise ratio of drone i is calculated as shown in the following expression:
[0093]
[0094] in, represents the channel gain of drone i, and i' represents the drone that has not been connected but has a worse channel status than drone i.
[0095] To determine whether drone i is successfully connected in the current time slot, the following conditions must also be met:
[0096] ε i >δ (4)
[0097] Where δ is the minimum signal to interference noise ratio threshold.
[0098] Step 3: If Figure 2 As shown in the figure, combined with the mathematical model of non-orthogonal multiple access of UAV cluster uplink, a non-orthogonal multiple access simulation communication scenario of UAV cluster uplink is built, and the following simulation parameters are clarified: the number of UAVs N, the channel state of UAVs The allocated UAV transmission power P = {P1, P2, ..., P N}, the maximum number of iterations in the simulation scenario T max , the maximum number of access time slots t in each iteration.
[0099] max
[0100] Step 4: First, define the state space and action space of each UAV agent. The state s of each agent is i represents the interaction between the agent and the environment. For agent i in the drone cluster, this information is crucial for the agent to make subsequent access decisions. Its state space s i for:
[0101]
[0102] Among them, i Represents the current observation information of agent i, mainly including its own access status and the access status of other observed agents.
[0103] Action Space a i is the behavior that the agent can take at each moment. For each agent i, the actions that can be taken are:
[0104]
[0105] Therefore, the action space A of each agent i ={0,1} is a binary decision, so that the agent only needs to make a decision to access or not access in each time slot. This binary decision method reduces the computational complexity while maintaining sufficient flexibility to adapt to different communication environments, which directly affects the overall communication efficiency and collision probability of the drone cluster.
[0106] Secondly, we design a strategy actor network and a value critic network for each agent. The actor-critic network combines the policy gradient and value function methods to optimize both the strategy and the value function. The actor network is used to select the action of the agent in a given state. The strategy of each agent i is is based on its current state s i Generate action a i The Critic network is responsible for evaluating the quality of the action and guiding the optimization of the Actor network through feedback. This structure enables the agent to learn effective access strategies more quickly and improve the overall performance of the system.
[0107] The policy function π of the actor network i It can be expressed as:
[0108]
[0109] Among them, π i is the strategy of agent i, Its network parameters.
[0110] Set the objective function J(π) of the Actor network for the expected cumulative reward of the i-th agent i ), by maximizing J(π i ) to adjust the strategy π i :
[0111]
[0112] in, is the expected cumulative reward, ri t is the immediate reward obtained by agent i in time slot t, and γ is the discount factor.
[0113] Set the Critic network to evaluate agent i in a given state s and all agent action sequences (a1,...,a N ), which is the expected cumulative reward under the policy estimation function
[0114]
[0115] in, are the parameters of the Critic network, s0 and a0 are the states to be given and the action sequences of all agents, s and a are the states given by the current agent i and the action sequences of all agents.
[0116] Step 5: At the beginning of each iteration, initialize the current state, action, Actor network, Critic network and target network parameters of each agent i The target network is a copy of the Actor and Critic networks, which is used to improve the stability of training, reduce the variance of gradient updates through soft updates, and avoid unstable behavior. All drones are sorted from good to bad according to channel status, and the transmission power is assigned to the drones in this order from large to small. Initialize the current time slot number t and the index set of connected drones and the experience replay buffer, where the experience replay buffer is used to store the experience of the agent's interaction with the environment, including the current state, the action in the current state, the reward obtained, the next state (s i ,a i ,r i ,s i '). The initialization process is the starting point for the agent to learn autonomously. In each subsequent time slot, the agent selects actions based on the current state, observes rewards, and updates its network parameters to optimize future decisions. The experience replay mechanism accelerates the learning process by storing and reusing past experience, thereby improving sample efficiency.
[0117] Step 6: At the current time slot t, the actions of each agent according to the Actor network strategy Count the number of drone agents that choose to access, and calculate the signal-to-interference-noise ratio ε of each agent i according to the access order i , if the minimum signal-to-interference-noise ratio threshold ε is met i >δ, it is considered as a successful access and an immediate reward r is given i t , and counted into the index set of drones that successfully accessed during time slot t No need to access again in this round of iteration; otherwise, it is considered as access failure and an immediate penalty of -r is given i t , and interrupt the access process of all drones in the current time slot. Then update the status s of all drones i ', including channel status, current transmit power and access status of all observed drones. At the same time, all drones that are not connected will follow the new action strategy after training. Re-access in a subsequent time slot.
[0118] Step 7: Perform experience replay and update the policy network, and update the current state, action, reward, and next state (s i ,a i ,r i ,s i ') Stored to the experience replay buffer and from A batch of experience samples are randomly selected from the network, and the Actor and Critic networks of the intelligent agent are updated through these samples.
[0119] First, the Critic network is updated by minimizing the loss function. The loss function is usually calculated based on the temporal difference (TD) method, which can more accurately evaluate the long-term return of the current strategy. The calculation method is shown in the following expression:
[0120]
[0121] The target value y is calculated by the following formula:
[0122]
[0123] Among them, Q i' represents the target Critic network, π i' Represents the target Actor network, is the target Critic network parameter, The target Actor network parameters.
[0124] The Actor policy network is then updated using the policy gradient method, which uses the policy estimate fed back by the Critic network Evaluate strategies and optimize parameters of the Actor network
[0125] The policy gradient update formula is:
[0126]
[0127] in, To calculate the objective function J(πi )right The gradient of To calculate the strategy π i right The gradient of To calculate the policy estimate function Q i For action a i gradient.
[0128] Finally, the parameters of the target network are adjusted through soft updating, and the parameters of the current network are gradually transferred to the target network in a smooth manner to reduce the instability and fluctuation during the training process.
[0129] The update method is:
[0130]
[0131] Among them, τ is the step size of soft update, which is used to smooth the parameter update of the target network.
[0132] Step 8: Repeat steps 5-7, and repeat the process of drone access, experience storage and network update in each time slot. max or drone index collection When it is equal to the number of drones N, initialize and retrain in the next iteration until the maximum number of iterations T is reached. max . Through multiple iterations of training, the UAV agent can gradually learn effective access strategies that can maximize the overall performance of the system. After training, the access decision rules for each UAV can be extracted from the Actor network. These rules can dynamically adjust the access behavior of the UAV according to the real-time channel status and the allocated transmit power. This intelligent access decision rule can significantly improve the communication efficiency and spectrum utilization of the UAV cluster, providing a new solution for the uplink communication of the UAV cluster.
[0133] The following is a simulation of the non-orthogonal multiple access method for the UAV cluster uplink based on reinforcement learning proposed in the present invention, and introduces the traditional orthogonal multiple access technology for the UAV cluster uplink based on reinforcement learning as a simulation comparison scheme for an embodiment of the present invention.
[0134] like Figure 3 The main steps of the traditional orthogonal multiple access method for uplink of drone clusters based on reinforcement learning are shown in the flowchart. This scheme also uses the Multi-Agent Deep Deterministic Policy Gradient (MADDPG) algorithm in reinforcement learning to train all drone agents. Figure 2The difference between the non-orthogonal multiple access method for uplink of drone cluster based on reinforcement learning of the present invention is that before access, the traditional orthogonal multiple access scheme does not dynamically allocate the transmission power according to the different channel states of each drone, but randomly allocates a fixed transmission power value to each drone at the time of initialization; and in each time slot t, this scheme only allows one drone to access. When the number of drones exceeds 1, a collision will occur, resulting in the interruption of the access process of the time slot, and an immediate penalty of -r will be given to all drones. i t Therefore, currently only when there is only one drone and the minimum signal to noise ratio threshold is it considered a successful access, and the drone is given an instant reward r i t . Then update the Actor and Critic network parameters of each drone until the training is completed.
[0135] According to the embodiment of the present invention and the simulation comparison scheme, various simulation parameters of the scheme are first clarified, and the specific parameter information is shown in Table 1. On this basis, the reward training results and access delay training results of the two schemes are compared and simulated to evaluate the performance and superiority of the present invention.
[0136] Table 1 Various simulation parameters of the embodiment of the present invention and the simulation comparison scheme
[0137]
[0138] The simulation training results are analyzed as follows:
[0139] First, the reward training result diagrams of the embodiment of the present invention and the simulation comparison scheme are analyzed, as shown in FIG. Figure 4 and 5 shown.
[0140] exist Figure 4 In the experiment, the reward curve shows a trend of first decreasing and then rising rapidly, and stabilizes after 1000 iterations. The overall training process is relatively stable. Especially in the initial training stage, although the reward decreases slightly, its fluctuation is small and the efficiency of algorithm training is high. This shows that the non-orthogonal multiple access method of the present invention effectively allocates power, reduces the probability of collision between drones, and thus reduces the unstable fluctuations in the trial and error process of the reinforcement learning algorithm. The intelligent agent can quickly optimize the decision and successfully converge during the training process. Figure 5 In the example, the reward curve first drops and then rises slowly, with a large fluctuation range, and it gradually stabilizes after the 900th iteration. This is because the traditional orthogonal multiple access method cannot support multiple drones to access concurrently in the same time slot, which causes the algorithm to perform a lot of trial and error during the training phase, resulting in large reward fluctuations, a relatively unstable training process, and low algorithm training efficiency.
[0141] from Figure 4-5 It can be seen from the comparison that the non-orthogonal multiple access method of the present invention shows a relatively stable convergence trend during the training process, proving its significant advantage in handling multi-UAV access. Through the efficient power allocation mechanism, the training instability caused by frequent access trial and error of the UAV intelligent body is avoided, thereby improving the overall training efficiency and stability.
[0142] Then, the access delay training result diagram of the embodiment of the present invention and the simulation comparison scheme is specifically analyzed, as shown in FIG. Figure 6 and 7 shown.
[0143] exist Figure 6 In the experiment, the access delay curve of the present invention shows a trend of no fluctuation and rapid decline, especially after 1000 iterations, the curve gradually stabilizes, and the access delay of the five drones is basically stable at 3.5 time slots. This shows that the present invention can dynamically adjust the access strategy of the drones according to different communication environments. During the entire training process, the fluctuation of the access delay is small. By reasonably allocating the transmission power and optimizing the access order of the drones, it can ensure that collisions are avoided when multiple drones access at the same time, thereby significantly reducing the access delay of the overall communication system. And in Figure 7 In the figure, the curve shows obvious fluctuations in the first 900 iterations, first dropping sharply, then rising suddenly, and then dropping again, and finally stabilizing after 900 iterations, and the access delay of 5 drones is basically stable at 5 time slots. This is because the orthogonal multiple access scheme can only allow one drone to access each time slot, and the channel state and transmission power of each drone are uneven, resulting in frequent collisions during training, causing fluctuations in access delay. It was not until the late stage of training that the algorithm gradually converged and finally reached the ideal state of accessing 1 drone in 1 time slot.
[0144] from Figure 6-7 It can be seen from the comparison that the non-orthogonal multiple access method of the present invention significantly reduces the access delay required by the drone, improves the performance by about 30%, and is also significantly better than the traditional orthogonal multiple access scheme in terms of training stability. This fully demonstrates the superiority of non-orthogonal multiple access in drone cluster communication, especially in reducing access delay, reducing the probability of drone collision and improving resource utilization efficiency.
[0145] like Figure 8 As shown, this embodiment provides a non-orthogonal multiple access system for uplink of a drone cluster based on reinforcement learning, which is used to execute the above method and includes the following modules:
[0146] Communication architecture determination module: determines the communication architecture of non-orthogonal multiple access of the UAV cluster uplink;
[0147] Mathematical modeling module: Based on the communication architecture, the UAV parameters are quantified to complete the mathematical modeling of the UAV cluster uplink non-orthogonal multiple access;
[0148] Simulated communication scenario building module: Combined with the mathematical model of non-orthogonal multiple access of the uplink of the drone cluster, a simulated communication scenario of non-orthogonal multiple access of the uplink of the drone cluster is built;
[0149] Define space and network design module: define the state space and action space of each UAV agent, and design the strategy actor network and value critic network of each agent;
[0150] Initialization module: initialize the current state, action, and parameters of the Actor network and Critic network of each agent;
[0151] Judgment module: In the current time slot, the signal interference and noise ratio of each intelligent agent is calculated. If the signal interference and noise ratio is greater than the minimum signal interference and noise ratio threshold, it is considered as access success, otherwise, it is considered as access failure;
[0152] Update module: to replay experiences and update strategies and value networks;
[0153] Iteration module: In each time slot, the process of drone access, experience storage and network update is repeated until the maximum number of iterations is reached.
[0154] For other contents of this embodiment, please refer to the above method embodiment.
[0155] The preferred embodiments and principles of the present invention are described in detail above. For those skilled in the art, according to the ideas provided by the present invention, there may be changes in the specific implementation methods, and these changes should also be regarded as the protection scope of the present invention.
Claims
1. A non-orthogonal multiple access method for uplink of UAV cluster based on reinforcement learning, characterized in that: The specific steps are as follows: Step 1: Determine the communication architecture of non-orthogonal multiple access uplink of drone cluster; Step 2: Based on the communication architecture, quantify the UAV parameters and complete the mathematical modeling of the UAV cluster uplink non-orthogonal multiple access; Step 3: Combined with the mathematical model of non-orthogonal multiple access of the UAV cluster uplink, a simulation communication scenario of non-orthogonal multiple access of the UAV cluster uplink is constructed; Step 4: Define the state space and action space of each UAV agent, and design the strategy actor network and value critic network of each agent; Step 5: Initialize the current state, action, and parameters of the Actor network and Critic network of each agent; Step 6: In the current time slot, the signal interference and noise ratio of each agent is calculated. If the signal interference and noise ratio is greater than the minimum signal interference and noise ratio threshold, it is considered as access success, otherwise, it is considered as access failure; Step 7: Replay experience and update strategy and value network; Step 8: Return to step 5 until the maximum number of iterations is reached.
2. The non-orthogonal multiple access method for uplink of a drone cluster based on reinforcement learning as claimed in claim 1, characterized in that: In step 1, the non-orthogonal multiple access communication architecture of the UAV cluster uplink consists of one receiving UAV and N sending UAVs, where each sending UAV has a different channel state and allocated transmission power.
3. The non-orthogonal multiple access method for uplink of drone cluster based on reinforcement learning as claimed in claim 2, characterized in that: In step 2, the parameters include channel status and transmit power; The specific mathematical modeling is: set up is the channel coefficient between the sending UAV i and the receiving UAV, where h i The parameter is λ i The small-scale Rayleigh fading channel coefficient, |h i | 2 is the corresponding channel power gain, which follows an exponential distribution h i | 2 ~Exp(λ i );L i is the large-scale fading path loss factor; set up is the index set of drones that successfully access the receiving drone during time slot t, represents the set of drones that have successfully accessed in the past t time slots; therefore, the signal received by the receiving drone at the current time slot k is: Among them, p i is the transmission power of UAV i, n is the noise of the channel; Calculate the signal-to-interference-to-noise ratio of drone i as shown below: in, represents the channel gain of UAV i; If the following conditions are met, it is determined that drone i successfully accesses the current time slot; e i >d (4) Where δ is the minimum signal to interference noise ratio threshold.
4. The non-orthogonal multiple access method for uplink of drone cluster based on reinforcement learning as claimed in claim 3, characterized in that: In step 3, the following simulation parameters are specified: the number of drones N, the state of the drone channel The allocated UAV transmission power P = {P1, P2, ..., P N }, the maximum number of iterations in the simulation scenario T max , the maximum number of access time slots per iteration t max .
5. The non-orthogonal multiple access method for uplink of drone cluster based on reinforcement learning as claimed in claim 4, characterized in that: In step 4, the state s of each UAV agent is i Represents the interaction between the agent and the environment. For the drone agent i in the drone cluster, its state s i for: Among them, i Represents the current observation information of agent i; Action Space a i is the behavior that the UAV agent can take at each moment. For each UAV agent i, the actions that can be taken are: Therefore, the action space A of each agent i ={0,1} is a binary decision; Design the Actor network and Critic network of each agent i. The Actor network is used to select the action of the agent in a given state. The strategy of each agent i Based on the current state s i Generate action a i Function of Among them, π i is the strategy of agent i, is the network parameter; Set the objective function J(π) of the Actor network for the expected cumulative reward of the i-th agent i ), by maximizing J(π i ) to adjust the strategy π i : in, is the expectation of cumulative reward, is the immediate reward obtained by agent i in time slot t, and γ is the discount factor; Set the Critic network to evaluate agent i in a given state s and all agent action sequences (a1,…,a N ) under the expected cumulative reward, the policy estimation function as follows: in, are the parameters of the Critic network, s0 and a0 are the state to be given and the action sequence of all agents respectively, s and a are the state given by the current agent i and the action sequence of all agents respectively.
6. The non-orthogonal multiple access method for uplink of a drone cluster based on reinforcement learning as claimed in claim 5, characterized in that: In step 5, at the beginning of each iteration, the current state, action, actor network, critic network and target network parameters of each agent i are initialized. All drones are sorted according to channel status from best to worst, and transmit powers are assigned to the drones in this order from large to small; At the same time, initialize the current time slot number t and the index set of connected drones and the experience replay buffer, where the experience replay buffer is used to store the experience of the agent's interaction with the environment, including the current state, the action in the current state, the reward obtained, the next state (s i ,a i ,r i ,s i ').
7. The non-orthogonal multiple access method for uplink of a drone cluster based on reinforcement learning as claimed in claim 6, characterized in that: In step 6, at the current time slot t, according to the actions of each agent's Actor network Count the number of drone agents that choose to access, and calculate the signal-to-interference-noise ratio ε of each agent i according to the access order i , if ε is satisfied i >δ, it is considered as a successful access and an immediate reward is given And counted into the index set of drones that successfully accessed during time slot t No need to access again in this round of iteration; otherwise, it will be considered as access failure and given immediate penalty And stop all drones from accessing the current time slot; Then update the status s of all drones i ', including channel status, current transmit power and current observation information.
8. The non-orthogonal multiple access method for uplink of a drone cluster based on reinforcement learning as claimed in claim 7, characterized in that: In step 7, the current state, action, reward, and next state (s i ,a i ,r i ,s i ') Stored to the experience replay buffer from Randomly extract experience samples from the network, and update the agent's Actor network and Critic network through the experience samples; The Critic network is updated by minimizing the loss function, which is calculated based on the temporal difference method: The target value y is calculated by the following formula: Among them, Q i' represents the target Critic network, π i' Represents the target Actor network, is the target Critic network parameter, The target Actor network parameters; The Actor network is updated using the policy gradient method. The policy gradient update formula is: in, To calculate the objective function J(π i )right The gradient of To calculate the strategy π i right The gradient of To calculate the policy estimate function Q i For action a i The gradient of The parameters of the target network are adjusted through soft update, and the update rule is: Among them, τ is the step size of soft update.
9. The non-orthogonal multiple access method for uplink of a drone cluster based on reinforcement learning as claimed in claim 8, characterized in that: In step 8, when the maximum number of access slots per iteration is met, max or drone index collection When it is equal to the number of drones N, repeat steps 5-7 until the maximum number of iterations T is reached. max The training is over.
10. A non-orthogonal multiple access system for uplink of a drone cluster based on reinforcement learning, used to execute the method according to any one of claims 1 to 9, characterized in that: Includes the following modules: Communication architecture determination module: determines the communication architecture of non-orthogonal multiple access of the UAV cluster uplink; Mathematical modeling module: Based on the communication architecture, the UAV parameters are quantified to complete the mathematical modeling of the UAV cluster uplink non-orthogonal multiple access; Simulated communication scenario building module: Combined with the mathematical model of non-orthogonal multiple access of the uplink of the drone cluster, a simulated communication scenario of non-orthogonal multiple access of the uplink of the drone cluster is built; Define space and network design module: define the state space and action space of each UAV agent, and design the strategy actor network and value critic network of each agent; Initialization module: initialize the current state, action, and parameters of the Actor network and Critic network of each agent; Judgment module: In the current time slot, the signal interference and noise ratio of each intelligent agent is calculated. If the signal interference and noise ratio is greater than the minimum signal interference and noise ratio threshold, it is considered as access success, otherwise, it is considered as access failure; Update module: to replay experiences and update strategies and value networks; Iteration module: In each time slot, the process of drone access, experience storage and network update is repeated until the maximum number of iterations is reached.
Citation Information
Patent Citations
Distributed channel competition method based on multi-agent reinforcement learning
CN114375066A
Air-ground non-orthogonal multiple access uplink transmission method based on intelligent reflecting surface
CN114422056A
Distributed task unloading method for unmanned aerial vehicle assisted mobile edge computing
CN117440341A
Multi-unmanned aerial vehicle cooperative auxiliary communication optimization method based on deep reinforcement learning
CN117933517A
Unmanned aerial vehicle communication secrecy energy efficiency optimization method and system based on deep reinforcement learning
CN119255227A