A Method for Discovering UAV Neighbor Nodes Based on DQN Network under 6G Space-Air-Ground Integrated Network
The DQN-based method optimizes drone neighbor node discovery in 6G networks by dynamically adjusting beacon message intervals, addressing mobility and channel conflicts to enhance accuracy and reduce overhead.
Patent Information
- Application Number
- CN202310259591.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-16
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2043-03-16
AI Technical Summary
The existing drone neighbor node discovery method has low accuracy under high dynamics and channel conflicts, and the traditional Q-learning algorithm is slow to learn, making it unable to adapt to the high mobility and complex communication environment of the drone network.
Using the DQN network model, by constructing state, action and reward functions, combining two-dimensional Markov chains and CSMA/CA protocols, the broadcast interval of beacon messages is optimized, and the drone neighbor node discovery strategy is dynamically adjusted, system overhead is reduced and discovery accuracy is improved.
It improves the accuracy and efficiency of drone neighbor node discovery, reduces system overhead, and adapts to the high mobility and complex communication environment of drone networks.
Smart Images

Figure CN116390077B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of unmanned aerial vehicle (UAV) communication, and particularly relates to a method for discovering adjacent nodes of UAVs based on a deep Q-network (DQN) in a 6G integrated air-ground-space network. Background Art
[0002] With the continuous development of mobile communication technology, mankind has entered the era of mobile interconnection and interoperability. The booming development of 5G has made human life more convenient and colorful. However, it has also led to an exponential growth in data transmission. Currently, the widely developed 5G-related technologies can no longer meet the requirements of emerging services for the current communication performance. Therefore, accelerating the research on future communication network technologies has increasingly attracted the attention of the industrial and academic communities. The 6G network currently regards the integrated air-ground-space multi-access capability as an important key capability, which is based on the mobile communication network and expands the user access mode.
[0003] UAVs have a wide range of applications in fields such as environmental monitoring and disaster management. A self-organizing network composed of multiple UAVs can more effectively and economically complete tasks. With the rapid development of UAV technology, the UAV communication network will become a key component of the integration of the 6G integrated air-ground-space network and play an important role in civil and military fields such as battlefield reconnaissance, field rescue, and Internet of Things (IoT) information transmission.
[0004] Discovering adjacent nodes is a key step in discovering neighboring nodes and constructing the topology of the UAV network. Using the constructed topology, UAV network and routing schemes can be designed. The traditional method for discovering adjacent nodes is to set the broadcast packet time interval as a constant, that is, to send the status information of the local node to surrounding nodes at fixed time intervals. However, due to the three-dimensional deployment and high mobility of UAVs, the relative positions are constantly changing. The fixed broadcast packet time interval cannot adapt to the high mobility characteristics of the UAV self-organizing network. In addition, if the broadcast time interval is too short, the system overhead will increase; if the broadcast time interval is too long, adjacent nodes will be missed, resulting in a significant reduction in the discovery accuracy.
[0005] There are existing UAV adjacent node discovery methods based on Q-learning, which adjust the node's own state by continuously detecting the number of discovered adjacent nodes and change the beacon message sending interval to discover all neighbor nodes as much as possible. However, this method is not suitable for the high dynamicity of the UAV network, its state space will be large, and the Q-learning algorithm needs to discretize the problem during training, resulting in a very slow speed of learning the optimal strategy. And none of the existing UAV adjacent node discovery methods consider the situation where channel conflicts occur when two or more nodes broadcast discovery messages simultaneously, reducing the accuracy of adjacent node discovery. Summary of the Invention
[0006] To solve the above problems existing in the prior art, the present invention provides a method for discovering neighboring nodes of unmanned aerial vehicles (UAVs) based on a deep Q-network (DQN) in a 6G integrated air-ground-space network. The technical problems to be solved by the present invention are achieved through the following technical solutions:
[0007] A method for discovering neighboring nodes of UAVs based on a DQN in a 6G integrated air-ground-space network, comprising the following steps:
[0008] Step 1, constructing an initial system model of the UAV network; wherein, the initial system model includes a plurality of network nodes, a plurality of reference nodes randomly selected from the network nodes, and a communication channel;
[0009] Step 2, calculating the probability p of each reference node successfully competing for the channel and successfully sending information by using a two-dimensional Markov chain for the binary exponential backoff algorithm of the CSMA / CA protocol s ;
[0010] Step 3, setting the state of the DQN network according to the probability p of successfully competing for the channel and successfully sending information, setting the action of the DQN network according to the constant broadcast interval of sending beacon messages, and constructing a reward function of the DQN network according to the UAV neighboring node discovery reward and the beacon message sending times reward; s
[0011] Step 4, training the DQN network based on the state, the action, and the reward function to obtain a trained DQN network;
[0012] Step 5, inputting the current state corresponding to the network node into the trained DQN network and outputting the action corresponding to the maximum reward value.
[0013] In an embodiment of the present invention, the probability p of successfully competing for the channel and successfully sending information s is calculated according to the following formula:
[0014]
[0015] wherein,
[0016]
[0017] p c = 1 - (1 - p tr ) n ,
[0018]
[0019]
[0020]
[0021] p tr represents the probability that a reference node competes successfully for the channel; p c represents the probability of collision occurring in the channel; p b represents the probability that the channel is busy within a time slot; p a represents the probability that at least one data packet of a reference node is waiting to be sent; p tr0 represents p a = 1, the probability that a reference node competes successfully for the channel; m represents the maximum number of backoffs allowed by the binary exponential backoff algorithm; n represents the actual number of neighboring nodes of a reference node at a certain moment; r represents the number of nodes among n reference nodes that have at least one data packet waiting to be sent; W i represents the contention window size when the number of backoffs is i; represents the average contention window size of all states of the binary exponential backoff algorithm; λ a represents the arrival intensity of data packets other than hello packets.
[0022] In an embodiment of the present invention, the expression of the state is:
[0023] state = <X, Y, Z, R, V, p s >;
[0024] wherein, X, Y, and Z respectively represent the geographical positions of the unmanned aerial vehicle in three-dimensional space, R represents the communication range of the unmanned aerial vehicle, and V represents the flight speed of the unmanned aerial vehicle at the current moment;
[0025] The expression of the action is:
[0026] action s = {...τ - 0.1, τ, τ + 0.1...};
[0027] wherein, τ represents the constant broadcast interval for sending beacon messages;
[0028] The expression of the reward function is:
[0029]
[0030] wherein, r discovery represents the neighboring node discovery reward, r discovery = (N R - N D ) g r overhead represents the beacon message sending times reward, r overhead = τ', N R represents the number of actual neighboring nodes, N DLet $n$ denote the number of discovered neighboring nodes, $\tau'$ denote the broadcast interval for the current beacon message transmission, $g$ be the weight factor, and both $\alpha$ and $\beta$ denote constant coefficients.
[0031] In one embodiment of the present invention, step four includes:
[0032] Step 41, initialize the key parameters of the DQN network;
[0033] Step 42, select the action with the maximum reward value based on the current state using the greedy algorithm;
[0034] Step 43, execute the action with the maximum reward value and calculate the current reward value according to the reward function;
[0035] Step 44, obtain a new state;
[0036] Step 45, store the state transition result in the memory pool;
[0037] Step 46, when the number of state transition results in the memory pool is greater than the memory pool size, train the DQN network to obtain the trained DQN network.
[0038] In one embodiment of the present invention, the key parameters include: the memory pool size $D$, the training pool size $d$, the DQN network weights, the greedy algorithm probability $\varepsilon$, and the state space $DP$.
[0039] Advantages of the present invention:
[0040] The present invention takes the multi-service requirements of the UAV network as a consideration factor for neighboring node discovery, and proposes a UAV neighboring node discovery method through a multi-service competition model and based on the DQN network. In the present invention, the action corresponding to the maximum value of the reward function is also the optimal broadcast interval for sending beacon messages, and the maximum value of the reward function corresponds to the situation where the number of neighboring nodes discovered by the UAV is minimized while the broadcast interval for sending beacon messages is maximized. Therefore, it reduces the system overhead of the UAV network while improving the accuracy of neighboring node discovery.
[0041] The following will further elaborate on the present invention in detail in conjunction with the drawings and embodiments. Brief Description of the Drawings
[0042] Figure 1 It is a flowchart of a method for discovering neighboring nodes of UAVs based on the DQN network in a 6G air-ground-space integrated network provided by an embodiment of the present invention;
[0043] Figure 2 It is a schematic diagram of an initial UAV network system based on a 6G air-ground-space integrated network provided by an embodiment of the present invention. Detailed Embodiments
[0044] The present invention will be further described in detail below in conjunction with specific embodiments, but the implementation manners of the present invention are not limited thereto.
[0045] As Figure 1 and Figure 2 shown, a method for discovering neighboring nodes of unmanned aerial vehicles (UAVs) based on a deep Q-network (DQN) in a 6G integrated space-air-ground network includes the following steps:
[0046] Step 1: Construct an initial system model of the UAV network; wherein, the initial system model includes multiple network nodes, multiple reference nodes randomly selected from the network nodes, and a communication channel.
[0047] In this step, the network nodes refer to the active UAV nodes in the entire network, and their distribution in space satisfies a three-dimensional Poisson distribution. Their flight speeds vary due to different flight tasks, the communication range changes in real time according to the channel conditions, and the probability of successfully competing for the channel and successfully sending information changes with the number of nodes within the communication range. The reference nodes are randomly selected from the network nodes and are also ordinary nodes performing tasks in the UAV network. The communication channel is a characterization of the complex communication environment of the UAVs.
[0048] Step 2: Use a two-dimensional Markov chain to calculate the probability p of each reference node successfully competing for the channel and successfully sending information for the binary exponential backoff algorithm of the CSMA / CA protocol s . Specifically, the probability p s of successfully competing for the channel and successfully sending information is calculated according to the following formula:
[0049]
[0050] wherein,
[0051]
[0052] p c = 1 - (1 - p tr ) n ,
[0053]
[0054]
[0055]
[0056] p tr represents the probability of a reference node successfully competing for the channel; p c represents the probability of a collision occurring in the channel; p b represents the probability that the channel is busy within a time slot; p a represents the probability that at least one data packet of a reference node is waiting to be sent; ptr0 Indicates p a =1, the probability of the reference node successfully competing for the channel; m represents the maximum number of backoffs allowed by the binary exponential backoff algorithm; n represents the number of real neighbor nodes of the reference node at a certain moment; r represents the number of nodes among the n reference nodes that have at least one data packet waiting to be sent; W i Indicates the contention window size when the number of backoffs is i; represents the average contention window size of all states of the binary exponential backoff algorithm; λ a represents the arrival intensity of other data packets except hello packets. Combining the above equations, we can solve p s and p tr .
[0057] Step 3: According to the probability p of successfully sending information after successfully competing for the channel s Set the state of the DQN network, set the action of the DQN network according to the constant broadcast interval of sending beacon messages, and construct the reward function of the DQN network based on the drone neighbor node discovery reward and the beacon message sending number reward.
[0058] Define the states, actions, and rewards in DQN theory; in the DQN algorithm, a neural network is used to replace the Q table in the Q-learning algorithm to solve the problem of large space occupation and high computational complexity of the Q table in large-scale continuous state space.
[0059] Modeling the drone neighbor node discovery problem as a problem that can be solved by the DQN algorithm requires defining the state, action, and reward.
[0060] The expression for setting the state is:
[0061] state= <X,Y,Z,R,V,p s >
[0062] Among them, X, Y, and Z represent the geographical location of the drone in three-dimensional space respectively. The unit can be changed according to the movement range of the drone. The unit can be km or m. R represents the communication range of the drone, and V represents the flight speed of the drone at the current moment.
[0063] Action: The behavior in different states is the choice of broadcast interval. Based on the constant broadcast interval τ for sending beacon messages, several interval values are selected on the left and right sides of the value to make changes according to the change of state.
[0064] action s ={…τ-0.1,τ,τ+0.1…};
[0065] Among them, τ represents the constant broadcast interval for sending beacon messages;
[0066] The setting of the reward value will be evaluated according to the number of neighboring node drones found in state state. For example, in the actual drone flight scenario, if there are more neighboring node drones around in some states, then broadcasting needs to be performed more frequently. When the number of surrounding drones is small, the broadcast interval is increased, so as to achieve dynamic broadcast interval changes.
[0067] The expression of the reward function is:
[0068]
[0069] Among them, r discovery represents the neighboring node discovery reward, r discovery =(N R -N D ), g r overhead represents the beacon message sending times reward, r overhead =τ', N R represents the actual number of neighboring nodes, N D represents the number of neighboring nodes discovered, τ' represents the current broadcast interval for sending beacon messages, g represents the weight factor, and α and β both represent constant coefficients.
[0070] Reward: The reward function is a non-sparse reward function, including two parts: neighboring node discovery reward and beacon sending times reward.
[0071] The neighboring node discovery reward is used to represent the number of neighboring nodes discovered by the drone, which can enable the drone to obtain corresponding rewards for each broadcast interval selection. It is one of the main parts of the reward function and is of great significance for the efficient learning of the intelligent agent. For the convenience of counting, the number of missing nodes r discovery is used to characterize the neighboring node discovery reward. The number of missing nodes is the actual number of neighboring nodes minus the number of neighboring nodes discovered. The actual number of neighboring nodes is also the number of nodes within the communication range of the reference node during the broadcast interval for sending beacon messages.
[0072] The beacon sending times reward is an important indication part of the neighboring node discovery efficiency of the drone and can measure the communication overhead for neighboring node discovery by the drone. The beacon sending times reward is defined as r overhead =τ', where τ' characterizes the communication overhead. The smaller its value, the more frequently the beacon message is sent, and the greater the communication overhead. Conversely, the smaller the communication overhead.
[0073] Step 4: Train the DQN network based on the state, action, and reward function to obtain the trained DQN network. Specifically, Step 4 includes Step 41 - Step 46:
[0074] Step 41, initialize the key parameters of the DQN network; the key parameters include: the memory pool size D, the training pool size d, the DQN network weights, the greedy algorithm probability ε, and the state space DP.
[0075] Step 42, select the action with the maximum reward value based on the current state using the greedy algorithm;
[0076] Step 43, execute the action with the maximum reward value and calculate the current reward value according to the reward function;
[0077] Step 44, obtain the new state;
[0078] Step 45, store the state transition result in the memory pool;
[0079] Step 46, when the number of state transition results in the memory pool is greater than the memory pool size, train the DQN network to obtain the trained DQN network.
[0080] Step Five, input the current state corresponding to the network node into the trained DQN network, and output the action corresponding to the maximum reward value. In this step, when using the DQN network to solve the optimal beacon interval, based on the trained DQN network, solve the action with the maximum reward value in the current state, adjust the state, and the action taken is the optimal beacon message broadcast interval.
[0081] In this embodiment, each time the action corresponding to the maximum reward value is taken, the maximum value of the reward function corresponds to the action when the number of neighboring nodes discovered by the drone is the smallest and the broadcast interval of sending beacon messages is the largest. That is, the broadcast interval that makes the entire neighboring node discovery process accurate and low-cost. When the optimal broadcast interval is determined, it also means that the current state is determined, that is, the neighboring nodes are discovered.
[0082] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present invention, "a plurality" means two or more unless otherwise specifically defined.
[0083] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "examples", "specific examples", or "some examples", etc., mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, those skilled in the art can combine and combine the different embodiments or examples described in this specification.
[0084] The above content is a further detailed description of the present invention in combination with specific preferred embodiments, and it cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention pertains, without departing from the concept of the present invention, several simple deductions or substitutions can still be made, and all should be regarded as belonging to the protection scope of the present invention.
Claims
1. A method for discovering neighboring nodes of unmanned aerial vehicles based on a DQN network under a 6G integrated space-air-ground network, characterized in that It includes the following steps: Step 1, construct an initial system model of the UAV network; wherein, the initial system model includes multiple network nodes, multiple reference nodes randomly selected from the network nodes, and communication channels; Step 2: Use a two-dimensional Markov chain to calculate the probability that each of the reference nodes competes successfully for the channel and successfully sends information in the binary exponential backoff algorithm of the CSMA / CA protocol ; Step 3, according to the probability of successfully sending information on the competing channel Set the state of the DQN network, set the actions of the DQN network according to the constant broadcast interval for sending beacon messages, and construct the reward function of the DQN network based on the drone neighbor discovery reward and the beacon message sending times reward; Step 4, train the DQN network based on the state, the action, and the reward function to obtain a trained DQN network; Step 5, input the current state corresponding to the network node into the trained DQN network, and output the action corresponding to the maximum reward value, The expression of the state is: state=<X,Y,Z,R,V, p s >; wherein, X, Y, and Z respectively represent the geographical positions of the UAV in three-dimensional space, R represents the communication range of the UAV, and V represents the flight speed of the UAV at the current moment; The expression of the action is: action s ={…τ - 0.1, τ, τ + 0.1…}; Among them, represents the constant broadcast interval for sending beacon messages; The expression of the reward function is: ; Among them, represents the neighbor node discovery reward, , represents the beacon message sending times reward, , represents the number of true neighbor nodes, represents the number of discovered neighbor nodes, represents the broadcast interval of the current beacon message sending, is the weight factor, α and β both represent constant coefficients; The said Step 4 includes: Step 41, initialize the key parameters of the DQN network; Step 42, select the action with the maximum reward value based on the current state using the greedy algorithm; Step 43, execute the action with the maximum reward value and calculate the current reward value according to the reward function; Step 44, obtain a new state; Step 45, store the state transition result in the memory pool; Step 46, when the number of state transition results in the memory pool is greater than the memory pool size, train the DQN network to obtain a trained DQN network.
2. The method for discovering neighboring nodes of an unmanned aerial vehicle based on a DQN network in a 6G integrated air, space, and ground network according to claim 1, characterized in that, The probability that the competing channel is successful and the information is successfully transmitted is calculated according to the following formula: ; Wherein, , , , , ; Indicates the probability that a reference node successfully competes for the channel; Indicates the probability of collision occurring in the channel; Indicates the probability that the channel is busy within a time slot; Indicates the probability that at least one packet is waiting to be sent by a reference node; Indicates When = 1, the probability that a reference node successfully competes for the channel; Indicates the maximum number of backoffs allowed by the binary exponential backoff algorithm; Indicates the actual number of neighboring nodes of a reference node at a certain moment; r Indicates The number of nodes among Indicates the contention window size when the number of backoffs is i; Indicates the average contention window size for all states of the binary exponential backoff algorithm; Indicates the arrival intensity of other packets except hello packets.
3. A method for discovering neighboring nodes of an unmanned aerial vehicle based on a DQN network in a 6G integrated space-air-ground network according to claim 2, characterized in that, The key parameters include: memory pool size D, training pool size d, DQN network weights, greedy algorithm probability ε, and state space DP.