A Channel Access Method for Unmanned Aerial Vehicle Ad Hoc Networks

By applying deep reinforcement learning technology in the drone ad hoc network, each node can learn interactively and obtain a channel access strategy with strong adaptability, it solves the problem that the channel access protocol in the existing technology cannot meet the high-speed movement and topological dynamic changes of nodes, and achieves more efficient channel utilization and lower transmission delay.

CN114599115BActive Publication Date: 2025-06-10SOUTHEAST UNIV +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210142923.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-16
Publication Date
2025-06-10
Estimated Expiration
2042-02-16

AI Technical Summary

Technical Problem

The channel access protocol in the existing drone ad hoc network cannot meet the performance requirements of nodes' high-speed movement and topological dynamic changes, resulting in low channel utilization, poor fairness and high transmission delay.

Method used

Deep reinforcement learning technology is adopted to use each drone node as a decision-maker, and a highly adaptable channel access strategy is obtained through interactive learning, and the channel access probability is adjusted to improve channel utilization and fairness and reduce transmission delay.

Benefits of technology

It improves channel utilization and fairness, reduces transmission delay, and is suitable for channel access scenarios in drone ad hoc networks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114599115B_ABST
    Figure CN114599115B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for channel access in an unmanned aerial vehicle (UAV) ad-hoc network, which includes: when a UAV node has a transmission requirement, the node first listens to the channel. If there is no idle channel, the transmission is postponed; otherwise, the node decides whether to access according to the probability of accessing during idle time and selects which idle channel to access; after the node makes a decision, it obtains the corresponding feedback, modifies the reward value of the feedback according to the decision similarity between the current node and the surrounding nodes, and trains the neural network; before the next decision, the node takes the historical decision and feedback as the state to input into the neural network, and the network calculates and outputs the probability of accessing during idle time to guide the next decision of the node; repeat the above process at each time step, and the UAV decision-making body continuously interacts and learns with the environment, and finally obtains an access strategy with adaptability, channel utilization rate and node fairness.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of channel access, and particularly to a method for channel access in an unmanned aerial vehicle (UAV) ad hoc network. Background Art

[0002] The UAV ad hoc network has received increasing research attention due to its broad application prospects. Among them, the MAC protocol controls the rules followed by all network nodes when accessing the wireless channel, and determines how to maximize the use of limited channel bandwidth. Therefore, the quality of the channel technology directly determines the utilization rate of the wireless channel and the overall network performance. At present, UAV ad hoc networks mainly rely on competitive protocols in traditional Ad Hoc networks to manage channel access. However, UAV ad hoc networks have the characteristics of high-speed node movement and dynamic topology changes, and existing protocols cannot meet the performance requirements. Therefore, it is of great significance to study the MAC technology in UAV ad hoc networks.

[0003] The impact of network dynamics on competitive protocols lies in the fact that the channel competition environment will change, such as changes in the number of active nodes and the access strategies of other nodes. This requires each node to have a certain feedback and adjustment ability, and be able to access the channel in a way of dynamic strategy adjustment. Existing competitive protocols based on the CSMA mechanism support fast node access / withdrawal from the network, but the access collision probability increases with the increase in the number of nodes and lacks self-adaptability.

[0004] Based on deep reinforcement learning technology, the present invention takes each UAV node as a decision-making entity and proposes a distributed adaptive MAC algorithm, enabling the node to interact with the environment and learn until a more self-adaptive access strategy is obtained, improving channel utilization and fairness, reducing transmission delay, and having considerable application prospects. Summary of the Invention

[0005] In view of this, the purpose of the present invention is to provide a method for channel access in a UAV ad hoc network, which is used to adaptively adjust the channel access probability p with the change of the environment on the basis of the p-persistent CSMA protocol, so as to improve channel utilization and fairness and reduce transmission delay.

[0006] To achieve the above purpose, the present invention adopts the following technical solutions:

[0007] A method for channel access in a UAV ad hoc network, the access method is applied to a communication scenario based on the UAV ad hoc network. Each node in the UAV ad hoc network is used as a decision-making entity, and then based on the deep reinforcement learning algorithm, the decision-making entity interacts with the environment to learn and obtain a self-adaptive access strategy. The access method specifically includes the following steps:

[0008] Step S1: At a certain time slot, when one or more nodes in the UAV ad-hoc network are performing data transmission tasks, first perform carrier sensing on the channel to determine whether there is an idle channel. If all channels are occupied, choose to postpone access and make a decision again in the next time slot;

[0009] If there is at least one idle channel, select one of the idle channels for access according to the idle-time access probability, occupy several time slots to send data packets to the receiving node, or choose to postpone access and continue carrier sensing;

[0010] Step S2: Define the channel feedback and the corresponding reward value obtained when the node makes different decisions, including: if the node chooses to postpone access, the channel feedback is the busy / idle state of the channel, and the reward value is 0; if the node chooses to access an idle channel and the channel feedback is successful access, the reward value is 1; if the node chooses to access an idle channel and the channel feedback is node collision and access failure, the reward value is C, where -1 < C < 0;

[0011] Step S3: Select a certain node to interact and learn with other nodes, and compare the decisions of this node with those of the surrounding neighboring UAV nodes. Modify the reward value in Step S2 according to the similarity of the decisions. The higher the similarity, the greater the reward value for successful access, and the other reward values remain unchanged;

[0012] Step S4: Construct a deep Q-network and an experience replay pool for training. Use this experience replay pool as the input to train the deep Q-network, update the parameters in the network by the gradient descent method, and fix the network parameters after multiple iterations to obtain a channel allocation model. The experience replay pool includes the node selected in Step S3, its decisions and feedback in the current and past several time steps;

[0013] Step S5: For the node selected in Step S3, use its historical experience as the current state and input it into the channel allocation model obtained in Step S4. Calculate the different probabilities corresponding to different decisions of the node in the next step through this model, that is, the idle-time access probability;

[0014] Step S6: For all nodes with data transmission tasks in this communication scenario, repeat Step S1 - Step S5, make the next decision according to the idle-time access probability, until each node obtains an adaptive access strategy.

[0015] Further, for the communication scenario based on the UAV ad-hoc network, in this scenario, it includes N nodes and M channels. Each channel has the same bandwidth and access conditions, and each channel is divided into multiple time slots. The sets of nodes and channels are respectively denoted as: and

[0016] Further, step S3 specifically includes:

[0017] Step S301: When a node sends on an access channel, attach the idle-time access probability corresponding to the current decision to the data packet and send it out;

[0018] Step S302: Each node records the idle-time access probability p received from surrounding nodes, where p min is the minimum value received, and p max is the maximum value received;

[0019] Step S303: Divide the interval [p min , p max evenly into 8 sub-intervals, and sort the 8 sub-intervals in descending order according to the number of the interval where p is located as {[It 0 , It 1 , [It 1 , It 2 , ······, [It 7 , It 8}, that is, the p value appears most frequently in the interval [It 0 , It 1 , and the reward values corresponding to the 8 intervals are

[0020] Step S304: When the current decision of the node is to access the channel and the access is successful, change the reward value of this decision from 1 to R ACE .

[0021] Further, in step S4, two deep Q-networks with the same structure but different parameters are used for training, which are respectively named the main network and the target network, and the network parameters are respectively initialized as θ and θ - , and the parameters of the main network are assigned to the target network every F time steps to reduce the correlation between data. Among them, the deep Q-network adopts a recurrent neural network (RNN) structure, including an input layer, two hidden layers and an output layer, and the two hidden layers are respectively a long short-term memory layer (LSTM) and a forward propagation layer (FNN).

[0022] Further, in step S4, before training, an initial set {s t , a t , r t+1 , s t+1} needs to be established, where s t is the state at time step t, a t is the decision taken at time step t, and r t+1The reward obtained after making a decision at time step t, s t+1 The state of the next time step for time step t; the action a that the node may take at time step t t ∈ {0, 1, 2, ..., M}, a t When a is 0, the node chooses to defer access, a t When a is m and m is not 0, the node chooses to access channel m; the state s t+1 = [c t-Ω+2 , ..., c t , c t+1 , where c t+1 = [a t , z t T , z t is the feedback obtained by the node after making a decision at time step t, and the expression is: respectively represent the result of carrier sensing and the result of accessing the channel, and Ω is the length of the state history.

[0023] Furthermore, for the length of the state history, its value satisfies: 16 ≤ Ω ≤ 32.

[0024] Furthermore, in step S4, during training, using s t as the network input, its network output is:

[0025]

[0026] In formula (1), a represents all possible actions in the action set , θ is the weight of the edge in the neural network, and Q is the score corresponding to all possible decisions at the state at time t;

[0027] Processing Q obtained through formula (1) with ε-greedy and softmax algorithms and integrating it into the probability of accessing in idle time, and its expression is:

[0028] σ n (t) = {p n,0 (t), p n,1 (t), ..., p n,M (t)} (2)

[0029] In formula (2), when m is not 0, p n,M (t) represents the probability that node n chooses to access channel m at time step t, and when m is 0, p n,0 (t) represents the probability that the node defers access.

[0030] ​Further, in the step S4, when training, the non-uniform distribution characteristic of time steps needs to be considered to calculate the loss function, and the neural network parameters are updated by the gradient descent method, which specifically includes:

[0031] Each decision corresponds to a time step. The carrier sensing decision requires 1 time slot, and the channel access decision requires several time slots. Then, the calculation expression of the actual reward value at the τ time step is:

[0032]

[0033] In formula (3), γ is the time discount factor, 0 < γ < 1, and d(a τ ) is the duration that the decision made by the node at the time step τ needs to last;

[0034] The calculation expression of the loss function is:

[0035]

[0036] In formula (4), N E is the number of samples taken from the experience pool E when training the neural network, and e τ is a discrete sample;

[0037] Using the gradient descent method for the loss function L(θ) to update the neural network parameters, the calculation method is as follows:

[0038]

[0039] In formula (5), is the gradient function of Q(s τ , a τ ; θ).

[0040] The beneficial effects of the present invention are:

[0041] Compared with the prior art, based on the deep reinforcement learning technology, the present invention takes each UAV node as a decision-making body and proposes a distributed adaptive channel access algorithm, enabling the node to interact with the environment and learn until a more adaptable access strategy is obtained, improving channel utilization and fairness, reducing transmission delay, and having an observable application prospect. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1 It is a schematic flowchart of a method for channel access in a UAV ad hoc network provided in Embodiment 1;

[0043] Figure 2 It is a network model diagram of the UAV ad hoc network provided in Embodiment 1;

[0044] Figure 3Schematic diagram of the interaction learning between the drone node and the environment provided in Embodiment 1. Detailed implementation manners

[0045] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0046] Embodiment 1

[0047] Participate Figures 1 - 3 , this embodiment provides a method for channel access in a drone ad-hoc network. Based on the deep reinforcement learning technology, each drone node is used as a decision-making entity, and the node interacts with the environment to learn until a channel access strategy with strong adaptability is obtained. The method includes the following processes:

[0048] Step 1. When a drone node has a data transmission requirement, it first performs carrier sensing on the channels. If all channels are occupied, the node can only choose to postpone access and make a decision in the next time slot;

[0049] If an idle channel is detected, the node can choose to access an idle channel according to the probability of accessing during idle time, occupy several subsequent time slots to send data packets to specific nodes, or choose to postpone access and continue to perform carrier sensing;

[0050] Step 2. Different decisions of the drone node result in different feedbacks and reward values of the channel, specifically including:

[0051] If the drone node chooses to postpone access, the channel feedback is the busy / idle state of the channel, and the reward value is 0;

[0052] If the node chooses to access an idle channel and the channel feedback is successful access, the reward value is 1;

[0053] If the node chooses to access an idle channel and the channel feedback is node collision and access failure, the reward value is C (-1 < C < 0);

[0054] Step 3. The node interacts and learns with other surrounding nodes, compares the decision of the current decision-making drone node with the decisions of the surrounding neighboring drone nodes, and modifies the reward value in Step 2 according to the similarity of the decisions. The higher the similarity, the greater the reward value for successful access, and the other reward values remain unchanged;

[0055] Step 4: Pairwise store the decisions and feedbacks in the current and several past time steps into the experience replay pool of the deep reinforcement learning algorithm. Each time the neural network is trained, a cluster of data is sampled from the experience pool, and the weights of the edges in the network are updated by the gradient descent method. Use the historical decisions and feedbacks of the nodes as the current state to input into the neural network, and the neural network will correspondingly calculate the different probabilities corresponding to different decisions of the nodes in the next step, that is, the probability of accessing in idle time.

[0056] Step 5: The UAV nodes repeat the above steps and make the next decision according to the probability of accessing in idle time until each node obtains a more adaptable access method.

[0057] Specifically, in this embodiment, the network model of the UAV ad hoc network is as Figure 2 shown. Assume that there are N nodes in the network, and the nodes are divided into multiple clusters according to their locations. The UAV nodes in different clusters share limited bandwidth resources, and this resource can be divided into M channels according to the spectrum. Each channel has the same bandwidth and access conditions. The sets of nodes and channels are respectively denoted as: and

[0058] According to the requirements of the CSMA mechanism based on random competition, each channel can be further divided into multiple time slots. When each UAV node has a data transmission requirement, it will select an available channel to access in a certain time slot. When two or more neighboring nodes select to access the same channel in the same time slot, collision interference will occur, resulting in invalid access.

[0059] Specifically, in this embodiment, due to the limited transmission range of the nodes, two nodes at a relatively long distance accessing the same channel simultaneously will not cause interference. Therefore, the wireless channel has a certain degree of reusability. In Figure 2 , the UAV nodes in Cluster 1 and Cluster 2 are relatively close, so UAV No.1 - 5 form an interference domain. Similarly, UAV No.4 - 6 form another interference domain. Nodes in the same interference domain cannot access the same channel in the same time slot, while nodes in different interference domains are not restricted.

[0060] Specifically, in this embodiment, at the beginning of each time slot, the UAV node needs to make a decision, which specifically includes:

[0061] If the node has no data transmission requirement, the node can only choose to postpone accessing the channel and continue to maintain carrier sensing. If the node has a data transmission requirement, the node will first refer to the decision and feedback in the previous time slot. If the node performed sensing in the previous time slot and found an idle channel, the node will select one of the idle channels to access, occupy several subsequent time slots to send data packets to specific nodes. Otherwise, the node can only choose to postpone accessing the channel and continue to maintain carrier sensing.

[0062] Specifically, in this embodiment, which channel a node specifically selects to access is determined according to the access probability during idle time. The decision that a node can make at time step t can be expressed as a t ∈ {0, 1, 2,..., M}, where a t being 0 means that the node selects to postpone access, and a t being m (m ≠ 0) means that the node selects channel m for access, and m is one of the idle channels.

[0063] Specifically, in this embodiment, after an unmanned aerial vehicle (UAV) node makes a decision each time, it will obtain corresponding feedback from the channel. The node obtains corresponding rewards according to different feedbacks, which specifically include:

[0064] If the UAV node selects to postpone access, the channel feedback is the busy / idle state of the channel, and the reward value obtained by the node is 0; if the node selects to access an idle channel and the channel feedback is successful transmission, the reward value obtained by the node is 1;

[0065] If the node selects to access an idle channel but other nodes in the same interference domain select to access simultaneously, and the channel feedback is access failure, the reward value obtained by the node is C (-1 < C < 0).

[0066] The feedback obtained by the node after making a decision at time step t is z t , which is expressed as:

[0067]

[0068] Specifically, in this embodiment, an initial set {s t , a t , r t+1 , s t+1} is established for the deep reinforcement learning process, where s t is the state at time step t, a t is the decision made at time step t, r t+1 is the reward obtained after making a decision at time step t, and s t+1 is the state of the next time step at time step t, that is, the current situation of the channel that the node can observe and the historical experience obtained, which can be expressed as s t+1 = [c t-Ω+2 ,..., c t , c t+1 , where c t+1 = [a t , z t T , Ω is the length of the state history, that is, the length of the experience window. The larger Ω is, the more references the node can obtain when making decisions, and the greater the possibility of making a reasonable decision. However, the growth of the state space will lead to a deterioration of the algorithm's convergence. ​

[0069] More specifically, in this embodiment, Ω selects a compromise value, specifically 16 ≤ Ω ≤ 32.

[0070] Specifically, in this embodiment, the interactive learning process between the node and the environment is as follows Figure 3 As shown, at t - 4, the node selects channel 2 for access and the transmission is successful; since the node was in the transmission state in the previous time slot and did not perform carrier sensing, at t - 3, the node can only choose to listen, and the result is that there is an idle channel; because the listening result in the previous time slot was idle, at t - 2, the node can choose to access, but the node chooses to postpone access according to the idle - time access probability, and the listening result is that all channels are busy; after completing the listening at t - 1, at time t, the node selects channel 1 for access and successfully completes the transmission.

[0071] Specifically, in this embodiment, in addition to interacting and learning with the environment, the node can also improve the algorithm performance by interacting and learning with surrounding nodes. The method of interacting and learning with other surrounding nodes is as follows:

[0072] Step 301: When the node accesses the channel for transmission, it attaches the idle - time access probability corresponding to the current decision to the data packet and sends it out;

[0073] Step 302: Each node records the idle - time access probabilities p received from surrounding nodes, where p min is the minimum value received, and p max is the maximum value received;

[0074] Step 303: Divide the interval [p min , p max evenly into 8 sub - intervals, and sort the 8 sub - intervals in descending order according to the number of times p is in the interval as {[It 0 , It 1 , [It 1 , It 2 , ······, [It 7 , It 8}, that is, the p value appears most frequently in the interval [It 0 , It 1 , and the reward values corresponding to the 8 intervals are

[0075] Step 304: When the current decision of the node is to access the channel and the access is successful, the reward value of this decision is changed from 1 to R ACE .

[0076] Specifically, in this embodiment, during the deep reinforcement learning process, the decisions and feedbacks within a number of past time steps are stored in pairs in the experience replay pool of the deep reinforcement learning algorithm, and a cluster of data is sampled from the experience pool each time the neural network is trained.

[0077] Specifically, in this embodiment, the above neural network is the deep Q-network, which adopts a recurrent neural network (RNN) structure, including an input layer, two hidden layers, and an output layer. The two hidden layers are a long short-term memory layer (LSTM) and a forward propagation layer (FNN), respectively. The deep reinforcement learning algorithm uses the past Ω time steps as part of the current state for reference. The input of the neural network is s t , and the output is:

[0078]

[0079] In formula (2), a represents all possible actions in the action set , θ is the weight of the edges in the neural network, Q is the score corresponding to all possible decisions at the state at time t, and the higher the score, the greater the possibility that the decision is suitable for the current environment;

[0080] The output Q is processed using the ε-greedy and softmax algorithms and integrated into the probability of accessing during idle time. The expression is:

[0081] σ n (t) = {p n,0 (t), p n,1 (t),..., p n,M (t)} (3)

[0082] In formula (3), when m is not 0, p n,m (t) is the probability that node n selects to access channel m at time step t; when m is 0, p n,0 (t) is the probability that the node delays access.

[0083] Specifically, in this embodiment, the algorithm uses two neural networks with the same structure but different parameters, named the main network and the target network respectively. The network parameters are initialized as θ and θ - respectively, and the parameters of the main network are assigned to the target network every F time steps to reduce the correlation between data.

[0084] Specifically, in this embodiment, when considering the non-uniform time step characteristics, the training and updating process of the neural network is as follows:

[0085] Each decision of the node corresponds to a time step, and the duration corresponding to each time step is not necessarily the same. It takes 1 time slot for the node to make a deferred access and carrier sensing decision, and several time slots are required to make a channel access decision to send a data packet. In deep reinforcement learning, the value of the current decision is measured by the discounted sum of the values of the future states that may evolve. The influence of the value of the future state on the value of the current state gradually decreases as the time step progresses. Therefore, for those time steps with a longer duration, the discounted sum process needs to be modified.

[0086] When not considering the non-uniform characteristics of the time steps, the expression for calculating the real reward value at time step τ is:

[0087]

[0088] In this embodiment, the expression for calculating the real reward value is updated to:

[0089]

[0090] In formula (5), γ is the time discount factor, 0 < γ < 1, and d(a τ ) is the duration required for the decision made by the node at time step τ. The update content is to perform a discounted average processing on the reward, spread the reward value evenly over each time slot, and then also perform a discount processing on the maximum value of the subsequent state, because the next time step τ + 1 is already d(a τ ) time slots later.

[0091] The expression for calculating the error function is:

[0092]

[0093] In formula (6), N E is the number of samples taken from the experience pool E during the training of the neural network, and e τ is the discrete sample.

[0094] Using the gradient descent method for the error function L(θ) to update the neural network parameters, the calculation method is as follows:

[0095]

[0096] In the next time slot, the neural network will obtain a new state again, train and update the parameters, output the probability of accessing during idle time, and the UAV node selects the next decision according to the probability of accessing during idle time, and inputs the decision and feedback into the neural network, repeating continuously until a better adaptive access strategy is obtained.

[0097] In summary, based on the deep reinforcement learning technology, the present invention takes each drone node as a decision-making entity and proposes a distributed adaptive channel access algorithm, enabling the nodes to interact with the environment and learn until a more adaptive access strategy is obtained, improving channel utilization and fairness, reducing transmission delay, and having considerable application prospects.

[0098] Where the present invention is not described in detail, it is the well-known technology of those skilled in the art.

[0099] The preferred specific embodiments of the present invention have been described in detail above. It should be understood that those of ordinary skill in the art can make many modifications and variations based on the concept of the present invention without creative efforts. Therefore, all technical solutions that can be obtained by those skilled in the art in the technical field of the present invention through logical analysis, reasoning, or limited experiments based on the concept of the present invention on the basis of the prior art should fall within the protection scope determined by the claims.

Claims

1. A method for channel access in an unmanned aerial vehicle (UAV) ad-hoc network, characterized in that, the access method is applied to a communication scenario based on a UAV ad-hoc network. Each node in the UAV ad-hoc network is regarded as a decision-making entity. Then, based on a deep reinforcement learning algorithm, the decision-making entity interacts with the environment for learning to obtain an adaptive access strategy. The access method specifically includes the following steps: Step S1: At a certain time slot, when one or more nodes in the UAV ad-hoc network are performing data transmission tasks, first perform carrier sensing on the channel to determine whether there is an idle channel. If all channels are occupied, select to postpone access and make a decision again in the next time slot; If there is at least one idle channel, select one of the idle channels for access according to the idle-time access probability, occupy several time slots to send data packets to the receiving node, or select to postpone access and continue carrier sensing; Step S2: Define the channel feedback and the corresponding reward value obtained when the node makes different decisions, including: if the node selects to postpone access, the channel feedback is the busy / idle state of the channel, and the reward value is 0; if the node selects to access an idle channel and the channel feedback is successful access, the reward value is 1; if the node selects to access an idle channel and the channel feedback is node collision and access failure, the reward value is C, where -1 < C < 0; Step S3: Select a certain node to interact and learn with other nodes, and compare the decisions of this node with those of the surrounding neighboring UAV nodes. Modify the reward value in Step S2 according to the similarity of the decisions. The higher the similarity, the greater the reward value for successful access, and the other reward values remain unchanged; Step S4: Construct a deep Q-network and an experience replay pool for training. Use this experience replay pool as the input to train the deep Q-network, and update the parameters in the network by the gradient descent method. After multiple iterations, fix the network parameters to obtain a channel allocation model. The experience replay pool includes the node selected in Step S3, its decisions and feedback in the current and past several time steps; Step S5: For the node selected in Step S3, use its historical experience as the current state and input it into the channel allocation model obtained in Step S4. Calculate the different probabilities corresponding to different decisions of the node in the next step, that is, the idle-time access probability; Step S6: For all nodes with data transmission tasks in this communication scenario, repeat Step S1 - Step S5, make the next decision according to the idle-time access probability until each node obtains an adaptive access strategy; The communication scenario based on the UAV ad-hoc network, in which there are N nodes and M channels. Each channel has the same bandwidth and access conditions, and each channel is divided into multiple time slots. The sets of nodes and channels are respectively denoted as: and The specific content of Step S3 includes: Step S301: Set that when the node accesses the channel for sending, attach the idle-time access probability corresponding to the current decision to the data packet and send it out; Step S302: Each node records the idle access probability p received from the surrounding nodes, where p min is the minimum value received, and p max is the maximum value received; Step S303: Evenly divide the interval [p min , p max into 8 sub - intervals, and sort the 8 sub - intervals in descending order according to the number of p in each interval as {[It 0 , It 1 , [It 1 , It 2 , ······, [It 7 , It 8}, that is, the p value appears most frequently in the interval [It 0 , It 1 . The reward values corresponding to the 8 intervals are Step S304: When the current decision of the node is to access the channel and the access is successful, change the reward value of this decision from 1 to R according to the interval where the idle time access probability p of the decision lies ACE ; In the step S4, two deep Q networks with the same structure but different parameters are used for training, which are respectively named as the main network and the target network, and the network parameters are respectively initialized as θ and θ - , and the parameters of the main network are assigned to the target network every F time steps to reduce the correlation between data. Among them, for the deep Q network, a recurrent neural network (RNN) structure is adopted, which includes an input layer, two hidden layers and an output layer, and the two hidden layers are respectively a long short-term memory layer (LSTM) and a forward propagation layer (FNN); In the step S4, before training, it is necessary to establish an initial set {s t , a t , r t+1 , s t+1}, where s t is the state at time step t, a t is the decision taken at time step t, r t+1 is the reward obtained after taking the decision at time step t, and s t+1 is the state of the next time step at time step t; the possible actions a t ∈ {0, 1, 2,..., M} that the node may take at time step t. When a t is 0, the node chooses to defer access. When a t is m and m is not 0, the node chooses to access channel m; the state s t+1 = [c t-Ω+2 ,..., c t , c t+1 , where c t+1 = [a t , z t T , and z t is the feedback obtained by the node after taking the decision at time step t. The expression is: respectively represent the results of carrier sensing and accessing the channel, and Ω is the state history length;​ In the step S4, during training, s t is used as the network input, and its network output is: In formula (1), a represents all possible actions in the action set , θ is the weight of the edges in the neural network, and Q is the score corresponding to all possible decisions in the state at time t; Process the Q obtained by formula (1) using the ε-greedy and softmax algorithms and integrate it into the idle-time access probability. Its expression is: σ n (t) = {p n,0 (t), p n,1 (t),..., p n,M (t)} (2) In formula (2), when m is not 0, p n,M (t) represents the probability that node n selects to access channel m at time step t. When m is 0, p n,0 (t) represents the probability that the node delays access; In Step S4, when training, it is necessary to consider the non-uniform distribution characteristics of time steps to calculate the loss function and update the neural network parameters by the gradient descent method. Specifically, it includes: Each decision corresponds to a time step. The carrier sensing decision requires 1 time slot, and the channel access decision requires several time slots. Then, the calculation expression of the real reward value at time step τ is as follows: In formula (3), γ is the time discount factor, 0 < γ < 1, and d(a τ ) is the duration required for the decision made by the node at time step τ; The calculation expression of the loss function is as follows: In formula (4), N E is the number of samples taken from the experience pool E when training the neural network, and e τ is a discrete sample; Apply the gradient descent method to the loss function L(θ) to update the neural network parameters. The calculation method is as follows: In formula (5), is the gradient function of Q(s τ , a τ ; θ).

2. A method for channel access in an unmanned aerial vehicle self-organizing network according to claim 1, characterized in that the length of the state history satisfies: 16 ≤ Ω ≤ 32.

Citation Information

Patent Citations

  • Unmanned aerial vehicle CSMA access method based on adaptive adjustment strategy

    CN111050413A

  • Air-ground cooperative self-organizing network data transmission method

    CN114025330A