Unmanned aerial vehicle ad hoc network adaptive channel access method
Through the multi-channel, multi-competition window mechanism and DDQN algorithm, channel selection and competition window adjustment are optimized, and channel congestion and conflict problems of wireless communication networks in high dynamic and high interference environments are solved, and network performance is improved.
Patent Information
- Application Number
- CN202510467924.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-07-11
AI Technical Summary
In the high dynamic and high interference environment, the channel selection and competition window adjustment mechanisms are insufficient, resulting in channel congestion, increased conflicts, data loss and throughput, and the inability to effectively utilize multi-channel resources.
The multi-channel and multi-competition window mechanism are used to combine the deep reinforcement learning (DDQN) algorithm to dynamically adjust the competition window value and backoff time through factors such as channel busyness rate, competition window popularity and throughput to optimize the channel access method.
It improves network throughput, reduces latency, optimizes network fairness, and enhances anti-interference capabilities, achieving efficient data transmission in complex environments.
Smart Images

Figure CN120302455A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of wireless communication technologies, and particularly to an adaptive channel access method for unmanned aerial vehicle (UAV) ad-hoc networks. Background Art
[0002] In modern wireless communication networks, especially in highly dynamic networks such as UAV ad-hoc networks and vehicle-to-everything (V2X) networks, nodes have high mobility, the network topology often changes, and the environmental interference is relatively complex. Traditional Medium Access Control (MAC) protocols (such as IEEE 802.11p) are mainly designed for low-dynamic and low-interference scenarios and cannot effectively cope with these dynamically changing environments, resulting in problems such as channel congestion, increased collisions, data loss, and throughput degradation. Especially in high-interference environments, fixed contention windows and single-channel selection mechanisms are difficult to adapt to changes in real time and cannot effectively utilize channel resources, thus affecting the overall performance of the network.
[0003] In the prior art, although deep reinforcement learning has been applied to a certain extent in the field of wireless communication, most of them focus on single-channel scenarios and lack effective perception and adaptive capabilities for complex interference environments. In addition, existing network protocols lack strategies for making full use of multi-channel resources and do not have a flexible mechanism for channel selection and contention window adjustment, and cannot optimize resource scheduling in real time according to the network state, resulting in poor system performance. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to overcome the deficiencies of the prior art and provide an adaptive channel access method for UAV ad-hoc networks. In particular, the present invention optimizes the adaptive channel access method in a high-dynamic communication environment, aiming to improve the throughput of the network, reduce latency, optimize fairness, and enhance anti-interference capabilities.
[0005] The present invention adopts the following technical solutions to solve the above technical problems:
[0006] An adaptive channel access method for UAV ad-hoc networks according to the present invention includes:
[0007] Initializing the contention window value of each channel;
[0008] UAV nodes in the UAV ad-hoc network select the optimal channel and start the backoff process for data transmission; wherein, the backoff time is randomly generated according to the current contention window value; if the data transmission fails, the UAV node dynamically adjusts the contention window value according to the reinforcement learning method and regenerates the backoff time.
[0009] As a further optimization of the adaptive channel access method for UAV ad-hoc networks according to the present invention, the UAV node dynamically adjusts the contention window value according to the reinforcement learning method, including:
[0010] Construct a state space based on the channel busy rate CBR, the popularity of the contention window, and the degree of channel interference sensed by the UAV nodes, and design a reward function based on the throughput, fairness, and channel interference degree of the UAV nodes. The reward function is used to dynamically adjust the contention window to achieve adaptive channel access optimization.
[0011] As a further optimization scheme of the UAV ad-hoc network adaptive channel access method described in the present invention, the reinforcement learning method includes a double deep Q network DDQN; the UAV nodes in the UAV network select the optimal channel according to the available channel table jointly maintained, and the available channel table refers to a list of channels that are not interfered or occupied and are in an idle state; the data includes the transmission picture, video, and text data of the UAV nodes.
[0012] As a further optimization scheme of the UAV ad-hoc network adaptive channel access method described in the present invention, the backoff time includes the backoff time of the i-th channel The backoff time of the i-th channel The calculation method is:
[0013]
[0014] where CW i is the current contention window value of the i-th channel, τ is the slot time, and random(*) is a random function, that is, an integer value is randomly selected from 0 to CW i as the backoff time.
[0015] As a further optimization scheme of the UAV ad-hoc network adaptive channel access method described in the present invention, CBR is calculated by the following formula:
[0016]
[0017] where t idle is the idle time, and the idle time refers to the time when the channel is not occupied, and t totdl is the total observation time, and the total observation time refers to the difference between the current time and the previous observation time.
[0018] As a further optimization scheme of the UAV ad-hoc network adaptive channel access method described in the present invention, the popularity of the contention window CW is obtained by the following method:
[0019] Statistically count the contention window values of the neighbor UAV nodes and their occurrence frequencies;
[0020] Sort the contention window values in descending order according to the occurrence frequency;
[0021] Calculate its popularity based on the position of the current UAV node's contention window value in the sorting.
[0022] As a further optimized solution of the self-organizing network adaptive channel access method for UAVs described in the present invention, the state space includes historical actions and historical rewards, and the historical actions and historical rewards are used to enhance the agent's perception ability of the dynamic environment.
[0023] As a further optimized solution of the self-organizing network adaptive channel access method for UAVs described in the present invention, the reward function is:
[0024]
[0025] Among them, is the contention window reward, and the contention window reward is a reward set based on the contention window popularity; is the average channel busy / idle rate reward, and the average channel busy / idle rate reward is a reward set based on the average channel busy / idle rate and the ratio. The ratio refers to the ratio of the current node throughput to the maximum throughput. success is the success flag, reward is the total reward, and β is the balance and coefficient.
[0026] As a further optimized solution of the self-organizing network adaptive channel access method for UAVs described in the present invention,
[0027]
[0028] Among them, rank(CW i ) is the contention window popularity of the i-th channel, CW max is the maximum contention window value, and CW i is the current contention window value of the i-th channel.
[0029] As a further optimized solution of the self-organizing network adaptive channel access method for UAVs described in the present invention, Calculate through the following formula:
[0030]
[0031] Among them, T is the current throughput, T max is the maximum throughput, is the average CBR of neighbor UAV nodes.
[0032] Compared with the prior art by adopting the above technical solutions, the present invention has the following technical effects:
[0033] (1) The realization of multi-channel and multi-contention window mechanisms supports dynamic channel selection and contention window adjustment;
[0034] (2) The dynamic adjustment mechanism of the contention window based on the DDQN algorithm optimizes the contention window size in real time according to network load, historical transmission success rate, and channel status, minimizing collisions and improving throughput to the greatest extent;
[0035] (3) By calculating the average channel busy rate (CBR) and the contention window popularity, the scheduling of channel selection and data transmission is further optimized;
[0036] (4) In the data retransmission mechanism, the backoff time is dynamically adjusted by combining the historical success rate and the current network state to improve the data transmission success rate;
[0037] (5) The present invention can optimize channel selection and contention window adjustment in a complex interference environment, reduce channel collisions, improve data transmission efficiency, and has good anti-interference ability; combined with the multi-channel, multi-contention window mechanism and the multi-agent reinforcement learning framework, nodes can automatically adjust the transmission strategy according to the network state to ensure the optimization of network performance. Description of the Drawings
[0038] Figure 1 It is a flowchart of the multi-channel multi-contention window mechanism;
[0039] Figure 2 It is a DDQN framework diagram;
[0040] Figure 3 It is a schematic diagram of the network model;
[0041] Figure 4 It is a schematic diagram of the frame structure. Detailed Embodiment
[0042] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be described in detail below in conjunction with the drawings and specific embodiments.
[0043] The present invention urgently needs an adaptive channel access method that can combine the multi-channel, multi-contention window mechanism with deep reinforcement learning to address the network performance challenges in a high-interference and high-dynamic environment, and improve the efficiency, robustness, and fairness of the wireless communication system. By combining the Double Deep Q-Network (DDQN) algorithm, the dynamic adjustment of channel selection and contention window is realized to solve the network performance optimization problem in a complex interference environment.
[0044] The present invention includes:
[0045] (1) The implementation of the multi-channel, multi-contention window mechanism, supporting dynamic channel selection and contention window adjustment.
[0046] (2) The dynamic adjustment mechanism of the contention window based on the DDQN algorithm optimizes the contention window size in real time according to network load, historical transmission success rate, and channel status, minimizing collisions and increasing throughput to the greatest extent.
[0047] (3) By calculating the average channel busy rate (CBR) and contention window popularity, the scheduling of channel selection and data transmission is further optimized.
[0048] (4) In the data retransmission mechanism, the backoff time is dynamically adjusted by combining the historical success rate and the current network state to improve the data transmission success rate.
[0049] 1. Multi-channel and multi-contention window mechanism
[0050] 1.1 Initialization phase
[0051] When each node is initialized, a minimum contention window value (CW min ) is assigned to each channel, forming an initial list of contention window sizes. For the sake of concise and clear description, the nodes in the present invention refer to UAV nodes, which will not be elaborated hereinafter.
[0052]
[0053] Among them, CW init is the initial contention window list, is the minimum contention window size of the i-th channel, usually set to 32 or 64 to ensure the fairness of the network in the startup phase and avoid overly intense competition at the beginning. Through the channel allocation in the initialization phase, it is ensured that the network has good transmission performance in the starting phase.
[0054] 1.2 Data transmission phase
[0055] When a node needs to perform data transmission, the data includes transmission pictures, videos, and text data of UAV nodes. First, channel selection is performed, and the optimal i-th channel is selected according to the jointly maintained available channel table. The available channel table refers to a list of channels that are not interfered with, not occupied, and in an idle state, and the following steps are executed:
[0056] (1) Channel detection: The node selects the optimal channel for transmission according to the jointly maintained available channel table.
[0057] (2) Backoff time calculation: According to the contention window value CW i of the selected i-th channel, the node generates a backoff time:
[0058]
[0059] Among them, τ is the slot time, usually 20 microseconds, which ensures competition fairness and controls channel access.
[0060] (3) Channel state monitoring: During the backoff process, the node monitors the idle time t idle of the channel and the total observation time t total , and calculates the channel busy rate CBR based on this information:
[0061]
[0062] By monitoring the channel state, the node can sense the channel utilization rate in real time, so as to more reasonably adjust the contention window and select the channel.
[0063] (4) Data transmission: When the backoff time reaches zero, the node starts data transmission on the selected i-th channel to maximize the data transmission success rate.
[0064] 1.3 Data retransmission stage
[0065] If the data transmission fails, the node performs the following steps:
[0066] (1) Contention window adjustment: The node dynamically adjusts the contention window value CW i of the i-th channel according to the DDQN algorithm to reduce collisions and optimize the transmission efficiency:
[0067]
[0068] Among them, f DDQN represents the new contention window size optimized by the DDQN algorithm, is the state space of the current node on the i-th channel at the k-th time step.
[0069] (2) Backoff time update: Based on the adjusted contention window size the node recalculates the backoff time
[0070] By adjusting the contention window and backoff time, the node can reduce collisions, improve the transmission success rate, and ensure fairness during the retransmission process.
[0071] 2. Multi-agent reinforcement learning framework
[0072] The DDQN algorithm is an algorithm in reinforcement learning that approximates the Q-function through a deep neural network to improve the accuracy in the decision-making process. The core principle of DDQN is to avoid the overestimation problem in Q-learning by using two neural networks. One network is responsible for calculating the current Q-value, and the other network is responsible for calculating the target Q-value. This dual-network structure enables DDQN to more accurately evaluate the value of actions when facing a high-dimensional state space, thus better guiding the behavior of nodes.
[0073] In the present invention, DDQN is used to calculate the optimal contention window size according to the state of the current node (including channel state, contention window size, historical success rate, etc.). Specifically, the node continuously interacts with the environment to learn the contention window adjustment strategy most suitable for the current environment, so as to minimize transmission conflicts and improve the throughput of the network. The specific design of the multi-agent reinforcement learning framework is introduced below.
[0074] 2.1 Multi-agent State Space Design
[0075] The state space of each agent contains its own observation value and the average state of neighbors, and the formula is as follows:
[0076]
[0077] (1) Own state:
[0078] (1-1) CW i : The current contention window value of the node on the i-th channel
[0079] (1-2) The local channel busy / idle rate of the node on the i-th channel
[0080] (2) Neighbor state:
[0081] (2-1) The average contention window value of neighbor nodes on the i-th channel:
[0082]
[0083] (2-2) The average busy / idle rate of neighbor nodes on the i-th channel:
[0084]
[0085] (2-3) rank(CW i ) : The frequency of the own contention window value among all neighbors, and it is used as an important state information for decision-making.
[0086] By combining neighbor information, nodes can more accurately perceive the competition state of the network and adjust their strategies according to changes in the environment.
[0087] 2.2 Calculation of the popularity of the contention window
[0088] The popularity of the contention window reflects the difference in the contention window value CW of the current node on the i-th channel from that of other nodes in the neighbor group, and the calculation method is as follows: i
[0089] (1) The node collects the contention window values of all neighbors within one-hop range and sorts them by frequency.
[0090] (2) Calculate the popularity according to the sorting position. The higher the popularity, the more beneficial the current contention window value is to reducing collisions:
[0091]
[0092] where Position(CW i ) is the position of the contention window value of the current node on the i-th channel among its neighbors, and |L| is the length of the sorted list. By calculating the popularity of the contention window, the node can make more intelligent decisions based on the behavior of its neighbors, thereby reducing channel collisions.
[0093] 2.3 Design of the reward function
[0094] The reward function comprehensively considers throughput, the popularity of the contention window, and the average CBR. The specific calculation formula is as follows:
[0095]
[0096] (1) Success flag success:
[0097] If the data transmission is successful, success = 1; otherwise success = 0.
[0098] (2) Reward for the popularity of the contention window
[0099]
[0100] where, rank(CW i ) is the popularity of CW, CW max is the maximum contention window value
[0101] (3) Reward for the average CBR
[0102]
[0103] where, The average CBR of neighbor nodes, T is the current throughput, T max is the maximum throughput.
[0104] By comprehensively considering throughput, competition window popularity, and average CBR, the reward function can more comprehensively evaluate the transmission performance of nodes, thereby guiding the agent to learn the optimal transmission strategy.
[0105] As Figure 1 shown, the method includes the following processes:
[0106] Step 1. Initialization phase:
[0107] In the initialization phase, each node will complete the following steps:
[0108] (1) Initialize the channel table: Each node initializes the available channel table according to the requirements of the network environment.
[0109] (2) Set the competition window range: Set the minimum and maximum values of the competition window for each channel and define an initial competition window value for each channel
[0110] (3) Channel state detection: Each node will evaluate the channel quality of the channel it is on during initialization, including information such as the idle time and interference intensity of the channel. This information will be used for subsequent calculation of the channel selection strategy.
[0111] Step 2. Data transmission phase:
[0112] The data transmission phase is the core process of the present invention, involving channel selection, competition window adjustment, and data transmission. The specific steps are as follows:
[0113] (4) Channel selection: The node first makes a preliminary selection according to the jointly maintained available channel table.
[0114] (5) Backoff time calculation: The node randomly generates a backoff time according to the competition window value of the selected channel.
[0115] (6) Channel state monitoring: The node will monitor the idle and busy states of the channel in real time.
[0116] (7) Data transmission: When the backoff time ends, the node starts to transmit data on the selected channel.
[0117] Step 3. Data retransmission phase
[0118] When the data transmission fails, the node will enter the data retransmission phase. The key steps in the retransmission phase include dynamic adjustment of the competition window, retry count control, and recalculation of the backoff time.
[0119] (8) Competitive window adjustment: If data transmission fails, the node will dynamically adjust the competitive window according to the DDQN algorithm.
[0120] (9) Retry count: Whenever data transmission fails, the node will increase the retry count. If the retry count exceeds the preset maximum retry count, the node will start transmitting the next data packet.
[0121] (10) Recalculate backoff time: Based on the new competitive window size, the node will recalculate the backoff time and continue to monitor the channel status until the backoff time expires and data retransmission is performed.
[0122] Step 4: Collect its own status information, calculate the status with neighbor nodes, and calculate the reward.
[0123] Step 5: Start the next transmission and repeat the above steps 1 - 4.
[0124] As Figure 3 shown, consider a drone network consisting of N drones. Each node has a carrier sensing function that can identify and detect interference signals, which enables the node to sense the usage of the surrounding channels and adjust the communication strategy according to the interference situation. The network operates in a distributed manner without a central base station and host nodes, and each node communicates equally. The drone nodes move freely in a three - dimensional space of 1000m * 1000m * 100m to simulate the dynamic changes in the actual environment. The number of drone nodes in the network ranges from 10 to 50 to test the performance of different - scale networks. Each drone node can not only send and receive data but also act as a relay node to help other nodes transmit data.
[0125] As Figure 1 shown, when the network starts, initialization operations are first performed. Each node selects a suitable channel according to the network environment and initializes its competitive window. For example, assume there are 5 channels, the minimum competitive window CW min for channel 1 is 32, the maximum competitive window CW max for channel 1 is 1024, and the initialization values for the other channels are similar. Each node will evaluate the quality of each channel, including the busy - idle rate of the channel, by parsing the status information from the frame headers of each node (as Figure 4 shown) and determine which channel is most suitable for data transmission based on this information.
[0126] Meanwhile, the node will initialize the DDQN model. This model is used to dynamically adjust the competitive window size and determine how to select the most suitable channel and competitive window value in future time slots based on historical data and environmental feedback. All nodes initialize the parameters of the DDQN model to default values and wait for subsequent training and optimization.
[0127] During the data transmission phase, a node determines whether to perform data transmission based on the current channel state and the size of the contention window. Suppose a certain node needs to send data through the first channel. It first calculates the backoff time according to the contention window of the selected channel (assumed to be 32). The formula for generating the backoff time is as follows:
[0128]
[0129] Suppose the backoff time is 320 microseconds (based on the contention window of the channel and the slot time). The node starts waiting. If the first channel is idle during the waiting process, the node starts data transmission; if the first channel is busy, the node continues to wait and adjusts the backoff time.
[0130] The node continuously monitors the channel state during the backoff period. If the channel is in an idle state, the node reduces the backoff time and performs data transmission. If the channel state changes (e.g., becomes busy), the node adjusts the backoff time in real time to avoid collisions during multi-node contention.
[0131] Once the backoff time expires, the node starts sending the data packet. If the data packet is successfully sent, the receiver returns an ACK signal for confirmation. If the ACK is not received, the data transmission fails, and the node enters the retransmission phase.
[0132] If the data transmission fails, the node enters the retransmission phase. At this time, the node needs to adjust the contention window to reduce the probability of collisions. According to the DDQN algorithm, the node calculates the new contention window size based on the current state, including the busy / idle rate, throughput, signal strength, etc. of the channel. For example, the node can double the contention window value until it reaches the maximum value of 1024:
[0133]
[0134] where, is the new contention window size on the first channel calculated by the DDQN algorithm, is the state of the current channel at the k-th time step. Each time of retry, the increase of the contention window can effectively reduce contention collisions and improve the success rate of data transmission.
[0135] The node regenerates the backoff time during retransmission and continues to monitor the channel state until the channel is idle and the data is successfully sent. If the number of retry attempts exceeds the preset maximum number of retry attempts, the node abandons the transmission of the current data packet, discards the data, and starts the transmission of the next data packet.
[0136] In the present invention, each intelligent UAV node uses the DDQN model to dynamically adjust the contention window to achieve optimal data transmission performance. The decision-making process of the node not only depends on the current channel state but also combines historical experience. By training and updating the deep neural network, the strategy is continuously optimized. The following will be combined with Figure 2 Describe this process in detail:
[0137] 1. Node State Observation and Environmental Feedback
[0138] The intelligent UAV node obtains environmental information by real-time monitoring the state of the channel. In each time slot, each node records the following information:
[0139] (1) Channel Busy Rate: By calculating the ratio of the idle time of the channel to the total observation time, the current load state of the channel is obtained.
[0140]
[0141] Among them, t idle is the idle time of the channel, and t total is the total observation time of the channel. A lower CBR value indicates that the channel is relatively idle, while a higher CBR value indicates that the channel is relatively congested.
[0142] (2) Channel Quality: Each node also measures the quality of the signal, such as throughput (data transmission rate), to determine the current transmission quality of the channel.
[0143] (3) Interference Intensity: The node evaluates the current interference level of the channel, usually by analyzing information such as the number of interfering nodes and signal strength in the surrounding network environment.
[0144] The state S t of each node contains this channel information (such as the busy rate, throughput, interference situation, etc.) of the channel. The action taken by the node based on the current state is to select a channel and adjust the contention window accordingly.
[0145] 2. Accumulation and Replay of Historical Experience
[0146] When updating the Deep Neural Network (DNN), the node not only depends on the current state but also uses historical experience to optimize its strategy. Through the experience replay mechanism, the node saves the state, action, reward, and the transfer information of the next state it has experienced. The specific process is as follows:
[0147] (1) Experience Storage: The node stores each decision-making process (state, action, reward, next state) in the experience replay pool, and the experience pool contains the historical data of the node.
[0148] Experience = {S t , a t , R t , S t+1}
[0149] where S t is the current state, a t is the action of node selection (i.e., channel selection and contention window adjustment), R t is the reward after the node takes an action from the current state, and S
[0150] (2) Random Sampling: The node randomly samples a batch of historical experiences from the experience pool and uses these experiences to train the DNN. By random sampling, the node can avoid the overfitting problem that may occur during the training process.
[0151] 3. Training and Updating of the DNN Model
[0152] Through DDQN, the node predicts the future Q value using the current state and updates the DNN according to the experience. The specific process is as follows:
[0153] (1) Calculation of Q Value: For each state and the corresponding action, the node uses the DNN to calculate the Q value of this action:
[0154] Q(S t , a t ) = DNN(S t , a t )
[0155] The DNN predicts the future return (Q value) of each possible action according to the current state S t , and this return represents the long-term reward that the node can obtain after selecting a certain action in the current state.
[0156] (2) Q Value Update: Through the DQN algorithm, the node updates the DNN according to the actual reward and the target Q value. The target Q value is usually calculated by the Bellman equation:
[0157] Q target = R t + γ · Q(S t+1 , a)
[0158] where R t is the immediate reward after taking an action from the current state, γ is the discount factor, representing the weight of future rewards, and max a Q(S t+1 , a) is the maximum Q value obtained by the node when selecting the optimal action in the next state.
[0159] (3) Loss function: To minimize the error between the predicted Q-value and the target Q-value, the node uses the mean squared error as the loss function:
[0160] L = (Q target - Q(S t , a t )) 2
[0161] The node optimizes the DNN through gradient descent, enabling the DNN to more accurately predict the Q-value and gradually adjust the strategy.
[0162] 4. Policy Evaluation and Selection
[0163] In each time slot, the node calculates the Q-values of each channel and the contention window by observing the channel state and combining historical experience, and selects the channel with the highest Q-value and the corresponding contention window for data transmission. The process of policy evaluation and selection is as follows:
[0164] (1) Calculate the total reward: Based on the Q-values predicted by the DNN, the node can evaluate the total rewards for different channel and contention window adjustments.
[0165] (2) Select the optimal policy: The node selects the channel and contention window that maximize the Q-value as the policy for the current time slot. Specifically, the node will select the action with the maximum Q-value according to the current state and all possible actions.
[0166] (3) Execute the policy: The node executes the selected action and obtains an immediate reward during the transmission. Then, based on the error between the reward and the target Q-value, the DNN is updated.
[0167] 5. Model Update and Optimization
[0168] As the training progresses, the DNN model will gradually optimize and learn an optimal policy that can adjust the strategy according to the channel state, historical experience, and contention window to maximize the long-term reward. This process is continuously iterated until the network performance reaches the optimal state.
[0169] The present invention can dynamically optimize channel selection and contention window adjustment in a complex interference environment, thereby improving the network throughput, reducing the latency, optimizing fairness, and enhancing the anti-interference ability of the network.
[0170] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention.
Claims
1. An adaptive channel access method for an unmanned aerial vehicle self-organizing network, characterized in that Including: Initializing the contention window value for each channel; The UAV nodes in the UAV ad-hoc network select the optimal channel and start the backoff process for data transmission; among them, the backoff time is randomly generated according to the current contention window value; if the data transmission fails, the UAV node dynamically adjusts the contention window value according to the reinforcement learning method and regenerates the backoff time.
2. The method for an unmanned aerial vehicle self-organizing network to adaptively access a channel according to claim 1, characterized in that The UAV node dynamically adjusts the contention window value according to the reinforcement learning method, including: Constructing a state space based on the channel busy rate CBR, the contention window popularity, and the channel interference degree sensed by the UAV node, and designing a reward function based on the throughput, fairness, and channel interference degree of the UAV node. The reward function is used to dynamically adjust the contention window to achieve adaptive channel access optimization.
3. The method for self-organizing network adaptive channel access of an unmanned aerial vehicle according to claim 1, wherein The reinforcement learning method includes the double deep Q network DDQN; the UAV nodes in the UAV network select the optimal channel according to the available channel table jointly maintained. The available channel table refers to a list of channels that are not interfered with or occupied and are in an idle state; the data includes the transmission pictures, videos, and text data of the UAV nodes.
4. The self-organizing network adaptive channel access method for an unmanned aerial vehicle according to claim 1, wherein The backoff time includes the backoff time of the i-th channel The backoff time of the i-th channel The calculation method is as follows: where CW i is the current contention window value of the i-th channel, τ is the slot time, and random(*) is a random function, i.e., an integer value is randomly selected from 0 to CW i as the backoff time.
5. The method for self-organizing network adaptive channel access of an unmanned aerial vehicle according to claim 2, wherein CBR is calculated by the following formula: where t idle is the idle time, which refers to the time when the channel is not occupied, and t total is the total observation time, which refers to the difference between the current time and the previous observation time.
6. The method for an unmanned aerial vehicle (UAV) ad-hoc network to adaptively access channels according to claim 2, wherein The contention window CW popularity is obtained by the following method: Counting the contention window values of neighbor UAV nodes and their occurrence frequencies; Sorting the contention window values from high to low according to the occurrence frequency; Calculating its popularity according to the position of the contention window value of the current UAV node in the sorting.
7. The method for self-organizing network adaptive channel access of an unmanned aerial vehicle according to claim 2, wherein The state space includes historical actions and historical rewards, which are used to enhance the agent's perception ability of the dynamic environment.
8. The self-organizing network adaptive channel access method for an unmanned aerial vehicle according to claim 2, wherein The reward function is: Among them, is the contention window reward, which is set based on the contention window popularity; is the average channel busy / idle rate reward, which is set based on the average channel busy / idle rate and a ratio. The ratio refers to the ratio of the current node throughput to the maximum throughput. success is the success flag, reward is the total reward, and β is the coefficient for balancing and .
9. The adaptive channel access method for a UAV ad-hoc network according to claim 8, wherein where rank(CW i ) is the contention window popularity of the i-th channel, CW max is the maximum contention window value, and CW i is the current contention window value of the i-th channel.
10. The self-organizing network adaptive channel access method for an unmanned aerial vehicle according to claim 8, characterized in that Calculated by the following formula: where T is the current throughput, and T max is the maximum throughput, and is the average CBR of neighbor UAV nodes.