UAV D2D Scheduling via Sparse Convolution and Reinforcement Learning
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In D2D communication over IoT networks, determining D2D link scheduling that maximizes information transmission while efficiently sharing spectrum between links within a given area is challenging due to high computational complexity and the difficulty in obtaining channel state information (CSI).
Innovation Solution
A D2D scheduling method and apparatus in a UAV-based IoT network that schedules transmission links without CSI, using a sparse convolution model and reinforcement learning-based scheduling policy to extract feature maps and make scheduling decisions based on UAV support topology information, thereby reducing complexity and increasing speed.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If D2D resource scheduling is performed using all necessary network information (channel and interference status), then scheduling accuracy is improved, but computational complexity increases significantly
Solution Approach 1:
The patent extracts only the necessary features from the geographical map (transmitter and receiver locations) rather than processing complete channel state information. This selective extraction reduces the data volume and computational complexity while maintaining sufficient scheduling accuracy through the sparse convolution model that processes only location-based features.
Solution Approach 2:
The patent uses lightweight neural network models (actor and critic networks) that require minimal computational resources compared to traditional scheduling algorithms. These models are designed to be computationally efficient, enabling real-time scheduling decisions without requiring extensive processing power.
2Measurement precision
If CSI is collected and transmitted to control devices, then scheduling precision is improved, but power consumption and control overhead increase
Solution Approach 1:
The patent extracts only location information from the geographical map rather than collecting and transmitting complete CSI data. This selective extraction eliminates the need for extensive CSI measurement and transmission, significantly reducing power consumption and control overhead while maintaining scheduling precision through location-based feature extraction.
Solution Approach 2:
The UAV performs scheduling decisions autonomously using the geographical map and neural network models without requiring centralized control or extensive CSI transmission. This self-service approach eliminates the need for complex control signaling and reduces overall system power consumption.
3Measurement precision
If traditional D2D scheduling algorithms are used, then scheduling accuracy is maintained, but execution speed decreases due to high computational complexity
Solution Approach 1:
The patent employs lightweight neural network models that are computationally efficient and can be executed rapidly compared to traditional scheduling algorithms. These simplified models maintain scheduling accuracy while enabling real-time decision-making through their reduced computational requirements.
Solution Approach 2:
The actor and critic networks are pre-trained offline using historical data, so that during actual operation, scheduling decisions can be made quickly by applying the pre-learned models to current geographical map data. This preliminary training separates the computationally intensive learning phase from the fast execution phase.
Data Source
AI summary
A D2D scheduling method in a UAV-based IoT network includes: (a) acquiring a geographical map for all transmission links of a D2D network within a network coverage area; (b) applying the geographical map to a sparse convolution model to extract a feature map; (c) defining the feature map for a time slot t as a state St, and then inputting the feature map to actor and critic networks of a reinforcement learning-based scheduling policy learning model, respectively, and selecting a scheduling decision At for the D2D transmission link based on a scheduling output of the actor network and a greedy strategy; and (d) transmitting the scheduling decision At to the D2D network and then receiving reward when the scheduling decision At is applied by the D2D network, wherein the reward is calculated as a total achievable transmission rate.


