Deep Reinforcement Learning for Dynamic Spectrum Access

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Traditional optimization techniques for dynamic spectrum access and spectrum sharing in radio frequency networks are computationally burdensome and fail to capture environmental dynamics, such as user changes and terrain effects, leading to inefficient spectrum utilization as the number of mobile terminals and IoT devices increases, resulting in a shortage of frequency spectrum resources.

Innovation Solution

The implementation of deep reinforcement learning (DRL) techniques to train neural networks for dynamic spectrum access and spectrum sharing, where neural networks receive policies, observe features of telecommunication groups, assign them to channels, and adjust policies based on throughput changes, enabling efficient spectrum allocation and utilization.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If traditional optimization techniques are used for spectrum allocation, then computational burden increases, but spectrum sharing efficiency does not improve sufficiently

Engineering Contradiction:
Improvespectrum sharing efficiencyVSAvoidcomputational burden
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent replaces traditional mathematical optimization algorithms with deep reinforcement learning neural networks. The neural network learns optimal spectrum allocation policies through environmental interaction and feedback, substituting complex analytical optimization with a data-driven adaptive system that achieves better performance with reduced computational burden during operation.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent changes the fundamental parameters of the optimization approach by using neural network weights and policies instead of traditional optimization variables. The system learns optimal parameters through training with simulated environmental feedback, transforming the problem from analytical optimization to adaptive learning, thereby improving spectrum sharing efficiency while reducing real-time computational complexity.

Inventive Principle:
Principle #35Parameter changes

2Adaptability or versatility

If traditional optimization techniques are used, then environmental dynamics such as user changes and terrain effects cannot be captured, but implementing adaptive techniques increases computational complexity

Engineering Contradiction:
Improveenvironmental dynamics captureVSAvoidcomputational complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The neural network performs self-learning and self-adaptation through interaction with the environment. It automatically captures environmental dynamics including user changes, terrain effects, and spectrum conditions by processing sensory inputs and adjusting its policies based on received feedback, eliminating the need for explicit programming of environmental models.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system implements continuous feedback loops where the neural network observes environmental states, executes spectrum allocation actions, receives throughput feedback, and updates its policies accordingly. This feedback mechanism enables the system to adapt to changing environmental conditions dynamically without requiring complex real-time computational models of the environment.

Inventive Principle:
Principle #23Feedback

3Quantity of substance

If more spectrum is allocated to accommodate increasing mobile terminals and IoT devices, then spectrum shortage is addressed, but spectrum utilization efficiency decreases due to lack of efficient allocation

Engineering Contradiction:
Improvespectrum resourcesVSAvoidspectrum utilization efficiency
Core Design Contradiction:
Quantity of substanceVSProductivity

Solution Approach 1:

The patent implements dynamic spectrum allocation where the neural network continuously adapts channel assignments and power levels based on real-time environmental conditions, user demands, and spectrum usage patterns. This dynamic approach allows efficient utilization of available spectrum resources by constantly optimizing allocations rather than using static assignments, thereby maintaining high utilization efficiency as the number of devices increases.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system segments the spectrum into multiple channels and dynamically assigns different channels to different user groups based on their specific needs, locations, and interference conditions. The neural network learns to optimally partition and allocate spectrum segments to maximize overall utilization efficiency while accommodating a growing number of mobile terminals and IoT devices.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS12067487B2Method and apparatus employing distributed sensing and deep learning for dynamic spectrum access and spectrum sharing
Publication Date: 2024.08.20 CACI INC FEDERAL
  • US12067487B2 patent drawing
  • US12067487B2 patent drawing
  • US12067487B2 patent drawing

AI summary

The present application describes a method for training a neural network via deep reinforcement learning (DRL) in a RF network. The method includes a step of receiving, via the neural network, a policy from a third party. The method also includes a step of receiving, via the neural network, features of plural telecommunication groups located in an RF network. The method also includes a step of observing, via the neural network, a graphical representation of the received features of the plural telecommunication groups in the RF network. The method further includes a step of assigning, based on the observation, one of the plural telecommunication groups to one of plural channels in the RF network. The method even further includes a step of determining, via the neural network, a change in throughput of the RF network based on the assignment. The method yet even further incudes as step of adjusting, based on the determined change in throughput, the policy received from the third party.