Prior-Knowledge Double-Action Reinforcement Learning for Spectrum Access

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Reinforcement learning for spectrum management faces challenges such as the 'cold start' problem and the inability to leverage prior knowledge of the environment, leading to inefficient spectrum access in complex electromagnetic environments.

Innovation Solution

A spectrum access method using prior knowledge-based double-action reinforcement learning, which involves evaluating and screening prior knowledge, initializing a Q-table, and performing Q-learning by decomposing the action space into two dimensions, updating the Q-table with biased information, and adjusting it using reward values to guide the agent towards optimal actions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If reinforcement learning is used for spectrum management, then the system can adapt to complex electromagnetic environments, but the cold start problem reduces spectrum access efficiency

Engineering Contradiction:
Improveadaptability to electromagnetic environmentVSAvoidspectrum access efficiency
Core Design Contradiction:
Adaptability or versatilityVSProductivity

Solution Approach 1:

The patent applies preliminary action by evaluating and screening prior knowledge before the reinforcement learning process begins. The Q-table is initialized using this screened prior knowledge, which prepares the system in advance to avoid the cold start problem and improve spectrum access efficiency from the beginning of operation.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent implements feedback mechanisms through the reinforcement learning process where the agent receives reward signals based on its spectrum access decisions. The Q-table is continuously updated using these feedback rewards, allowing the system to adapt and improve its spectrum access strategy over time while maintaining efficiency.

Inventive Principle:
Principle #23Feedback

2Device complexity

If prior knowledge is not utilized, then the reinforcement learning process is simpler, but the convergence speed of the algorithm is slower

Engineering Contradiction:
Improvesimplicity of reinforcement learning processVSAvoidalgorithm convergence time
Core Design Contradiction:
Device complexityVSLoss of time

Solution Approach 1:

The patent performs preliminary evaluation and screening of prior knowledge before the main reinforcement learning process. This preliminary action initializes the Q-table with useful prior information, which accelerates convergence without significantly complicating the overall learning process structure.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent changes the initial state parameters of the Q-table from random values to values derived from screened prior knowledge. This parameter change in the initialization stage reduces the number of iterations needed for convergence while maintaining the fundamental reinforcement learning algorithm structure.

Inventive Principle:
Principle #35Parameter changes

3Productivity

If the action space is decomposed into two dimensions, then the learning process becomes more directed, but the action selection complexity increases

Engineering Contradiction:
Improvelearning efficiencyVSAvoidaction space structure
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent applies segmentation by dividing the action space into two distinct dimensions: channel selection and time slot selection. This segmentation makes the learning process more directed and efficient by breaking down the complex decision-making into manageable components, while the Q-table structure handles the increased complexity systematically.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces a dimensional structure to the action space by creating a two-dimensional action framework. This dimensional organization improves learning efficiency by providing clearer guidance for the agent, while the structured Q-table manages the complexity through its multi-dimensional state-action value storage.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

4Reliability

If biased information is used to update the Q-table, then wrong actions are reduced, but the update process becomes more complex

Engineering Contradiction:
Improveaccuracy of action selectionVSAvoidQ-table update process
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent uses biased information as a feedback mechanism during Q-table updates. The screening results of prior knowledge provide biased guidance that reduces wrong actions by emphasizing more promising directions in the learning process, while the update rules incorporate this bias in a systematic way that manages complexity.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent modifies the Q-table update parameters by incorporating biased information from screened prior knowledge. This parameter modification improves action selection accuracy by guiding the learning toward more promising actions, while the structured update rules maintain manageable complexity through consistent application of the bias.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS12402015B2Spectrum access method and system using prior knowledge-based double-action reinforcement learning
Publication Date: 2025.08.26 NAT UNIV OF DEFENSE TECH
  • US12402015B2 patent drawing
  • US12402015B2 patent drawing
  • US12402015B2 patent drawing

AI summary

The present disclosure provides a spectrum access method and system using prior knowledge-based double-action reinforcement learning, and belongs to the technical field of electromagnetic spectrum. The method includes evaluating and screening prior knowledge, initializing a Q-table, and confirming a current state; and performing Q-learning by: firstly, decomposing an action space into two dimensions with an action in one dimension defined as a channel chosen by an agent, and an action in the other dimension defined as a number of time slots of an access channel, and choosing actions in turn according to the dimensions; then performing spectrum access according to the actions chosen; and finally, updating the Q-table in combination with biased information, wherein the biased information is a reward value. The system is configured to implement the proposed method. By adoption of the method, better performance is achieved, and the efficiency of spectrum access can be improved.