Deep Double-Q Reinforcement Learning for Adaptive Anti-Jamming

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Cognitive radio devices face significant communication performance deterioration due to radio-jamming attacks from smart jammers, and existing adaptive anti-jamming methods are spectral inefficient, high in energy cost, and complex, with traditional Q-learning being inefficient for large state and action problems.

Innovation Solution

The implementation of a Deep Double-Q Reinforcement learning system for adaptive anti-jamming communications, comprising a wideband spectrum sensing block, an anti-jamming strategy generating block using a prediction Q-neural network, and an anti-jamming strategy implementation block, which selects optimal transmission channels and power levels to maximize successful transmission rates while minimizing power consumption.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If traditional Q-learning is used for anti-jamming communication, then the system can learn optimal transmission strategies, but the system becomes inefficient when the number of states and actions is very large

Engineering Contradiction:
Improveanti-jamming communication reliabilityVSAvoidlearning efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent replaces the traditional Q-learning algorithm with Deep Double-Q Reinforcement Learning, substituting the mechanical search process through Q-tables with a neural network-based deep learning system. This substitution enables the system to handle large state and action spaces efficiently by using neural networks to approximate Q-values, thereby resolving the contradiction between learning reliability and productivity.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent changes the fundamental parameters of the learning system by introducing deep neural networks with multiple layers (input layer, hidden layers, output layer) to represent the Q-function. This parameter transformation allows the system to scale from small to large state-action spaces while maintaining learning efficiency, directly addressing the contradiction between reliability and productivity.

Inventive Principle:
Principle #35Parameter changes

2Reliability

If spread spectrum based techniques are used for anti-jamming, then communication robustness is improved, but spectral utilization efficiency deteriorates

Engineering Contradiction:
Improvecommunication robustnessVSAvoidspectral utilization efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent applies dynamics by enabling the cognitive radio device to dynamically select optimal transmission parameters (channel, power level, modulation scheme) based on real-time spectrum sensing and deep reinforcement learning. This dynamic adaptation allows the system to achieve robustness against jamming while efficiently utilizing available spectrum, resolving the contradiction between reliability and productivity.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent changes transmission parameters dynamically based on learned strategies from deep reinforcement learning. By continuously optimizing parameters such as transmission power, selected channel, and modulation scheme according to the current spectral environment, the system achieves both robustness and spectral efficiency, eliminating the trade-off between these two characteristics.

Inventive Principle:
Principle #35Parameter changes

3Reliability

If traditional adaptive anti-jamming methods are used, then jamming resistance is improved, but device complexity and energy consumption increase

Engineering Contradiction:
Improvejamming resistanceVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent substitutes complex traditional anti-jamming algorithms with a deep reinforcement learning framework that uses neural networks to automatically learn optimal strategies. This substitution reduces system complexity by replacing manual algorithm design with an automated learning process that adapts to jamming conditions, thereby improving jamming resistance while reducing device complexity.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent implements self-service by enabling the cognitive radio device to autonomously learn and adapt anti-jamming strategies through deep reinforcement learning without requiring complex external control systems. The system self-adjusts transmission parameters based on spectrum sensing and learned policies, reducing device complexity while maintaining high jamming resistance.

Inventive Principle:
Principle #25Self-service

4Reliability

If transmission power is increased to overcome jamming, then communication reliability is improved, but power consumption increases

Engineering Contradiction:
Improvecommunication reliabilityVSAvoidpower consumption
Core Design Contradiction:
ReliabilityVSUse of energy by moving object

Solution Approach 1:

The patent optimizes transmission power by dynamically adjusting it based on deep reinforcement learning decisions. The neural network learns to select optimal power levels that achieve reliable communication while minimizing energy consumption, resolving the contradiction between reliability and power usage by finding the optimal balance point through learned strategies.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent applies local quality by enabling the system to adjust transmission parameters locally based on specific spectral conditions and learned strategies. Instead of using fixed high power levels, the system adapts power consumption to local communication needs, achieving reliability only when necessary while minimizing overall energy usage.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS12149343B2Method and apparatus for adaptive anti-jamming communications based on deep double-Q reinforcement learning
Publication Date: 2024.11.19 VIETTEL GRP
  • US12149343B2 patent drawing
  • US12149343B2 patent drawing
  • US12149343B2 patent drawing

AI summary

In order to avoid various jamming attacks from intelligent jammers in modern complex wireless environments, a system and method is presented for a user radio to generate and implement an adaptive anti-jamming communication strategy. The said adaptive anti-jamming communication strategy is obtained via the training process for a specific neural network using Deep Double-Q Reinforcement learning algorithm in the strategy generation phase. The objective of this process is to discover a strategy to select the optimal radio action including transmission channel and transmission power for the user radio, which is changed adaptively to different jamming patterns to maximize the successful transmission rate (“jamming-free”) while retaining the power consumption of user radio as low as possible. In the strategy implementation phase, the user radio chooses an appropriate radio action based on output of trained neural network after the training process; thus, achieves robust and efficient communications against diverse complex jamming scenarios.