Visual Pattern Attacks on Reinforcement-Learning Driving Agents

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing deep reinforcement learning-based autonomous driving systems are vulnerable to adversarial attacks, particularly those using visually learned patterns, which can hijack the vehicle's control policies without direct access to perception modules, posing risks in real-world scenarios.

Innovation Solution

A targeted adversarial attack methodology using learned visual patterns is developed, where an attacker places static adversarial objects in the environment to mislead the agent into a specific state by optimizing a differentiable dynamical model and perturbing the environment's visual input, ensuring the attack is effective within a specified time window.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If targeted adversarial attacks are implemented using learned visual patterns, then the attack effectiveness against autonomous driving agents is improved, but the system vulnerability to manipulation increases

Engineering Contradiction:
Improveattack effectivenessVSAvoidsystem vulnerability
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The attack method performs preliminary actions by learning the environment dynamics model and identifying vulnerable visual patterns before executing the actual adversarial attack. The system pre-processes the environment by capturing state transitions, learning dynamics through multiple rollouts, and identifying key visual features that influence agent behavior, enabling more effective targeted attacks.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The attack methodology creates a copy of the environment dynamics through learned visual patterns and dynamical models. By replicating the environment's behavior through training data and learned models, the attacker can simulate and predict agent responses to adversarial inputs without directly manipulating the physical environment, thereby improving attack precision.

Inventive Principle:
Principle #26Copying

2Measurement precision

If the attack plan is generated using trained AI models with learned dynamics, then the precision of state manipulation is improved, but the computational complexity increases

Engineering Contradiction:
Improvestate manipulation precisionVSAvoidcomputational complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system performs preliminary computational work by training AI models to learn environment dynamics before the actual attack execution. The dynamics model is pre-trained using rollout data captured from the environment, enabling the attack generator to quickly compute precise manipulation strategies without performing complex real-time simulations during attack execution.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The attack methodology replaces direct environmental manipulation with AI model-based prediction and simulation. Instead of physically testing different attack configurations in the real environment, the system uses learned dynamical models to simulate agent responses and generate attack plans computationally, reducing the need for repeated physical experiments.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Stability of the object's composition

If multiple rollouts with variable noise are performed to generate training data, then the robustness of the attack model is improved, but the time required for data generation increases

Engineering Contradiction:
Improveattack model robustnessVSAvoiddata generation time
Core Design Contradiction:
Stability of the object's compositionVSLoss of time

Solution Approach 1:

The system employs periodic action by performing multiple structured rollouts with variable noise additions at different stages of model training. Instead of continuous data collection, the methodology uses discrete, periodic sampling of state transitions with controlled noise variations, efficiently building robust training data while managing computational time.

Inventive Principle:
Principle #19Periodic action

Solution Approach 2:

The attack methodology uses partial action by performing a limited number of rollouts with variable noise rather than exhaustive sampling. The system adds noise at specific thresholds and performs rollouts only when certain conditions are met, achieving sufficient model robustness without the time cost of complete environmental exploration.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS12632564B2Targeted attacks on deep reinforcement learning-based autonomous driving with learned visual patterns
Publication Date: 2026.05.19 THE ARIZONA BOARD OF REGENTS ON BEHALF OF THE UNIV OF ARIZONA
  • US12632564B2 patent drawing
  • US12632564B2 patent drawing
  • US12632564B2 patent drawing

AI summary

A system may be configured for implementing targeted attacks on deep reinforcement learning-based autonomous driving with learned visual patterns. In some examples, processing circuitry receives first input specifying an initial state for a driving environment and user configurable input specifying a target state. Processing circuitry may generate a representative dataset of the driving environment by performing multiple rollouts of the vehicle through the driving environment, including performing an action for the vehicle from the initial state with variable strength noise added to determine a next state for each rollout resulting from the action. Processing circuitry may train an artificial intelligence model to output a next predicted state based on the representative dataset as training input. In such an example, processing circuitry outputs from the artificial intelligence model, an attack plan against the autonomous driving agent to achieve the target state from the initial state.