Deep Reinforcement Learning for Autonomous Vehicle On-Ramp Merging

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Autonomous vehicles face challenges in safely navigating on-ramp merging zones due to unpredictable factors such as sudden stops and lane changes, leading to frequent collisions.

Innovation Solution

A deep reinforcement learning-based vehicle action decision method and apparatus that utilizes an information observation unit, policy execution unit, and reward determination unit to make decisions on acceleration control and lane changes, considering factors like speed, lane change, safety distance, and traffic density to enhance merging strategies.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If autonomous vehicles use traditional control methods in on-ramp merging zones, then the system structure is simple, but collision risk increases due to inability to handle sudden stops and lane changes

Engineering Contradiction:
Improvecollision riskVSAvoiddecision-making system complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent replaces traditional mechanical control systems with an intelligent decision-making system based on deep reinforcement learning. The neural network processes observation information (vehicle states, traffic conditions) and outputs optimized control actions (acceleration, lane changes) to navigate merging zones safely, substituting rule-based mechanical logic with adaptive intelligent processing.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent dynamically adjusts control parameters (acceleration rate, lane change timing, safety distance) based on real-time observation information. The deep reinforcement learning model learns optimal parameter combinations through training with reward functions that penalize collisions and reward successful merges, enabling adaptive parameter optimization for varying traffic conditions.

Inventive Principle:
Principle #35Parameter changes

2Reliability

If autonomous vehicles maintain large safety distances in merging zones, then collision risk decreases, but traffic flow efficiency deteriorates due to delayed merging

Engineering Contradiction:
Improvesafety distance complianceVSAvoidmerging efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent implements dynamic adjustment of safety distances based on real-time traffic conditions and predicted vehicle behaviors. The deep reinforcement learning model learns to adapt safety margins dynamically - maintaining larger distances when risk is high and reducing them when conditions permit - rather than using fixed conservative thresholds, enabling both safety and efficiency.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent employs reward functions that provide feedback on safety distance compliance and merging outcomes. The model receives positive rewards for successful merges and negative rewards for collisions or excessive delays, enabling it to learn optimal safety distance strategies that balance risk mitigation with merging efficiency through continuous feedback loops.

Inventive Principle:
Principle #23Feedback

3Reliability

If autonomous vehicles make conservative driving decisions in merging zones, then safety increases, but traffic flow efficiency decreases due to delayed merges and acceleration

Engineering Contradiction:
ImprovesafetyVSAvoidtraffic flow efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent uses prediction modules to anticipate future vehicle behaviors and traffic conditions before making decisions. By predicting potential hazards (sudden stops, aggressive lane changes) and planning ahead, the system can take preliminary actions that ensure safety while minimizing delays, such as early lane positioning or gradual acceleration strategies.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent optimizes driving parameters (acceleration rate, speed, lane change timing) dynamically based on learned patterns from training data. The deep reinforcement learning model identifies optimal parameter combinations that achieve successful merges with minimal delay, replacing conservative fixed-parameter approaches with adaptive optimized parameters.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS12534084B2Method and apparatus for determining behavior based on deep reinforcement learning for autonomous vehicle merging strategy
Publication Date: 2026.01.27 FOUND OF SOONGSIL UNIV IND COOP
  • US12534084B2 patent drawing
  • US12534084B2 patent drawing
  • US12534084B2 patent drawing

AI summary

A deep reinforcement learning-based vehicle action decision apparatus for merging strategy of an autonomous vehicle in an on-ramp merging zone is disclosed. The deep reinforcement learning-based vehicle action decision apparatus comprises an information observation unit for collecting observation information from a sensing module or roadside unit (RSU) of an autonomous vehicle; a policy execution unit for deciding on a current action, including acceleration control and lane change of the autonomous vehicle, based on the current observation information and policy; and a reward determination unit for determining a reward according to the current observation information, the current action, and the next observation information according to the current action, wherein reward in the reward determination unit is determined through a reward term related to speed, lane change, safety distance compliance, and an accident of the autonomous vehicle and a merge reward term related to merge of an autonomous vehicle in the on-ramp merging zone.