Autonomous Intersection Navigation Using Reinforcement Learning Feedback

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Self-driving vehicles face low traffic efficiency when navigating intersections without traffic lights due to conservative driving strategies, which result in slower passage and reduced traffic flow.

Innovation Solution

An intersection traffic control method utilizing an instruction learning model based on reinforcement learning, which acquires vehicle signals from a first vehicle and nearby vehicles to determine optimal action instructions, balancing safety, efficiency, and comfort by calculating scores for traffic indicators and performing weighted summation to determine the best course of action.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If self-driving vehicles adopt conservative strategies to pass through intersections without traffic lights, then safety is improved, but traffic efficiency deteriorates

Engineering Contradiction:
ImprovesafetyVSAvoidtraffic efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The system implements real-time feedback by continuously acquiring vehicle signals from multiple vehicles, evaluating safety dynamically, and adjusting navigation instructions accordingly. The instruction learning model processes current traffic conditions and provides actionable feedback to optimize both safety and efficiency at intersections without traffic lights.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system enables self-service by allowing self-driving vehicles to autonomously determine optimal navigation instructions through the instruction learning model without relying on external traffic light control. The vehicle independently evaluates its own safety and efficiency metrics to make real-time decisions at intersections.

Inventive Principle:
Principle #25Self-service

2Reliability

If self-driving vehicles pass through intersections at lower speeds, then safety is improved, but traffic flow deteriorates

Engineering Contradiction:
ImprovesafetyVSAvoidtraffic flow
Core Design Contradiction:
ReliabilityVSSpeed

Solution Approach 1:

The system applies dynamics by enabling real-time adjustment of vehicle speed and navigation instructions based on current traffic conditions. The instruction learning model dynamically optimizes speed recommendations to balance safety requirements with maintaining adequate traffic flow, rather than enforcing fixed conservative speed limits.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system implements parameter changes by modifying navigation instructions based on real-time evaluation of vehicle signals and traffic conditions. The instruction learning model adjusts speed and trajectory parameters dynamically to achieve optimal balance between safety and traffic flow efficiency at intersections without traffic lights.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS11348455B2Intersection traffic control method, apparatus and system
Publication Date: 2022.05.31 GUANGZHOU AUTOMOBILE GROUP CO LTD
  • US11348455B2 patent drawing
  • US11348455B2 patent drawing
  • US11348455B2 patent drawing

AI summary

An intersection traffic control method, apparatus and system are provided. The method includes that: a vehicle signal of a first vehicle at an intersection and a vehicle signal of a second vehicle located in a set zone in proximity to the intersection are acquired; the vehicle signal of the first vehicle and the vehicle signal of the second vehicle are input into an instruction learning model trained in advance based on a reinforcement learning principle, and a score of a preset traffic indicator of the first vehicle after executing a respective candidate action instruction is calculated; a reward of the first vehicle when executing the respective candidate action instruction is acquired according to the score of the preset traffic indicator, a candidate action instruction corresponding to a maximum reward is determined as an output result of the instruction learning model, and a next action instruction is determined according to the output result; and navigation of the first vehicle through the intersection is controlled according to the next action instruction.