Camera-Radar Fusion for Accurate Traffic Sign Classification

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing autonomous vehicles face challenges in accurately and efficiently detecting and classifying traffic signs in driving environments, particularly due to limitations in sensor modalities such as cameras and radars.

Innovation Solution

A system and method that utilize a combination of camera and radar images processed by respective neural networks to generate features, which are then aggregated and processed by a BEV model to identify and classify traffic signs accurately.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If camera and radar sensors are used for traffic sign detection, then cost is reduced compared to lidar, but measurement precision and detection accuracy are worsened

Engineering Contradiction:
ImprovecostVSAvoiddetection accuracy
Core Design Contradiction:
Ease of manufactureVSMeasurement precision

Solution Approach 1:

The patent combines camera and radar sensors into a fused sensing system where camera images provide visual information for traffic sign detection and radar images provide depth and spatial information. The neural networks process both modalities simultaneously and fuse their features to achieve accurate traffic sign detection and classification, resolving the accuracy limitation of individual sensors while maintaining cost effectiveness.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent creates a composite sensing approach by integrating multiple sensor types (camera and radar) with different characteristics. The camera captures optical information while radar captures electromagnetic reflection data, and their combined processing through neural networks creates a composite detection capability that overcomes the limitations of each individual sensor type.

Inventive Principle:
Principle #40Composite materials

2Device complexity

If traditional single-modality sensors are used, then device complexity is reduced, but detection speed and accuracy of traffic signs are worsened

Engineering Contradiction:
Improvesensor configurationVSAvoiddetection speed
Core Design Contradiction:
Device complexityVSProductivity

Solution Approach 1:

The patent implements dynamic processing through neural networks that adaptively process camera and radar images in real-time. The system dynamically adjusts feature extraction and fusion based on input data characteristics, enabling fast detection speeds while maintaining the multi-modality sensor configuration for improved accuracy.

Inventive Principle:
Principle #15Dynamics

3Productivity

If camera and radar features are processed separately, then processing efficiency is maintained, but classification accuracy of traffic signs is worsened

Engineering Contradiction:
Improveprocessing efficiencyVSAvoidclassification accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent merges the processing streams of camera and radar features through a fusion neural network. Instead of processing features separately to the end, the system combines features from both modalities at intermediate processing stages, allowing the network to leverage complementary information from both sensors for improved classification accuracy while maintaining processing efficiency through parallel feature extraction followed by fusion.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS20250199121A1Detection and classification of traffic signs using camera-radar fusion
Publication Date: 2025.06.19 WAYMO LLC
  • US20250199121A1 patent drawing
  • US20250199121A1 patent drawing
  • US20250199121A1 patent drawing

AI summary

The disclosed systems and techniques facilitate efficient detection and classification of traffic signs in driving environments. The disclosed techniques include, obtaining, using a sensing system of a vehicle a first set of perspective camera images of an environment and a second set of radar images of the environment. The techniques further include generating, using a first neural network, one or more camera features characterizing the first set of images, generating, using a second neural network, one or more radar features characterizing the second set of images, and processing the one or more camera features and the one or more radar features to obtain an identification of one or more traffic signs in the environment.