Aerial View Image Fusion for Occlusion-Aware Vehicle Perception

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing image processing methods for autonomous driving fail to accurately capture occluded objects and require separate models for each perception task, leading to reduced accuracy and prolonged processing times due to inefficient use of resources.

Innovation Solution

An image processing method that combines temporal and spatial information using a variability attention mechanism to generate an aerial view feature, allowing for simultaneous processing of multiple perception tasks with a shared model, thereby improving accuracy and efficiency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If separate models are trained for each perception task, then each task can be processed with dedicated optimization, but the image processing time is prolonged and processing efficiency is reduced

Engineering Contradiction:
Improveperception task accuracyVSAvoidimage processing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent merges multiple perception task models into a single unified model that processes multiple tasks simultaneously. The model architecture integrates different perception functions (object detection, lane detection, traffic light recognition, etc.) into one framework, allowing parallel processing of multiple tasks from the same input images, thereby reducing processing time while maintaining accuracy through shared feature extraction layers

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent creates a universal perception model that can handle multiple perception tasks through a single system. The unified model is designed with multi-functional capabilities to perform various perception operations (detection, segmentation, classification) on the same input data, eliminating the need for separate dedicated models for each task and improving overall processing efficiency

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Productivity

If two-dimensional image information is converted into aerial view feature, then perception tasks can be processed, but occluded objects cannot be captured well and accuracy is reduced

Engineering Contradiction:
Improveperception task processing capabilityVSAvoidaerial view feature accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent transitions from two-dimensional image processing to three-dimensional aerial view feature representation. By introducing the third dimension (depth/elevation) through aerial view features, the system can represent occluded objects and spatial relationships more accurately, allowing perception tasks to capture objects that would be hidden in traditional 2D images while maintaining processing capability

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS12469277B2Image processing method, apparatus and device, and computer-readable storage medium
Publication Date: 2025.11.11 SHANGHAI ARTIFICIAL INTELLIGENCE INNOVATION CENT
  • US12469277B2 patent drawing
  • US12469277B2 patent drawing
  • US12469277B2 patent drawing

AI summary

The embodiments of the present application discloses an image processing method, an image processing apparatus, an image processing device and a computer-readable storage medium. The method comprises: obtaining a plurality of view images captured by a plurality of acquisition apparatuses at different views on a vehicle and an aerial view feature at a previous moment of a current moment; extracting temporal information from the aerial view feature at the previous moment according to a preset aerial view query vector, extracting spatial information from a plurality of view image features corresponding to the plurality of view images, and combining the temporal information and the spatial information to generate an aerial view feature at the current moment, wherein the preset aerial view query vector corresponds to a three-dimensional physical world which is a preset range away from the vehicle in a real scene at the current moment.