Multi-Camera Image Fusion on a Virtual Surface for Distant Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Fusion-based environmental perception methods, such as bird's-eye view representation, suffer from limited accuracy and high memory requirements when detecting objects far away from image sensors, and face challenges in object assignment and tracking between cameras.
Innovation Solution
A method that projects images from multiple cameras onto a virtual surface, using a neural network to generate a virtual overall image, allowing for early fusion and overcoming the limitations of bird's-eye view representation by maintaining accuracy and reducing memory requirements, while also addressing object assignment and tracking issues.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If bird's-eye view representation is used for fusion, then object assignment and tracking between cameras is enabled, but memory requirements increase significantly and accuracy decreases for distant objects
Solution Approach 1:
The patent transitions from representing the environment in a 2D bird's-eye view plane to mapping images onto a 3D virtual surface that surrounds the agent. This dimensional change allows the virtual surface to be positioned close to distant objects, reducing the need to map large empty spaces and thereby reducing memory requirements while maintaining accurate object representation.
Solution Approach 2:
The patent introduces a virtual surface as an intermediary between the camera images and the final bird's-eye view representation. Images from multiple cameras are first projected onto this virtual surface, which acts as an intermediate mapping layer, before being transformed into the final bird's-eye view. This intermediary structure enables more efficient memory usage by allowing selective positioning and scaling.
2Reliability
If bird's-eye view representation is used for fusion, then a unified environment model is created, but accuracy is limited for objects far away from the agent
Solution Approach 1:
By moving from a flat 2D bird's-eye view to a 3D virtual surface that can be positioned at varying distances from the agent, the system can maintain appropriate scaling and resolution for objects at different ranges. The virtual surface allows distant objects to be represented with sufficient detail without requiring excessive memory to map the entire large-area space.
3Adaptability or versatility
If multiple cameras with different detection ranges are used, then comprehensive environment coverage is achieved, but object assignment between cameras becomes complex
Solution Approach 1:
The virtual surface serves as a common intermediary coordinate system where images from multiple cameras with different detection ranges are projected. This unified intermediate representation simplifies object assignment by providing a consistent reference frame, eliminating the need for complex direct mapping between different camera coordinate systems.
Data Source
AI summary
A method for detecting an environment using images from at least two image sensors. The method includes: providing a first image of the environment from a first image sensor; providing a second image of the environment from a second image sensor; wherein the first image sensor and the second image sensor are configured to detect the environment with different detection ranges; defining a virtual surface, which is arranged between the environment and the at least two image sensors; generating a virtual overall image on the virtual surface based on a projection transformation of respective pixels of the first image and a projection transformation of respective pixels of the second image from a relevant image plane of the relevant image sensor onto the virtual surface; and representing the environment based on the virtual overall image and on a neural network trained to represent the environment, to detect the environment.

