Spatial Geometry Estimation From Single-Frame Road Scene Images
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for acquiring 3D perception information of road surfaces in assisted and automated driving scenarios face challenges in accuracy and computational efficiency, particularly when dealing with moving objects and requiring interframe pose operations.
Innovation Solution
A method and apparatus for generating a spatial geometric information estimation model that utilizes a single frame of image to train a model by determining coordinates and annotation spatial geometric information, allowing for accurate representation of 3D space states and reducing the need for interframe pose operations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If interframe pose operations are used for 3D perception information estimation, then measurement precision may be improved, but device complexity and computational demands increase
Solution Approach 1:
The patent extracts and removes the interframe pose operation step from the traditional multi-frame processing pipeline. By using only single-frame images for model training and estimation, it eliminates the need for complex interframe pose calculations while maintaining 3D perception accuracy through a specially designed spatial geometric information estimation model
Solution Approach 2:
The patent segments the 3D perception problem into independent single-frame estimation tasks rather than requiring sequential multi-frame processing. Each frame can be processed independently through the trained model, dividing the complex interframe operation into simpler, parallelizable single-frame operations
2Measurement precision
If multi-frame image sequences are processed for 3D perception, then measurement precision improves, but productivity decreases due to higher computational demands
Solution Approach 1:
The patent removes the time-consuming interframe pose operation from the processing pipeline by using only single-frame images. This extraction of the essential estimation task from the complex multi-frame processing significantly reduces computational demands and improves processing speed while maintaining accuracy
Solution Approach 2:
The patent performs preliminary action by training the spatial geometric information estimation model in advance using multi-frame data and ground truth annotations. Once trained, the model can perform rapid single-frame estimation without requiring real-time interframe pose calculations, achieving both high accuracy and efficiency
3Device complexity
If single frame image is used for model training, then device complexity is reduced, but measurement precision may deteriorate
Solution Approach 1:
The patent changes the parameter representation by using spatial geometric information (depth, height, disparity) as direct training targets instead of requiring complex interframe pose parameters. The model learns to estimate these geometric parameters directly from single-frame images, achieving accurate 3D perception without complex operational parameters
Solution Approach 2:
The patent introduces a spatial geometric information estimation model as an intermediary that bridges single-frame images and 3D perception information. This mediator model, trained with ground truth annotations, enables accurate depth and height estimation from single frames without requiring direct interframe pose operations
Data Source
AI summary
Disclosed in embodiments of the present disclosure are a method and an apparatus for generating a spatial geometric information estimation model, a computer-readable storage medium, and an electronic device. The method includes: acquiring point cloud data collected for a preset scene and a scene image captured for the preset scene; determining coordinates corresponding to the point cloud data in a camera coordinate system corresponding to the scene image; determining, based on the coordinates, annotation spatial geometric information of a target pixel in the scene image corresponding to the point cloud data; and training an initial model by using the scene image as an input of a preset initial model and the annotation spatial geometric information of the target pixel as an expected output of the initial model, to obtain a spatial geometric information estimation model.


