Road Surface Height Estimation With Temporal Fusion for Autonomous Vehicles
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional methods for determining road surface information, such as those using monocular and stereo cameras, often produce inaccurate or less precise surface height estimations, especially in low-texture or contrast conditions, and lack temporal consistency in predictions.
Innovation Solution
A system utilizing a homography-based road surface estimation guided by stereo neural networks, which includes disparity estimation, track point height determination, and temporal fusion of track point heights to enhance accuracy and robustness, using a plane-parallax algorithm and ego motion-based multi-frame fusion.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If monocular cameras are used for road surface estimation, then device complexity is reduced, but measurement precision of surface height deteriorates
Solution Approach 1:
The patent introduces an intermediary processing system that combines monocular camera data with vehicle motion information (ego motion) and temporal data from multiple frames. This intermediary approach reconstructs 3D surface information by compensating for the monocular camera's depth estimation limitations through mathematical modeling and temporal fusion, achieving accuracy comparable to stereo systems without the added hardware complexity.
2Measurement precision
If stereo systems are implemented for accurate depth estimation, then measurement precision improves, but device complexity increases
Solution Approach 1:
The patent creates a virtual copy of the depth information that would normally require stereo cameras by using temporal fusion of monocular frames. Instead of physically duplicating sensors (stereo pairs), the system synthesizes depth data through temporal processing and geometric reconstruction, effectively copying the functional capability of stereo vision using a single camera across multiple time points.
3Device complexity
If conventional post-processing modules are used for monocular systems, then device complexity is minimized, but reliability of surface predictions deteriorates
Solution Approach 1:
The patent implements continuous temporal fusion by processing multiple consecutive video frames and integrating the surface height information over time. This continuous processing approach maintains reliability by averaging out temporal variations and inconsistencies in individual frames, providing smooth and consistent surface predictions without requiring complex post-processing intervention.
4Measurement precision
If conventional stereo systems are used, then measurement precision is improved, but reliability in low-texture conditions deteriorates
Solution Approach 1:
The patent introduces vehicle motion data and temporal information as intermediary elements that compensate for the lack of texture cues. By combining monocular depth estimation with ego motion compensation and multi-frame temporal fusion, the system creates additional constraints and information sources that replace the texture-dependent matching cues that stereo systems rely on, thereby maintaining reliability in low-texture environments.
Data Source
AI summary
In various examples, systems and methods are disclosed relating to determining first track point heights of a ground surface for each of a plurality of frames of a disparity image based on a plane parallax algorithm, the first track point heights including previous track point heights of the ground surface for each of the at least one previous frame of the plurality of frames of the disparity image and current track point heights of the ground surface for the current frame of the plurality of frames of the disparity image and determining second track point heights by temporally fusing the current track point heights for the current frame and the previous track point heights for each of the at least one previous frame.


