3D Object Detection via Coordinate Conversion for Monocular Depth Recovery
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Monocular imaging devices lack depth information, leading to low accuracy in obtaining camera coordinates of objects, which is a limitation in vehicle-road coordination monitoring.
Innovation Solution
A 3D object detection method that determines first plane coordinates of landing points of an object in a projection frame, establishes world coordinate information based on object size and position, and uses coordinate conversion processing to obtain camera coordinates using external and internal parameter information of the image acquisition device.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If a single monocular imaging device is used, then the cost is low and resolution is high, but the accuracy of obtaining camera coordinates is low due to lack of depth information
Solution Approach 1:
The patent introduces an intermediary coordinate conversion process that maps 2D image coordinates to 3D world coordinates through a projection frame. By establishing correspondence between landing points in the image and their 3D world coordinates, the system recovers depth information indirectly without requiring additional depth sensors, thus resolving the contradiction between using simple monocular imaging and achieving accurate 3D coordinates.
2Measurement precision
If coordinate conversion processing is performed to obtain camera coordinates, then measurement accuracy is improved, but device complexity increases
Solution Approach 1:
The patent performs preliminary actions by pre-establishing the projection frame and corresponding 3D world coordinate system before actual measurement. By preparing the coordinate transformation relationships in advance and storing the correspondence between 2D image points and 3D world points, the system reduces the complexity of real-time coordinate conversion while maintaining high measurement accuracy.
Data Source
AI summary
A three-dimensional object detection method includes: determining first plane coordinates of multiple landing points of a first object based on a projection frame of the 3D first object in a to-be-detected image, the to-be-detected image being captured by an image acquisition device; obtaining first world coordinate information based on a world coordinate system established based on position information and size information of the first object, the first world coordinate information including first world coordinates of the multiple landing points of the first object and first word coordinates of multiple vertices of the first object; and using coordinate conversion processing to obtain external parameter information that converts the world coordinates of the first object into camera coordinate based on the first plane coordinates, the first world coordinates of the multiple landing points of the first object, and the internal parameter information of the image acquisition device.


