3D Shape Data Generation Using Image Grid Recognition
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current technologies for recognizing 3D objects in autonomous driving and other industries face challenges such as interference, high construction costs, and inaccuracies in distance measurement, particularly in measuring the size and shape of objects, especially the invisible rear side, and require large-capacity 3D data for accurate object recognition.
Innovation Solution
A method and apparatus using image recognition to generate 3D entity shape data by recognizing a grid matching part in images, creating a cube-shaped 3D space grid, and generating shape data for external objects, which reduces construction costs and improves accuracy by obtaining real coordinates and distances of feature points, allowing for the representation of complete 3D information.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If radar or laser is used to measure 3D objects, then distance measurement capability is improved, but interference and crosstalk occur as the number of vehicles increases
Solution Approach 1:
The patent replaces radar and laser (electromagnetic wave-based active sensing systems) with a camera-based passive imaging system combined with image recognition algorithms. This substitution eliminates the interference and crosstalk problems inherent in active sensing systems while maintaining the ability to measure 3D shape information through 2D image analysis and geometric transformation.
2Measurement precision
If 3D cameras are used to capture 3D information, then depth perception is improved, but mechanical wear increases and focal length adjustment is limited
Solution Approach 1:
The patent replaces complex mechanical 3D camera systems with a simpler 2D camera system combined with computational image processing. By using image recognition to identify grid patterns and applying geometric transformation algorithms, the system derives 3D shape information without requiring mechanical focal length adjustment or complex lens mechanisms, thereby eliminating mechanical wear issues.
3Measurement precision
If road measurement with tape measure and checkerboard painting is used, then distance information accuracy is improved, but construction cost increases excessively
Solution Approach 1:
The patent extracts and utilizes naturally existing grid-like structures in the environment (such as road markings, building facades, or artificially placed simple grid markers) rather than requiring complete checkerboard painting of entire roads. This selective extraction of grid features from existing structures significantly reduces construction costs while maintaining the ability to calculate 3D shape information through image recognition.
Solution Approach 2:
The patent employs simple, inexpensive grid markers or utilizes existing road markings instead of expensive permanent checkerboard road surfaces. These minimal grid indicators can be easily applied and removed, providing sufficient reference information for 3D reconstruction without the excessive construction investment required by traditional checkerboard methods.
4Reliability
If multiple 2D images are combined to represent 3D objects, then object recognition capability is improved, but data capacity increases significantly
Solution Approach 1:
The patent segments the 3D object representation into discrete grid units (voxels or polyhedral elements) that can be independently processed and stored. By dividing the continuous 3D space into discrete grid cells and only storing information for occupied cells, the system reduces overall data capacity requirements while maintaining complete object representation for accurate recognition.
Solution Approach 2:
The patent transforms multiple 2D image data into a compressed 3D grid representation, effectively using the third dimension to organize and compress spatial information. This dimensional transformation allows the system to store 3D shape data more efficiently than storing multiple full-resolution 2D images, as the grid structure inherently compresses redundant information across different viewing angles.
Data Source
AI summary
A method for generating 3D entity shape data using image recognition which is performed by a computing device, the method includes the steps of: recognizing a grid matching part having four edge vertices of a quadrangle displayed in an image captured by a camera; generating a cube-shaped 3D space grid of a specific distance unit applied to the image by using the grid matching part; and generating shape data for an external object using the 3D space grid.


