3D Object Visual Perception Through RGB-D Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing computer vision techniques struggle to effectively handle non-rigid 3D objects due to their unpredictable shapes and occlusions, and lack of high-quality training data, making it challenging to develop robust visual perception systems for industrial automation.
Innovation Solution
A system and method utilizing a segmentation machine learning model to segment rigid and non-rigid objects, determine key points, track movements, and generate accurate visual perception by combining point cloud, RGB, and depth data, with modules for filtering and shape detection.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional computer vision techniques are used to handle non-rigid 3D objects, then the system structure remains simple, but the measurement precision and reliability of object identification and tracking deteriorate due to unpredictable shapes and occlusions
Solution Approach 1:
The patent segments the complex task of non-rigid object perception into multiple modules: point cloud segmentation module, RGB-D data processing module, key point detection module, and tracking module. Each module handles a specific aspect of the problem, improving measurement precision while managing system complexity through functional decomposition
Solution Approach 2:
The patent transitions from traditional 2D image processing to 3D point cloud processing by integrating depth information from RGB-D cameras. This dimensional enhancement allows the system to capture spatial characteristics of non-rigid objects more accurately, improving identification precision in complex environments
2Reliability
If traditional computer vision techniques are used, then the system is easier to implement, but the reliability of tracking under occlusion and with infinite configurations deteriorates
Solution Approach 1:
The system performs preliminary segmentation of point cloud data and extraction of spatial characteristics before tracking begins. By pre-processing the data and identifying key features in advance, the system improves tracking reliability even when objects undergo deformation or partial occlusion during operation
Solution Approach 2:
The patent introduces key point detection as an intermediary step between raw point cloud data and tracking algorithms. These detected key points serve as stable reference features that maintain tracking reliability even when the overall object shape changes or becomes partially obscured
3Measurement precision
If diverse training data for non-rigid objects is collected, then the visual perception accuracy improves, but the time and resources required for data acquisition and processing increase
Solution Approach 1:
The system employs unsupervised learning algorithms that automatically learn from the point cloud data without requiring extensive manual annotation or curation of training datasets. This self-service approach reduces data preparation time while maintaining visual perception accuracy by leveraging the inherent structure in the sensor data
Solution Approach 2:
The patent performs preliminary feature extraction and point cloud segmentation before the main perception task. By pre-processing the data to extract meaningful spatial characteristics and organize the point cloud structure in advance, the system reduces the computational burden during actual operation, effectively reducing overall processing time
Data Source
AI summary
A method for determining a visual perception of 3-dimensional (3D) objects in a real scene. The method includes segmenting the 3D objects into segmented data comprising of rigid objects and non-rigid objects. Further, the method includes determining a position and a shape for the segmented 3D objects. The position indicates a set of coordinates, and the shape indicates a sequence of a set of key points. Furthermore, the method includes tracking movement of the segmented 3D objects. Furthermore, the method includes determining the visual perception of the segmented 3D objects based on the tracked movement. The visual perception indicates the shape and location of the rigid objects and the non-rigid objects in the real scene.


