Stereovision Annotation Feedback for 3D-Coherent Training Data
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing machine learning models struggle to accurately identify three-dimensionally coherent locations of points on objects due to the lack of well-prepared training data with accurate labels, particularly in real-world applications where depth information is limited or occlusions occur.
Innovation Solution
A guided feedback loop process using stereo images to ensure three-dimensional coherence of annotations, where a user identifies a location in a first image and adjusts it based on a range of possible locations in a second image, iteratively refining the annotation until it matches the actual location, generating accurate training data for machine learning models.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional single-image annotation techniques are used, then annotation process is simple, but three-dimensional coherence of identified locations is poor
Solution Approach 1:
The patent uses a second image from a different perspective as an intermediary to verify and guide the annotation of the first image. By displaying possible locations in the second image based on the first annotation, the system provides visual feedback that acts as a mediator to ensure three-dimensional coherence without requiring complex manual 3D modeling.
Solution Approach 2:
The system implements a feedback loop where the annotation in the first image generates predicted locations in the second image, which are then displayed to the user. The user can verify these predictions and provide corrected annotations, creating a feedback mechanism that continuously improves three-dimensional coherence while keeping the process relatively simple.
2Measurement precision
If supervised machine learning techniques are used with pixel-level labels, then model training accuracy is improved, but data acquisition difficulty increases significantly
Solution Approach 1:
The system enables annotators to self-correct their annotations by providing visual feedback from the second image. The possible locations displayed in the second image serve as self-service guidance, allowing annotators to verify and adjust their own work without requiring expert supervision or complex validation processes.
Solution Approach 2:
The second image acts as an intermediary that bridges the gap between simple annotation and high accuracy. It provides visual evidence that mediates the annotation process, making it easier to produce pixel-level accurate labels without requiring extremely difficult data collection procedures.
3Measurement precision
If iterative annotation refinement is performed, then annotation accuracy is improved, but time consumption increases
Solution Approach 1:
The system performs partial verification by only displaying possible locations in the second image that are relevant to the first annotation, rather than requiring complete verification of all possible points. This partial action approach maintains accuracy while reducing the time required for iterative refinement.
Solution Approach 2:
The system performs preliminary computation to generate the range of possible locations in the second image before presenting them to the user. This preliminary action prepares the feedback in advance, allowing for faster iterative refinement without requiring real-time complex calculations during the annotation process.
Data Source
AI summary
Certain aspects of the present disclosure provide techniques for generating three-dimensionally coherent training data for image-detection machine learning models. Embodiments include receiving a first image of an object from a first perspective and a second image of the object from a second perspective. Embodiments include receiving user input identifying a location in the first image corresponding to a point on the object. Embodiments include displaying a range of possible locations in the second image corresponding to the point on the object based on the location in the first image. Embodiments include generating training data for a machine learning model based on updated user input associated with the range of possible locations.


