Automatic Image Annotation Using Environmental Feature Points
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current image annotation methods in AR/VR applications are inefficient, labor-intensive, and lack universal applicability, especially when CAD models are not available or when objects are solid-colored, highly reflective, or transparent, leading to inaccurate automatic annotation.
Innovation Solution
A method for automatically annotating objects in images using a computer-implemented process that generates a three-dimensional space model based on environmental feature points, allowing for accurate annotation of objects without relying on CAD models, even for complex or difficult-to-track objects.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If manual annotation is used to annotate target objects in images, then annotation flexibility and adaptability to different object types are maintained, but annotation efficiency and productivity are severely reduced
Solution Approach 1:
The system performs automatic annotation by having the target object serve itself through its own visual features. The annotation algorithm automatically detects, tracks, and annotates target objects without requiring manual intervention for each image, enabling the system to self-service the annotation process while maintaining adaptability to different object types through feature-based detection
Solution Approach 2:
The patent replaces the mechanical manual annotation process with an automated computer vision system. Instead of manual delineation of rectangular boxes, the system uses algorithmic detection and tracking based on visual features, substituting human mechanical action with automated image processing and pattern recognition
2Measurement precision
If CAD model based automatic annotation is used, then annotation precision and automation are improved, but universal applicability deteriorates when CAD models are unavailable
Solution Approach 1:
The patent extracts the essential visual features directly from the target objects in the images themselves, rather than relying on external CAD models. By taking out the object detection and tracking functionality and making it independent of pre-existing models, the system achieves universal applicability while maintaining precision through feature-based identification
Solution Approach 2:
Instead of starting with a CAD model and matching it to images (top-down approach), the system inverts the approach by detecting objects directly from images and creating annotations bottom-up. This inversion removes the dependency on CAD models while maintaining annotation precision through direct visual feature analysis
3Extent of automation
If model-based tracking is used for automatic annotation, then annotation automation is achieved, but tracking accuracy deteriorates for solid-color, highly reflective, or transparent objects
Solution Approach 1:
The system applies local quality by using different detection strategies for different types of objects. Instead of a uniform tracking approach, it adapts to local characteristics of target objects, using feature detection methods that work specifically for solid-color objects, reflective objects, or transparent objects based on their local visual properties
Solution Approach 2:
The patent changes detection parameters and algorithms based on object characteristics. By adjusting detection sensitivity, feature types, and tracking parameters according to the specific properties of target objects (solid-color, reflective, transparent), the system maintains high tracking accuracy across diverse object types while preserving automation
Data Source
Figure 1-1
Figure 1-2
Figure 2
AI summary
Embodiments of the disclosure disclose a method for automatically annotating a target object in images. In one embodiment, the method comprises: obtaining an image training sample including a plurality of images, wherein each image of the plurality of images is obtained by photographing a same target object, and the adjacent images share one or more same environmental feature points; using one of the plurality of images as a reference image to determine a reference coordinate system, and create a three-dimensional space model based on the three-dimensional reference coordinate system; determining the position information of the target object in the three-dimensional reference coordinate system upon the three-dimensional space model being moved to the position of the target object in the reference image; and mapping the three-dimensional space model to image planes of each image, respectively, based on respective camera pose information determined based on environmental feature points in each image.