Object Position Estimation via Fusion of Visual and Relationship Data
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing techniques fail to accurately identify the position of objects in images, especially when objects are partially hidden, due to insufficient training data and high labor costs in preparing object images and training information.
Innovation Solution
A position estimation device and method that fuse position information, visual information, and relationship information of a subject object with a target object to estimate the target object's position using an object position estimator, with a parameter update unit optimizing the estimation by reducing the distance between estimated and correct position information.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional object detection methods are used, then training data preparation is required, but labor costs and time consumption increase significantly
Solution Approach 1:
The system performs preliminary actions by pre-extracting visual features from images and pre-computing relationships between objects before actual detection. The visual feature extractor learns to represent objects, and the relationship extractor pre-establishes spatial and semantic relationships, so that during actual object detection, only new information needs to be processed rather than preparing all training data from scratch.
Solution Approach 2:
The patent introduces intermediary components including a visual feature extractor that mediates between raw images and object detection, and a relationship extractor that mediates between object pairs. These intermediaries process and transform information in a standardized way, reducing the need for manual training data preparation while maintaining detection accuracy.
2Reliability
If objects are partially hidden in images, then detection becomes more difficult, but conventional methods still fail to identify position
Solution Approach 1:
The system employs feedback mechanisms where the relationship between detected subject objects and target objects is continuously refined. The relationship extractor analyzes spatial and semantic relationships between object pairs, and this information feeds back to improve the detection of hidden objects by considering contextual relationships rather than relying solely on direct visual detection of the hidden object itself.
Solution Approach 2:
The patent transitions from two-dimensional image analysis to three-dimensional spatial relationship analysis by extracting and utilizing relationship information between objects. This adds a dimensional layer of understanding that helps identify hidden objects through their relationships with visible objects, effectively solving the detection problem for partially hidden objects.
3Measurement precision
If extensive training data is collected, then model accuracy improves, but preparation labor and cost increase
Solution Approach 1:
The system performs self-service by automatically extracting visual features and relationships from images without requiring manual annotation or preparation of training data. The visual feature extractor and relationship extractor automatically process image data, generating the necessary training information themselves, which eliminates the need for human labor in data preparation while maintaining model accuracy.
Solution Approach 2:
The patent changes the parameters of data processing by transforming raw image data into extracted visual features and relationship information. This parameter transformation allows the system to work with automatically generated data representations rather than requiring manually prepared training data, significantly reducing preparation labor while maintaining the accuracy needed for position estimation.
Data Source
AI summary
It is possible to identify a position of a target object that is difficult to recognize.A position estimation device includes: an information fusion unit that generates fusion information in which position information of a subject object that is an object corresponding to a subject, visual information of the subject object, and relationship information indicating a relationship with a target object paired with the subject object are fused; and an object position estimation unit that estimates a position of the target object by using an object position estimator learned in advance on the basis of the fusion information.


