Mobile Image Search Using Multi-Attribute Detectors
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current image recognition systems fall short of replicating the human ability to instantly retrieve information about objects, landscapes, and text from images, often requiring specific viewing angles and limited application in real-world scenarios due to limitations in pose invariance and contextual understanding.
Innovation Solution
A system that utilizes attribute detectors trained on example images to convert image information into symbolic data, allowing users to capture images which are then processed by a server to retrieve relevant information from databases, incorporating GPS and multi-angle image training for improved recognition and integration with search engines.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If traditional image recognition systems are used, then object recognition is achieved, but the system fails to replicate human visual memory capabilities and requires specific viewing angles
Solution Approach 1:
The patent segments the image recognition task into multiple attribute detectors that independently analyze different visual attributes (color, shape, texture, etc.). Each detector is trained on multiple views and angles, allowing the system to recognize objects regardless of viewing angle by aggregating results from multiple specialized detectors.
Solution Approach 2:
The patent transitions from traditional single-view image recognition to multi-view attribute-based recognition. By training attribute detectors on images from multiple angles and dimensions, the system creates a more comprehensive representation that replicates human visual memory's ability to recognize objects from any viewpoint.
2Measurement precision
If image recognition is performed with limited training data, then processing speed is maintained, but recognition accuracy and contextual understanding deteriorate
Solution Approach 1:
The patent implements attribute detectors that focus on specific visual attributes rather than attempting comprehensive object recognition. Each detector is trained on a focused set of attributes with multiple views, achieving high accuracy for that attribute without requiring exhaustive training data for all possible object characteristics.
Solution Approach 2:
The system changes the parameter space by breaking down complex object recognition into multiple attribute dimensions (color, shape, texture, etc.). Each attribute detector operates in its own parameter space with targeted training data, improving overall recognition accuracy while managing training complexity through specialization.
3Loss of information
If comprehensive image analysis is performed, then information retrieval capability is improved, but processing time and system complexity increase
Solution Approach 1:
The patent segments the comprehensive image analysis into parallel attribute detection tasks. Multiple attribute detectors operate simultaneously on different visual attributes, enabling complete information retrieval without sequential processing delays. The segmented approach maintains completeness while reducing overall processing time through parallelization.
Solution Approach 2:
The attribute detectors are pre-trained on extensive multi-view training data before deployment. This preliminary action allows the detectors to quickly match new images against learned attributes without requiring complex real-time analysis, reducing processing time while maintaining comprehensive information retrieval capability.
Data Source
AI summary
An increasing number of mobile telephones and computers are being equipped with a camera. Thus, instead of simple text strings, it is also possible to send images as queries to search engines or databases. Moreover, advances in image recognition allow a greater degree of automated recognition of objects, strings of letters, or symbols in digital images. This makes it possible to convert the graphical information into a symbolic format, for example, plain text, in order to then access information about the object shown.


