Multi-Angle Object Recognition With Guided Camera Feedback
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing object recognition software on mobile devices lacks real-time user feedback, leading to imperfect object recognition, including incorrect identification, no positive identification, or identification of undesired objects, due to the absence of direct visual indicators of the object recognition process.
Innovation Solution
The implementation of computer-implemented methods that generate user interface elements to indicate camera operations, such as capturing multiple images from different angles or zoom levels, allowing users to provide feedback and assist in the object recognition process, and sending images to an object recognition server for analysis.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If object recognition software operates automatically without user feedback, then the operation simplicity is improved, but the object recognition accuracy deteriorates
Solution Approach 1:
The patent implements real-time visual feedback by displaying candidate object identifications and confidence scores to users during the recognition process. This allows users to provide implicit feedback through their viewing behavior and explicit feedback through selections, which the system uses to iteratively improve recognition accuracy while maintaining automatic operation.
Solution Approach 2:
The patent introduces an intermediary interface layer between the automatic recognition system and the final identification result. This interface presents candidate objects with confidence metrics, allowing the system to maintain automatic operation while enabling user verification and correction when needed, thus resolving the contradiction between automation and accuracy.
2Measurement precision
If multiple images from different angles are required for object recognition, then the object recognition accuracy is improved, but the time required for recognition increases
Solution Approach 1:
The patent implements partial action by processing and presenting candidate identification results after analyzing a subset of captured images, rather than requiring all images to be processed. This allows the system to achieve satisfactory recognition accuracy within acceptable time limits by performing recognition on representative samples of the multi-angle images.
Solution Approach 2:
The patent applies preliminary action by pre-processing and pre-evaluating multiple captured images to identify the most informative ones for recognition. This preliminary selection and preparation of images before the actual recognition process reduces the time required for full analysis while maintaining accuracy by focusing computational resources on the most valuable images.
3Loss of information
If real-time visual indicators are displayed to show the recognition process, then the user understanding is improved, but the device complexity increases
Solution Approach 1:
The patent segments the complex recognition process into distinct visual components: candidate object displays, confidence score indicators, and process status markers. This segmentation presents comprehensive recognition information to users without overwhelming them, as each element is displayed separately and clearly, reducing perceived complexity while maintaining full information transparency.
4Loss of information
If the system processes and displays all candidate objects, then the completeness of identification is improved, but the information overload increases
Solution Approach 1:
The patent applies local quality by differentiating the visual presentation of candidate objects based on their confidence scores and relevance. High-confidence primary candidates are displayed with prominent formatting, while lower-confidence alternatives receive less prominent display. This selective emphasis maintains complete information availability while preventing information overload through hierarchical presentation.
Data Source
AI summary
Methods, systems, and apparatus for controlling smart devices are described. In one aspect a method includes capturing, by a camera on a user device, a plurality of successive images for display in an application environment of an application executing on the user device, performing an object recognition process on the images, the object recognition process including determining that a plurality of images, each depicting a particular object, are required to perform object recognition on the particular object, and in response to the determination, generating a user interface element that indicates a camera operation to be performed, the camera option capturing two or more images, determining that a user, in response to the user interface element, has caused the indicated camera operation to be performed to capture the two or more images, and in response, determining whether a particular object is positively identified from the plurality of images.


