Multi-Angle Object Recognition With Guided Camera Feedback

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing object recognition software on mobile devices lacks real-time user feedback, leading to imperfect object recognition, including incorrect identification, no positive identification, or identification of undesired objects, due to the absence of direct visual indicators of the object recognition process.

Innovation Solution

The implementation of computer-implemented methods that generate user interface elements to indicate camera operations, such as capturing multiple images from different angles or zoom levels, allowing users to provide feedback and assist in the object recognition process, and sending images to an object recognition server for analysis.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If object recognition software operates automatically without user feedback, then the operation simplicity is improved, but the object recognition accuracy deteriorates

Engineering Contradiction:
Improveoperation simplicityVSAvoidobject recognition accuracy
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The patent implements real-time visual feedback by displaying candidate object identifications and confidence scores to users during the recognition process. This allows users to provide implicit feedback through their viewing behavior and explicit feedback through selections, which the system uses to iteratively improve recognition accuracy while maintaining automatic operation.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent introduces an intermediary interface layer between the automatic recognition system and the final identification result. This interface presents candidate objects with confidence metrics, allowing the system to maintain automatic operation while enabling user verification and correction when needed, thus resolving the contradiction between automation and accuracy.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If multiple images from different angles are required for object recognition, then the object recognition accuracy is improved, but the time required for recognition increases

Engineering Contradiction:
Improveobject recognition accuracyVSAvoidrecognition time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent implements partial action by processing and presenting candidate identification results after analyzing a subset of captured images, rather than requiring all images to be processed. This allows the system to achieve satisfactory recognition accuracy within acceptable time limits by performing recognition on representative samples of the multi-angle images.

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The patent applies preliminary action by pre-processing and pre-evaluating multiple captured images to identify the most informative ones for recognition. This preliminary selection and preparation of images before the actual recognition process reduces the time required for full analysis while maintaining accuracy by focusing computational resources on the most valuable images.

Inventive Principle:
Principle #10Preliminary action

3Loss of information

If real-time visual indicators are displayed to show the recognition process, then the user understanding is improved, but the device complexity increases

Engineering Contradiction:
Improveuser understandingVSAvoidinterface complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent segments the complex recognition process into distinct visual components: candidate object displays, confidence score indicators, and process status markers. This segmentation presents comprehensive recognition information to users without overwhelming them, as each element is displayed separately and clearly, reducing perceived complexity while maintaining full information transparency.

Inventive Principle:
Principle #1Segmentation

4Loss of information

If the system processes and displays all candidate objects, then the completeness of identification is improved, but the information overload increases

Engineering Contradiction:
Improveidentification completenessVSAvoidinformation overload
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent applies local quality by differentiating the visual presentation of candidate objects based on their confidence scores and relevance. High-confidence primary candidates are displayed with prominent formatting, while lower-confidence alternatives receive less prominent display. This selective emphasis maintains complete information availability while preventing information overload through hierarchical presentation.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS12347154B2Multi-angle object recognition
Publication Date: 2025.07.01 GOOGLE LLC
  • US12347154B2 patent drawing
  • US12347154B2 patent drawing
  • US12347154B2 patent drawing

AI summary

Methods, systems, and apparatus for controlling smart devices are described. In one aspect a method includes capturing, by a camera on a user device, a plurality of successive images for display in an application environment of an application executing on the user device, performing an object recognition process on the images, the object recognition process including determining that a plurality of images, each depicting a particular object, are required to perform object recognition on the particular object, and in response to the determination, generating a user interface element that indicates a camera operation to be performed, the camera option capturing two or more images, determining that a user, in response to the user interface element, has caused the indicated camera operation to be performed to capture the two or more images, and in response, determining whether a particular object is positively identified from the plurality of images.