AR Visual Search Using Multi-Model Recognition for Pose Variation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing augmented reality visual search technologies are lacking in their ability to effectively recognize objects or persons within the wearer's field of view, particularly due to challenges in varying lighting conditions, poses/angles, and expressions, and the need for improved machine learning models for accurate facial and object recognition.
Innovation Solution
An apparatus and method utilizing augmented reality glasses with a network-connected machine learning model that includes a data comparator and search engine, capable of comparing live video data to a comparative database for enhanced recognition, incorporating 2D and 3D analyses, and employing multi-biometrics and pose/position estimation algorithms to improve accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If machine learning models are used for facial and object recognition in augmented reality, then recognition capability is provided, but accuracy is reduced under varying lighting conditions, poses/angles, and expressions
Solution Approach 1:
The patent segments the recognition task into multiple specialized components: a first ML model for face detection and landmark identification, a second ML model for pose estimation, and a third ML model for expression recognition. This segmentation allows each model to specialize in specific conditions and variations, improving overall accuracy under diverse lighting, poses, and expressions while maintaining reliable recognition capability.
2Ease of operation
If conventional visual search technology is used in augmented reality glasses, then basic object detection is provided, but the ability to search for recognizable objects or persons within the wearer's field of view is lacking
Solution Approach 1:
The patent implements a universal visual search system that can detect and recognize multiple types of targets including faces, objects, and persons within the augmented reality field of view. The system integrates face detection, pose estimation, expression recognition, and object search capabilities into a single platform, enabling the glasses to perform diverse visual search tasks across different scenarios and conditions.
3Speed
If single-model machine learning approaches are used, then processing speed is maintained, but accuracy in complex recognition scenarios is reduced
Solution Approach 1:
The patent divides the recognition system into multiple specialized ML models that operate in parallel: face detection model, pose estimation model, and expression recognition model. Each model processes specific aspects of the input data independently, then their results are integrated to produce the final recognition output. This segmented approach maintains processing speed while improving accuracy through specialized processing of different recognition tasks.
Data Source
AI summary
An apparatus, system and method for providing a visual search using augmented reality glasses. The apparatus, system and method include a network communicatively associated with the glasses capable of providing remote connectivity to an application programming interface (API); a machine learning (ML) model communicative with the API and having an input capable of receiving live video data indicative of a view field of the glasses, wherein the ML model includes at least a data comparator and platform-specific coding corresponded to the glasses; a search engine within the ML model and having a secondary input interfaced to a comparative database, wherein the search engine compares the live view field video data to the secondary input using the comparator; and a match output capable of outputting a match obtained by the search engine over the network to the glasses.


