Tap-less Visual Search with On-Device Object Tracking
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current visual search apps require users to manually tap on a screen to capture images, which can be cumbersome and interrupt the search process, especially in crowded scenes, necessitating a more intuitive and efficient method for product matching.
Innovation Solution
A 'tap-less' visual search system that uses on-device object detection and tracking to automatically identify and focus on products of interest, providing real-time visual cues and eliminating the need for user interaction, with data processed in the cloud for product matching.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If a manual tap-to-search approach is used, then the system can capture product images, but the user interaction becomes cumbersome and interrupts the search process
Solution Approach 1:
The system performs preliminary object detection and tracking to identify the product of interest before the user completes the search action. By continuously monitoring the camera feed and detecting product characteristics in advance, the system prepares the search query ready to be executed automatically when the product is identified, eliminating the need for manual tap confirmation and maintaining search process continuity
Solution Approach 2:
The system enables self-service by automatically capturing the product image and initiating the search process without requiring user confirmation. The object detection algorithm autonomously identifies the product, extracts relevant features, and triggers the search query automatically, allowing the system to serve itself and eliminating cumbersome user interactions while maintaining continuous search flow
2Extent of automation
If automatic object detection is implemented, then user interaction is minimized, but the system complexity increases
Solution Approach 1:
The patent introduces an intermediary object detection module that acts as a mediator between the camera input and the search engine. This intermediate layer continuously analyzes the camera feed, detects product characteristics, and filters relevant information before passing it to the search system. By placing this intermediary component, the system achieves automatic product identification while managing complexity through modular architecture, where the detection module handles the complex computer vision tasks separately from the core search functionality
3Measurement precision
If visual cues are provided to guide product focus, then product selection accuracy improves, but the interface complexity increases
Solution Approach 1:
The system employs color changes and visual transformations as cues to guide product focus. By applying color overlays, highlights, or visual effects to detected products in the camera view, the system clearly indicates which objects have been identified and are ready for search. These visual cues improve product selection accuracy by making the detected items stand out, while maintaining a relatively simple interface by using intuitive visual feedback rather than complex controls or multiple interface elements
Data Source
AI summary
A visual search uses information for identifying a product of interest to determine if the product of interest is present within a video being provided to the visual search engine. When the product of interest is determined by the visual search engine to be present within the video, a computing device in communication with the visual search engine is caused to provide a notification that the product of interest has been detected within the video.


