Visual Cue Detection False Positive Reduction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Vision sensor-equipped assistant devices face challenges with false positives due to visual content from sources like televisions and computer screens being mistaken for user-provided visual cues, leading to unintended actions and resource wastage.
Innovation Solution
Implementing excluded region classification techniques, such as using machine learning models and object recognition methods, to identify and ignore or downweight regions likely to contain visual noise, thereby reducing false positives by classifying areas with televisions, computer screens, and other potential noise sources.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If vision sensors analyze all regions in the field of view to detect visual cues, then detection completeness is improved, but false positive rate increases due to visual noise from televisions and computer screens
Solution Approach 1:
The patent divides the field of view into multiple regions of interest (ROIs), each corresponding to a potential visual noise source such as televisions or computer screens. By segmenting the analysis area, the system can selectively apply different processing strategies to different regions, analyzing user-facing areas thoroughly while excluding or downweighting regions containing visual noise sources.
Solution Approach 2:
The patent extracts and removes regions containing visual noise sources from the set of regions subject to full visual cue analysis. By identifying televisions, computer screens, and other potential noise sources and excluding them from analysis, the system eliminates the source of false positives while maintaining comprehensive analysis of remaining areas.
2Ease of operation
If vision sensors continuously monitor the entire field of view, then user interaction detection is improved, but computing resource consumption increases
Solution Approach 1:
The patent segments the field of view into regions of interest and non-interest areas. By identifying and excluding regions containing visual noise sources from continuous monitoring, the system reduces the computational load while maintaining effective user interaction detection in the remaining regions.
Solution Approach 2:
The patent applies partial action by monitoring only the necessary portions of the field of view rather than the entire area. By excluding regions with visual noise sources from full analysis, the system performs sufficient monitoring for user interaction detection while consuming fewer computing resources.
3Area of stationary object
If visual cues from all detected regions are processed equally, then detection coverage is improved, but false positive actions increase due to noise sources
Solution Approach 1:
The patent segments the field of view into multiple regions, identifying those containing visual noise sources such as televisions and computer screens. By applying different processing weights to different segments, the system maintains broad detection coverage while reducing the influence of regions that generate false positives.
Solution Approach 2:
The patent applies local quality by assigning different processing weights to different regions of the field of view. Regions containing visual noise sources are assigned lower weights or excluded from analysis, while user-facing regions receive full processing attention. This localized differentiation maintains coverage while eliminating harmful false positives.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
Techniques are described herein for reducing false positives in vision sensor-equipped assistant devices. In various implementations, initial image frame(s) may be obtained from vision sensor(s) of an assistant device and analyzed to classify a particular region of the initial image frames as being likely to contain visual noise. Subsequent image frame(s) obtained from the vision sensor(s) may then be analyzed to detect actionable user-provided visual cue(s), in a manner that reduces or eliminates false positives. In some implementations, no analysis may be performed on the particular region of the subsequent image frame(s). Additionally or alternatively, in some implementations, a first candidate visual cue detected within the particular region may be weighted less heavily than a second candidate visual cue detected elsewhere in the one or more subsequent image frames. An automated assistant may then take responsive action based on the detected actionable visual cue(s).