Vision-Based Voice Activation for Smart Displays

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Smart display devices require users to repeatedly use a wake word for each voice command, leading to a cumbersome user experience and potential power consumption issues due to unnecessary activation of voice recognition.

Innovation Solution

Implementing a vision-based mechanism using a camera to detect the presence of a user's face and determine when to activate voice recognition, eliminating the need for a wake word and optimizing power usage by only activating voice recognition when the user is present.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If voice recognition is continuously activated to enable immediate voice command processing, then responsiveness and user experience are improved, but power consumption increases

Engineering Contradiction:
Improvevoice command responsivenessVSAvoidpower consumption
Core Design Contradiction:
Ease of operationVSUse of energy by moving object

Solution Approach 1:

The system performs preliminary detection using the camera to identify user presence before activating voice recognition. This preliminary action (visual detection) prepares the system in advance, allowing voice recognition to be activated only when needed, thus resolving the contradiction between responsiveness and power consumption.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent replaces continuous acoustic monitoring (mechanical voice recognition activation) with optical detection (camera-based user presence detection). This substitution allows the system to determine when voice recognition should be activated based on visual cues rather than continuous audio processing, reducing power consumption while maintaining responsiveness.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Reliability

If wake word is required for each voice command, then false activation is reduced, but user experience becomes cumbersome

Engineering Contradiction:
Improvefalse activation preventionVSAvoidcommand execution simplicity
Core Design Contradiction:
ReliabilityVSEase of operation

Solution Approach 1:

The camera acts as an intermediary between the user and the voice recognition system. By detecting user presence visually, the camera mediates the activation process, allowing the system to distinguish between genuine user intent and background noise, thus preventing false activation while simplifying the interaction process.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system performs preliminary user presence detection before processing voice commands. This preliminary visual verification ensures that voice recognition is activated only when a user is actually present, maintaining reliability while eliminating the need for repetitive wake words and improving operational simplicity.

Inventive Principle:
Principle #10Preliminary action

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

Enhances user experience by allowing seamless voice command execution without the need for wake words and reduces power consumption by intelligently managing voice recognition activation.

Implementation Method 1

A smart display device may include a camera that can capture one or more images of the surroundings of the smart display device

Methodology Applied
Scientific EffectLight reflection and detection: Reflection

Data Source

PatentUS11151993B2Activating voice commands of a smart display device based on a vision-based mechanism
Publication Date: 2021.10.19 BAIDU USA LLC
  • US11151993B2 patent drawing
  • US11151993B2 patent drawing
  • US11151993B2 patent drawing

AI summary

An image is received from a light capture device associated with the smart display device. A determination is made as to whether to activate voice recognition of a recording device associated with the smart display device based on a face being in the image. In response to determining to activate the voice recognition of the recording device associated with the smart display device based on the face being in the image, the voice recognition of the recording device associated with the smart display device is activated.