Facial Feature Detection for Speech Activity Start and End

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current voice-controlled applications face challenges in determining the start and end of speech in noisy environments, leading to unintended noise processing and unnatural user interactions, particularly in environments with background noise or multiple speakers.

Innovation Solution

Facial feature detection is used to determine the user's intention to interact with a device by detecting the presence of a face looking at the screen and mouth movement, allowing for hands-free operation and natural command inputs without the need for trigger words or button presses.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual button press is used to control speech start and stop, then speech processing accuracy is improved, but ease of operation deteriorates

Engineering Contradiction:
Improvespeech processing accuracyVSAvoidease of operation
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The patent replaces the mechanical button-pressing system with an automated facial feature detection system. The camera captures facial images, and software algorithms detect mouth movement patterns to automatically determine speech start and end points, eliminating the need for manual mechanical input while maintaining accurate speech processing.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system enables self-service operation by automatically detecting when the user is speaking through facial feature analysis. The device monitors its own operational state using the camera and algorithms, eliminating the need for external manual control while maintaining precise speech processing.

Inventive Principle:
Principle #25Self-service

2Measurement precision

If trigger words are used to signal speech commands, then speech start detection is improved, but naturalness of interaction deteriorates

Engineering Contradiction:
Improvespeech start detectionVSAvoidnaturalness of interaction
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The patent replaces the linguistic trigger-word system with a visual facial detection system. Instead of requiring users to speak specific trigger words, the camera-based system directly detects mouth movement patterns that indicate speech initiation, enabling more natural interactions without artificial linguistic constraints.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Device complexity

If audio-only processing is used, then device complexity is reduced, but reliability deteriorates in noisy environments

Engineering Contradiction:
Improvedevice complexityVSAvoidreliability
Core Design Contradiction:
Device complexityVSReliability

Solution Approach 1:

The patent merges audio processing with visual facial feature detection. The system combines data from both the microphone (audio signals) and camera (facial images), using facial mouth movement detection to validate and enhance speech detection reliability, particularly in noisy environments where audio alone may be insufficient.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS10930303B2System and method for enhancing speech activity detection using facial feature detection
Publication Date: 2021.02.23 MICROSOFT TECHNOLOGY LICENSING LLC
  • US10930303B2 patent drawing
  • US10930303B2 patent drawing
  • US10930303B2 patent drawing

AI summary

Disclosed herein are systems, methods, and non-transitory computer-readable storage media for processing audio. A system configured to practice the method monitors, via a processor of a computing device, an image feed of a user interacting with the computing device and identifies an audio start event in the image feed based on face detection of the user looking at the computing device or a specific region of the computing device. The image feed can be a video stream. The audio start event can be based on a head size, orientation or distance from the computing device, eye position or direction, device orientation, mouth movement, and/or other user features. Then the system initiates processing of a received audio signal based on the audio start event. The system can also identify an audio end event in the image feed and end processing of the received audio signal based on the end event.