Video-Guided Voice Recognition for Hands-Free Speech Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing voice recognition technologies require separate user gestures to initiate and end voice recording, which can be inconvenient in situations like driving or carrying burdens, and struggle with accurately distinguishing speech from noise in various environments.

Innovation Solution

A method and apparatus that utilize video recognition to detect speech start and end times without requiring user gestures, using a combination of video and audio data analysis to convert the device into a voice recognition mode and generate relevant audio or video data for accurate voice command recognition.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If voice recognition requires separate user gestures to initiate and end recording, then the system can accurately control recording timing, but the ease of operation deteriorates in situations like driving or carrying burdens

Engineering Contradiction:
Improveease of operationVSAvoidreliability
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The system uses the user's own video image to automatically detect speech start and end times through mouth movement analysis, eliminating the need for separate gesture operations. The device serves itself by using its camera to capture the user's face and automatically triggering voice recognition based on detected mouth movements, making the system both easy to operate and reliable.

Inventive Principle:
Principle #25Self-service

2Measurement precision

If the system uses only audio data for voice recognition, then the device complexity is reduced, but the measurement precision of speech start and end times deteriorates in noisy environments

Engineering Contradiction:
Improvemeasurement precisionVSAvoiddevice complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system merges video data from the camera with audio data from the microphone to detect speech start and end times. By combining these two data sources, the system achieves higher measurement precision in determining when speech begins and ends, while the video component helps distinguish actual speech from background noise in challenging acoustic environments.

Inventive Principle:
Principle #5Merging (Combining)

3Loss of information

If the system records all audio data continuously, then no speech data is missed, but the loss of information increases due to including noise and silent sections

Engineering Contradiction:
Improveloss of informationVSAvoidproductivity
Core Design Contradiction:
Loss of informationVSProductivity

Solution Approach 1:

The system performs preliminary analysis of video data to detect mouth movements before processing audio data. By预先 detecting when the user's mouth moves (indicating speech), the system can selectively record only the relevant audio segments, avoiding the inclusion of noise and silent sections while ensuring no actual speech is missed.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS10049665B2Voice recognition method and apparatus using video recognition
Publication Date: 2018.08.14 SAMSUNG ELECTRONICS CO LTD
  • US10049665B2 patent drawing
  • US10049665B2 patent drawing
  • US10049665B2 patent drawing

AI summary

Provided are a method and an apparatus for performing exact start and end recognition of voice based on video recognition. The method includes determining whether a speech starts based on at least one of first video and audio data before conversion into a voice recognition mode, converting into the voice recognition mode and generating second audio data including a voice command, when it is determined that speech starts, and determining whether the speech is terminated based on at least one of second video and audio data after conversion into the voice recognition mode.