Voice Recognition Control via Output Volume and Face Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current human-machine interaction methods are limited by single interaction modes, requiring specific gestures or voice commands, leading to unnatural and inconvenient operation.

Innovation Solution

A method and apparatus that dynamically control voice recognition based on output volume, starting and stopping voice recognition functions when the volume is below or above preset thresholds, and using face detection to initiate voice control, ensuring accurate response to user voice operations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If a specific voice word activation method is used (e.g., say 'Hello, Xiao Bing' before dialogue), then the apparatus can recognize voice, but the interaction process is not natural and requires specific gestures to be set in advance

Engineering Contradiction:
Improvevoice recognition accuracyVSAvoidinteraction naturalness
Core Design Contradiction:
ReliabilityVSEase of operation

Solution Approach 1:

The system performs preliminary actions by detecting user presence through face detection and estimating user attention through gaze direction analysis before activating voice recognition. This prepares the system in advance to recognize voice commands without requiring users to say specific activation words, making the interaction more natural while maintaining reliable voice recognition.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent replaces the mechanical system of specific gesture-based activation (requiring users to perform predetermined hand movements or say specific words) with an optical and acoustic system that uses face detection, gaze estimation, and voice recognition. This substitution enables more natural interaction by detecting user attention through visual cues rather than requiring explicit mechanical activation gestures.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Extent of automation

If 'Raise your hand to speak' gesture method is used, then the apparatus can start voice recognition, but certain specific gestures need to be set in advance and the interaction process is not very natural

Engineering Contradiction:
Improveautomatic voice recognition activationVSAvoidgesture setting complexity
Core Design Contradiction:
Extent of automationVSDevice complexity

Solution Approach 1:

The patent replaces the mechanical gesture-based activation system with an optical system using cameras for face detection and gaze estimation. This substitution automatically detects user attention without requiring users to perform specific hand gestures, thereby increasing automation while reducing the complexity of gesture configuration and recognition.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system performs self-service by automatically detecting user presence and attention state through face and gaze detection, then autonomously activating or deactivating voice recognition functionality. This eliminates the need for users to learn and perform specific activation gestures, making the system more automated and easier to use.

Inventive Principle:
Principle #25Self-service

3Productivity

If voice recognition function is continuously active, then the apparatus can respond to user voice operations, but noise interference increases and user operation accuracy decreases

Engineering Contradiction:
Improvevoice operation response speedVSAvoidnoise interference
Core Design Contradiction:
ProductivityVSObject-affected harmful factors

Solution Approach 1:

The system performs preliminary detection of user presence through face detection and user attention through gaze estimation before activating voice recognition. This preliminary action ensures that voice recognition is only active when the user is actually present and paying attention, thereby improving response speed to legitimate user operations while reducing noise interference from periods when no user is present.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system uses feedback from face detection and gaze estimation results to dynamically control the activation state of voice recognition. When the system detects that the user is present and paying attention through visual cues, it activates voice recognition to respond quickly to commands. When the user is not present or not paying attention, the system deactivates voice recognition to reduce noise interference, creating a feedback-controlled adaptive system.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS11483657B2Human-machine interaction method and device, computer apparatus, and storage medium
Publication Date: 2022.10.25 LIU GUOHUA
  • US11483657B2 patent drawing
  • US11483657B2 patent drawing
  • US11483657B2 patent drawing

AI summary

The present application relates to a human-machine interaction method and device, a computer apparatus, and a storage medium. The method comprises: measuring the current output volume, and if the output volume is less than a first preset threshold, enabling a voice recognition function; acquiring a user's voice message, and measuring the size of the user's voice volume and responding to a user's voice operation; and if the user's voice volume is greater than a second preset threshold, turning down the output volume, and returning to the step of measuring the current output volume. In the entire process, the voice recognition function is controlled to be enabled by means of the output volume of an apparatus itself, thereby accurately responding to the user's voice operation, and if the user's voice is greater than a specified value, turning down the output volume, so that a user's subsequent voice message can be highlighted and accurately acquired so as to bring convenience to a user's operation and implement good human-machine interaction.