Voice Recognition Control via Output Volume and Face Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current human-machine interaction methods are limited by single interaction modes, requiring specific gestures or voice commands, leading to unnatural and inconvenient operation.
Innovation Solution
A method and apparatus that dynamically control voice recognition based on output volume, starting and stopping voice recognition functions when the volume is below or above preset thresholds, and using face detection to initiate voice control, ensuring accurate response to user voice operations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If a specific voice word activation method is used (e.g., say 'Hello, Xiao Bing' before dialogue), then the apparatus can recognize voice, but the interaction process is not natural and requires specific gestures to be set in advance
Solution Approach 1:
The system performs preliminary actions by detecting user presence through face detection and estimating user attention through gaze direction analysis before activating voice recognition. This prepares the system in advance to recognize voice commands without requiring users to say specific activation words, making the interaction more natural while maintaining reliable voice recognition.
Solution Approach 2:
The patent replaces the mechanical system of specific gesture-based activation (requiring users to perform predetermined hand movements or say specific words) with an optical and acoustic system that uses face detection, gaze estimation, and voice recognition. This substitution enables more natural interaction by detecting user attention through visual cues rather than requiring explicit mechanical activation gestures.
2Extent of automation
If 'Raise your hand to speak' gesture method is used, then the apparatus can start voice recognition, but certain specific gestures need to be set in advance and the interaction process is not very natural
Solution Approach 1:
The patent replaces the mechanical gesture-based activation system with an optical system using cameras for face detection and gaze estimation. This substitution automatically detects user attention without requiring users to perform specific hand gestures, thereby increasing automation while reducing the complexity of gesture configuration and recognition.
Solution Approach 2:
The system performs self-service by automatically detecting user presence and attention state through face and gaze detection, then autonomously activating or deactivating voice recognition functionality. This eliminates the need for users to learn and perform specific activation gestures, making the system more automated and easier to use.
3Productivity
If voice recognition function is continuously active, then the apparatus can respond to user voice operations, but noise interference increases and user operation accuracy decreases
Solution Approach 1:
The system performs preliminary detection of user presence through face detection and user attention through gaze estimation before activating voice recognition. This preliminary action ensures that voice recognition is only active when the user is actually present and paying attention, thereby improving response speed to legitimate user operations while reducing noise interference from periods when no user is present.
Solution Approach 2:
The system uses feedback from face detection and gaze estimation results to dynamically control the activation state of voice recognition. When the system detects that the user is present and paying attention through visual cues, it activates voice recognition to respond quickly to commands. When the user is not present or not paying attention, the system deactivates voice recognition to reduce noise interference, creating a feedback-controlled adaptive system.
Data Source
AI summary
The present application relates to a human-machine interaction method and device, a computer apparatus, and a storage medium. The method comprises: measuring the current output volume, and if the output volume is less than a first preset threshold, enabling a voice recognition function; acquiring a user's voice message, and measuring the size of the user's voice volume and responding to a user's voice operation; and if the user's voice volume is greater than a second preset threshold, turning down the output volume, and returning to the step of measuring the current output volume. In the entire process, the voice recognition function is controlled to be enabled by means of the output volume of an apparatus itself, thereby accurately responding to the user's voice operation, and if the user's voice is greater than a specified value, turning down the output volume, so that a user's subsequent voice message can be highlighted and accurately acquired so as to bring convenience to a user's operation and implement good human-machine interaction.


