Voice Control via Image Pose Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current voice control systems for smart home devices have low interaction efficiency, requiring a wake-up word and relying on complex positioning and alignment processes, which can lead to inaccurate and inefficient user interactions.
Innovation Solution
A method and device for voice control that determines the pose attribute of a target object based on image information, allowing direct execution of voice commands without the need for a wake-up word, thereby improving interaction efficiency and accuracy by analyzing facial expressions, face orientation, and gestures.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If wake-up word and complex positioning alignment processes are used, then voice control reliability is improved, but voice interaction efficiency deteriorates
Solution Approach 1:
The patent extracts and removes the wake-up word requirement and complex positioning alignment processes from the voice control system. By using image recognition to directly identify user identity and intent, the system eliminates these intermediate steps, thereby improving interaction efficiency while maintaining control reliability through accurate image-based user verification.
Solution Approach 2:
The system performs preliminary image acquisition and analysis before voice processing. By pre-identifying the user through image recognition and predicting their intent based on visual data, the system prepares the control context in advance, allowing for more efficient and reliable voice command execution without requiring wake-up words or complex alignment procedures.
2Measurement precision
If wake-up word and complex positioning alignment processes are used, then voice control accuracy is improved, but interaction time increases
Solution Approach 1:
The patent removes the time-consuming wake-up word detection and complex positioning alignment processes from the interaction flow. Image recognition instantly captures user identity and contextual information, enabling the system to process voice commands more quickly while maintaining accuracy through visual verification of user intent.
Solution Approach 2:
The system performs preliminary image analysis to identify user identity and predict intent before voice processing begins. This pre-preparation of contextual information from visual data allows for faster and more accurate voice command execution, reducing overall interaction time while improving control precision.
3Measurement precision
If continuous monitoring and complex positioning are performed, then control precision is improved, but power consumption increases
Solution Approach 1:
The patent extracts and eliminates continuous monitoring and complex positioning operations from the system. Image recognition captures essential user identity and contextual information in a single acquisition, providing sufficient control precision without the need for continuous power-intensive monitoring or complex positioning calculations.
Solution Approach 2:
The system performs partial action by using image recognition to capture only the essential information needed for control (user identity and basic contextual cues) rather than continuously monitoring all possible parameters. This selective approach achieves adequate control precision while significantly reducing power consumption compared to comprehensive continuous monitoring.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
A method for voice control includes that: a voice is acquired to obtain a voice signal; image information is obtained; whether a pose attribute of a target object that utters the voice satisfies a preset condition is determined based on the image information; and responsive to that the pose attribute of the target object satisfies the preset condition, an operation indicated by the voice signal is performed. A terminal and a non-transitory computer-readable storage medium are also provided.