Voice Control via Image Pose Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current voice control systems for smart home devices have low interaction efficiency, requiring a wake-up word and relying on complex positioning and alignment processes, which can lead to inaccurate and inefficient user interactions.

Innovation Solution

A method and device for voice control that determines the pose attribute of a target object based on image information, allowing direct execution of voice commands without the need for a wake-up word, thereby improving interaction efficiency and accuracy by analyzing facial expressions, face orientation, and gestures.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If wake-up word and complex positioning alignment processes are used, then voice control reliability is improved, but voice interaction efficiency deteriorates

Engineering Contradiction:
Improvevoice control reliabilityVSAvoidvoice interaction efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent extracts and removes the wake-up word requirement and complex positioning alignment processes from the voice control system. By using image recognition to directly identify user identity and intent, the system eliminates these intermediate steps, thereby improving interaction efficiency while maintaining control reliability through accurate image-based user verification.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system performs preliminary image acquisition and analysis before voice processing. By pre-identifying the user through image recognition and predicting their intent based on visual data, the system prepares the control context in advance, allowing for more efficient and reliable voice command execution without requiring wake-up words or complex alignment procedures.

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If wake-up word and complex positioning alignment processes are used, then voice control accuracy is improved, but interaction time increases

Engineering Contradiction:
Improvevoice control accuracyVSAvoidinteraction time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent removes the time-consuming wake-up word detection and complex positioning alignment processes from the interaction flow. Image recognition instantly captures user identity and contextual information, enabling the system to process voice commands more quickly while maintaining accuracy through visual verification of user intent.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system performs preliminary image analysis to identify user identity and predict intent before voice processing begins. This pre-preparation of contextual information from visual data allows for faster and more accurate voice command execution, reducing overall interaction time while improving control precision.

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If continuous monitoring and complex positioning are performed, then control precision is improved, but power consumption increases

Engineering Contradiction:
Improvecontrol precisionVSAvoidpower consumption
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The patent extracts and eliminates continuous monitoring and complex positioning operations from the system. Image recognition captures essential user identity and contextual information in a single acquisition, providing sufficient control precision without the need for continuous power-intensive monitoring or complex positioning calculations.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system performs partial action by using image recognition to capture only the essential information needed for control (user identity and basic contextual cues) rather than continuously monitoring all possible parameters. This selective approach achieves adequate control precision while significantly reducing power consumption compared to comprehensive continuous monitoring.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentEP3793202B1Method and device for voice control,terminal, and storage
Publication Date: 2024.09.04 BEIJING XIAOMI MOBILE SOFTWARE CO LTD
  • EP3793202B1 patent drawingFigure 1
  • EP3793202B1 patent drawingFigure 2
  • EP3793202B1 patent drawingFigure 3

AI summary

A method for voice control includes that: a voice is acquired to obtain a voice signal; image information is obtained; whether a pose attribute of a target object that utters the voice satisfies a preset condition is determined based on the image information; and responsive to that the pose attribute of the target object satisfies the preset condition, an operation indicated by the voice signal is performed. A terminal and a non-transitory computer-readable storage medium are also provided.