Home Appliance Speech Control via Vision-Based Intent Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speech recognition systems require a wake-up word to initiate speech recognition, leading to inconvenient user interactions, excessive power and processing resource consumption, and inability to naturally understand user intentions without explicit commands.
Innovation Solution
A method using a vision sensor to determine the presence and intention of a user, activating the speech recognition module only when necessary, allowing for natural speech interaction without a wake-up word and reducing resource usage, while also notifying users of device issues remotely.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If speech recognition function is activated at all times, then user can interact naturally without wake-up word, but power consumption and processing resource consumption increase excessively
Solution Approach 1:
The system performs preliminary detection using a vision sensor to identify user presence and interaction intent before activating the speech recognition module. This preliminary action allows the system to prepare for natural speech interaction only when needed, avoiding continuous activation and excessive power consumption.
2Ease of operation
If speech recognition function is activated at all times, then user can interact naturally without wake-up word, but machine may respond to voice when user does not intend to interact
Solution Approach 1:
The vision sensor acts as an intermediary between the user and the speech recognition module. It detects user presence, gaze direction, and facial expressions to determine interaction intent, serving as a mediator that filters out unintended voice inputs while allowing natural speech interaction when the user genuinely wants to communicate.
3Use of energy by moving object
If wake-up word is required for each command, then power consumption is reduced, but user experience becomes unnatural and inconvenient
Solution Approach 1:
The system dynamically adjusts the activation state of the speech recognition module based on real-time vision sensor input. When the user is detected to be present and intending to interact, the speech module is activated dynamically without requiring a wake-up word. When no interaction is detected, the module remains inactive, reducing power consumption while maintaining natural user experience.
4Productivity
If vision sensor is used to detect user presence and intent, then speech recognition can be activated selectively, but device complexity increases
Solution Approach 1:
The vision sensor serves multiple functions: detecting user presence, determining gaze direction, analyzing facial expressions, and identifying interaction intent. This multi-functionality allows the system to use a single sensor component for various detection tasks, reducing overall system complexity while improving speech interaction efficiency through selective activation of the speech recognition module.
Data Source
AI summary
A method of controlling a home appliance which operates in an Internet of Things environment through a 5G communication network and which is performed using a neural network model generated by machine learning, including determining whether there is a user in the vicinity of the home appliance, capturing a motion of the user using a vision sensor based on a determination that there is a user in the vicinity of the home appliance, identifying an intention of the user based on the captured motion, and activating a speech module of the home appliance based on the intention of the user.


