Speech Recognition Wake-Up Using Gaze-to-Speech Timing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech recognition devices require a keyword to awaken, increasing use complexity and degrading user experience.

Innovation Solution

A method and apparatus that awaken a speech recognition device based on the interval between the user staring at the device and making a speech, eliminating the need for a wakeup word.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If a keyword awakening mode is used for speech recognition devices, then the device can be activated through voice commands, but the use complexity increases and user experience deteriorates

Engineering Contradiction:
Improveease of awakening speech deviceVSAvoidcomplexity of awakening mode
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The system performs preliminary actions by detecting user gaze direction and anticipating speech intent before the user actually speaks. The electronic device monitors the user's line of sight using camera modules, and when the user stares at the device for a predetermined time, the system pre-activates speech recognition functions, eliminating the need for explicit keyword awakening.

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If the device monitors user gaze and speech interval to determine awakening, then the awakening accuracy improves, but the processing complexity increases

Engineering Contradiction:
Improveaccuracy of awakening detectionVSAvoidcomplexity of monitoring system
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system introduces an intermediary temporal parameter (predetermined time interval) between gaze detection and speech recognition activation. This intermediary mechanism simplifies the complex task of intent recognition by using time-based filtering: the system only activates speech recognition when speech occurs within a specific time window after gaze detection, effectively filtering out false positives while maintaining simple processing logic.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentEP4498365B1Control method and apparatus for speech recognition device, and electronic device and storage medium
Publication Date: 2026.04.08 GREE ELECTRIC APPLIANCE INC OF ZHUHAI
  • EP4498365B1 patent drawingFigure 1
  • EP4498365B1 patent drawingFigure 2
  • EP4498365B1 patent drawingFigure 3

AI summary

Disclosed are a method and apparatus for controlling a speech recognition device, an electronic device, and a storage medium. The method includes: recording, when the speech recognition device is in a dormant state, current time as first time in response to detecting that a target object stares at the speech recognition device (S201); recording the current time as second time in response to detecting that the target object makes a speech (S202); and awakening the speech recognition device in response to determining that an interval between the first time and the second time satisfies a preset condition, such that the speech recognition device enters a speech control mode (S203). According to the method, the speech recognition device is awakened according to an interval between time when the target object stares at the speech recognition device and time when the speech is made, and the target object does not need to firstly say a wakeup word to awaken the speech recognition device. In this way, complexity of awakening a speech device is reduced, and user experience is improved.