AI Speech Recognition Without Wake-Up Word

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current speech recognition technologies require a wake-up word to initiate interaction, limiting user interaction with machines unless the word is explicitly uttered, and fail to facilitate natural conversation without pre-defined triggers.

Innovation Solution

An information providing method and device that gathers situational information from user behavior through home monitoring devices and electronic devices, generates spoken sentences, and converts them into spoken utterance information to initiate and sustain conversations without the need for a wake-up word, using deep neural networks and feedback analysis to improve interaction naturalness and efficiency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If speech recognition is constantly activated to enable natural interaction, then user interaction capability is improved, but power consumption and processing resource usage increase excessively

Engineering Contradiction:
Improveuser interaction capabilityVSAvoidpower consumption
Core Design Contradiction:
Ease of operationVSUse of energy by moving object

Solution Approach 1:

The system performs preliminary analysis by extracting features from user utterances and comparing them against predefined wake-up word features before fully activating speech recognition. This preliminary filtering action enables the system to prepare for potential interaction without committing full computational resources, thus maintaining ease of operation while controlling power consumption.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The speech recognition process is segmented into multiple stages: initial feature extraction, wake-up word detection, and full speech recognition activation. By dividing the continuous speech processing into discrete segments triggered by specific conditions, the system achieves natural interaction capability only when needed, significantly reducing overall power consumption while maintaining user interaction effectiveness.

Inventive Principle:
Principle #1Segmentation

2Use of energy by moving object

If speech recognition is activated only by wake-up word to save power, then power consumption is reduced, but user interaction is limited and less natural

Engineering Contradiction:
Improvepower consumptionVSAvoiduser interaction flexibility
Core Design Contradiction:
Use of energy by moving objectVSAdaptability or versatility

Solution Approach 1:

The system dynamically adjusts its speech recognition activation strategy based on real-time analysis of user utterances. Instead of static wake-up word triggering, the system continuously evaluates speech features and adapts its activation threshold, enabling more flexible and natural user interaction while maintaining power efficiency through condition-based activation rather than constant operation.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system changes the activation parameter from fixed wake-up word matching to dynamic feature-based detection. By monitoring speech features such as phoneme patterns, pitch contours, and temporal characteristics, the system adjusts its activation criteria to recognize diverse user intents beyond predefined wake-up words, thereby improving interaction flexibility while maintaining power consumption benefits.

Inventive Principle:
Principle #35Parameter changes

3Reliability

If wake-up word recognition is used to trigger speech recognition, then system activation is controlled, but continuous conversation without wake-up word is not enabled

Engineering Contradiction:
Improvesystem activation controlVSAvoidconversation continuity
Core Design Contradiction:
ReliabilityVSEase of operation

Solution Approach 1:

The system maintains continuous monitoring of speech features even in idle states, enabling seamless conversation flow without requiring repeated wake-up words. The useful action of speech feature analysis continues at a low computational level, allowing the system to detect and respond to user speech continuously while maintaining reliable activation control through configured thresholds and state management.

Inventive Principle:
Principle #20Continuity of useful action

Solution Approach 2:

The system implements feedback mechanisms where the outcome of each speech recognition cycle influences subsequent activation behavior. When conversation context is detected, the system adjusts its activation sensitivity to maintain continuous interaction. This feedback loop enables the system to remember conversation state and facilitate natural dialogue continuity while preserving reliable activation control through adaptive threshold adjustment based on contextual feedback.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS11322144B2Method and device for providing information
Publication Date: 2022.05.03 LG ELECTRONICS INC
  • US11322144B2 patent drawing
  • US11322144B2 patent drawing
  • US11322144B2 patent drawing

AI summary

Disclosed are an information providing device and an information providing method, which provide information enabling a conversation with a user by executing an artificial intelligence (AI) algorithm and/or a machine learning algorithm in a 5G environment connected for Internet-of-Things. An information providing method according to one embodiment of the present disclosure includes gathering first situational information from a home monitoring device, gathering, from the first electronic device, second situational information corresponding to the first situational information, gathering, from the home monitoring device, third situational information containing a behavioral change of the user after gathering the first situational information, generating a spoken sentence to provide to the user on the basis of the first situational information to the third situational information, and converting the spoken sentence to spoken utterance information to be output to the user.