Voice Input Apparatus Intent Estimation Without Wake Word

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing voice input systems require a wake word for enabling voice operations, which can be cumbersome and lead to erroneous operations, especially in situations where quick action is needed, such as when a user's hand covers their mouth or when operating an image capture device.

Innovation Solution

A voice input apparatus and method that allows voice operations to be performed without the need for a wake word by estimating user intent based on voice instruction timing and image recognition, enabling quick and accurate processing.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If a wake word is required to enable voice operations, then erroneous operations are reduced, but operation speed and convenience deteriorate

Engineering Contradiction:
Improveerror reductionVSAvoidoperation time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system performs preliminary actions by continuously monitoring voice inputs and pre-identifying potential wake words or operational intents before formal voice commands are issued. This allows the system to be pre-prepared to execute commands quickly without requiring the user to always speak the complete wake word sequence.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The wake word detection system dynamically adjusts its sensitivity and recognition thresholds based on contextual factors such as user proximity, environmental noise levels, and usage patterns. This dynamic adaptation allows the system to reduce false positives while maintaining fast response times for legitimate commands.

Inventive Principle:
Principle #15Dynamics

2Reliability

If a wake word is required to enable voice operations, then erroneous operations are reduced, but ease of operation deteriorates

Engineering Contradiction:
Improveerror reductionVSAvoidoperation convenience
Core Design Contradiction:
ReliabilityVSEase of operation

Solution Approach 1:

The system provides self-service by automatically detecting and confirming user intent through contextual analysis of voice patterns, device state, and environmental factors. This allows the system to execute operations without requiring explicit wake word confirmation in every case, making the interface more natural and convenient while maintaining reliability through intelligent判断.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system applies partial wake word detection by recognizing fragments or variations of wake words in context, rather than requiring complete and exact matches every time. This partial recognition approach, combined with contextual verification, maintains error reduction while significantly improving operational convenience and speed.

Inventive Principle:
Principle #16Partial or excessive action

3Loss of time

If image recognition is used to detect user intent, then quick operations are enabled, but device complexity increases

Engineering Contradiction:
Improveoperation timeVSAvoidsystem complexity
Core Design Contradiction:
Loss of timeVSDevice complexity

Solution Approach 1:

The image recognition system is segmented into modular components that perform specific functions: face detection, expression analysis, gesture recognition, and intent classification. Each module operates independently and can be selectively activated based on the situation, reducing overall system complexity while enabling quick operation detection.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The image recognition system is designed with multi-functionality to detect various types of user intent including facial expressions, hand gestures, and body language. This universal detection capability allows a single system to handle multiple quick operation scenarios without requiring separate dedicated systems for each function, thereby managing complexity efficiently.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS11394862B2Voice input apparatus, control method thereof, and storage medium for executing processing corresponding to voice instruction
Publication Date: 2022.07.19 CANON KK
  • US11394862B2 patent drawing
  • US11394862B2 patent drawing
  • US11394862B2 patent drawing

AI summary

A voice input apparatus inputs voice and performs control to, in a case where a second voice instruction for operating the voice input apparatus is input in a fixed period after a first voice instruction for enabling operations by voice on the voice input apparatus is input, execute processing corresponding to the second voice instruction. The voice input apparatus, in a case where it is estimated that a predetermined user issued the second voice instruction, executes processing corresponding to the second voice instruction when the second voice instruction is input, even in a case where the first voice instruction is not input.