Speech Recognition Using Gaze and Nonlexical Word Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current speech recognition systems require users to manually operate a start button, which can be inconvenient and may lead to accidents, such as during vehicle operation, as users need to divert attention from the road to initiate speech recognition.

Innovation Solution

A speech recognition apparatus and method that utilizes a camera to track the user's gaze and detects a nonlexical word, such as an interjection, to automatically initiate and operate the speech recognition system, allowing for hands-free and convenient speech recognition without the need to press a start button.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If a manual start button is used to initiate speech recognition, then the system can be controlled reliably, but the ease of operation deteriorates and user safety is compromised

Engineering Contradiction:
Improveease of operationVSAvoiduser safety
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The system performs preliminary detection of user gaze direction and presence before activating speech recognition. The camera detects whether the user is looking at the apparatus and whether the user's face is within a predetermined distance, preparing the system state in advance so that speech recognition can be activated automatically without requiring manual button pressing, thus improving ease of operation while maintaining reliability through pre-validated conditions

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The speech recognition system activates itself automatically based on detected user conditions (gaze direction and presence) without requiring manual intervention. The system serves itself by autonomously determining when speech recognition should be activated based on camera detection results, eliminating the need for users to press start buttons while maintaining controlled activation through self-monitored conditions

Inventive Principle:
Principle #25Self-service

2Ease of operation

If speech recognition is activated automatically without manual input, then ease of operation improves, but the device complexity increases

Engineering Contradiction:
Improveease of operationVSAvoiddevice complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The camera serves multiple functions: it captures images for user presence detection, determines gaze direction, and triggers speech recognition activation. By making the camera a multi-functional component that performs detection and control tasks, the system achieves automatic activation without adding separate dedicated components, thus improving ease of operation while limiting the increase in device complexity

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The camera acts as an intermediary device that mediates between the user and the speech recognition system. Instead of directly connecting user intent to system activation, the camera captures visual information and translates it into control signals that automatically activate speech recognition when conditions are met, simplifying the user interaction while managing system complexity through an intermediate detection layer

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentEP2871640B1Speech recognition apparatus and method
Publication Date: 2021.01.06 LG ELECTRONICS INC
  • EP2871640B1 patent drawingFigure 1
  • EP2871640B1 patent drawingFigure 2
  • EP2871640B1 patent drawingFigure 3

AI summary

The present specification relates to a speech recognition apparatus and method capable of accurately recognizing the speech of a user in an easy and convenient manner without the user having to operate a speech recognition start button or the like. The speech recognition apparatus according to embodiments of the present specification comprises: a camera for capturing a user image; a microphone; a control unit for detecting a preset user gesture from the user image, and, if a nonlexical word is detected from the speech signal which is input through the microphone from the point in time at which the user gesture was detected, determining the speed signal detected after the detected nonlexical word as an effective speech signal; and a speech recognition unit for recognizing the effective speech signal.