Long-Touch Gesture for Multi-Modal Speech Recognition Initiation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current multi-modal speech recognition systems require users to perform multiple steps and explicit actions to initiate and control speech recognition, leading to complexity, errors, and privacy concerns, especially in mobile devices.

Innovation Solution

A system that uses a long-touch gesture on a touch-sensitive display to simultaneously indicate the target of speech input and initiate speech recognition, eliminating the need for separate actions to activate speech recognition, thereby simplifying user interactions and reducing ambient noise and battery usage.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If a separate button press action is used to initiate speech recognition, then speech recognition can be activated, but the user must perform multiple steps which increases operation complexity and time

Engineering Contradiction:
Improveease of initiating speech recognitionVSAvoidtime to complete multi-modal command
Core Design Contradiction:
Ease of operationVSLoss of time

Solution Approach 1:

The patent combines the speech recognition activation function with the existing long-touch gesture used for object selection. When a user performs a long-touch gesture on an object, the system simultaneously identifies the object and activates speech recognition, merging two separate actions (object selection and speech activation) into one unified gesture. This eliminates the need for a separate button press and reduces the number of steps required to issue a multi-modal command.

Inventive Principle:
Principle #5Merging (Combining)

2Ease of operation

If the microphone is left on continuously to enable open mic functionality, then speech recognition can be initiated instantly, but battery life is reduced and privacy concerns arise

Engineering Contradiction:
Improveease of initiating speech recognitionVSAvoidbattery consumption
Core Design Contradiction:
Ease of operationVSUse of energy by moving object

Solution Approach 1:

The system transitions from continuous microphone monitoring to periodic activation based on gesture detection. The microphone remains off by default and is activated only when a long-touch gesture is detected, creating a periodic on-demand activation pattern rather than continuous operation. This significantly reduces battery consumption while maintaining the ability to instantly respond to user intent when the gesture is performed.

Inventive Principle:
Principle #19Periodic action

3Ease of operation

If the microphone is left on continuously to enable open mic functionality, then speech recognition can be initiated instantly, but ambient noise and privacy concerns increase

Engineering Contradiction:
Improveease of initiating speech recognitionVSAvoidambient noise capture
Core Design Contradiction:
Ease of operationVSObject-affected harmful factors

Solution Approach 1:

The system activates the microphone only during specific periods when a long-touch gesture is detected, rather than continuously monitoring the environment. This periodic activation pattern ensures the microphone captures audio only when the user explicitly intends to speak, preventing the continuous capture of ambient noise and addressing privacy concerns while maintaining instant response capability.

Inventive Principle:
Principle #19Periodic action

4Reliability

If multiple separate actions are required to issue a multi-modal command, then explicit activation is achieved, but user confusion and errors increase

Engineering Contradiction:
Improveaccuracy of command issuanceVSAvoidnumber of interaction steps
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent merges object selection and speech recognition activation into a single long-touch gesture action. Instead of requiring separate actions for object selection and speech activation, the system interprets a long-touch gesture as both selecting the touched object and simultaneously activating speech recognition. This unified approach reduces the number of interaction steps, decreases user confusion, and maintains reliable command issuance by clearly indicating user intent through the gesture duration.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS10276158B2System and method for initiating multi-modal speech recognition using a long-touch gesture
Publication Date: 2019.04.30 AT&T INTELLECTUAL PROPERTY I L P
  • US10276158B2 patent drawing
  • US10276158B2 patent drawing
  • US10276158B2 patent drawing

AI summary

A system, method and computer-readable storage devices are disclosed for multi-modal interactions with a system via a long-touch gesture on a touch-sensitive display. A system operating per this disclosure can receive a multi-modal input comprising speech and a touch on a display, wherein the speech comprises a pronoun. When the touch on the display has a duration longer than a threshold duration, the system can identify an object within a threshold distance of the touch, associate the object with the pronoun in the speech, to yield an association, and perform an action based on the speech and the association.