Long-Touch Gesture for Multi-Modal Speech Recognition Initiation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current multi-modal speech recognition systems require users to perform multiple steps and explicit actions to initiate and control speech recognition, leading to complexity, errors, and privacy concerns, especially in mobile devices.
Innovation Solution
A system that uses a long-touch gesture on a touch-sensitive display to simultaneously indicate the target of speech input and initiate speech recognition, eliminating the need for separate actions to activate speech recognition, thereby simplifying user interactions and reducing ambient noise and battery usage.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If a separate button press action is used to initiate speech recognition, then speech recognition can be activated, but the user must perform multiple steps which increases operation complexity and time
Solution Approach 1:
The patent combines the speech recognition activation function with the existing long-touch gesture used for object selection. When a user performs a long-touch gesture on an object, the system simultaneously identifies the object and activates speech recognition, merging two separate actions (object selection and speech activation) into one unified gesture. This eliminates the need for a separate button press and reduces the number of steps required to issue a multi-modal command.
2Ease of operation
If the microphone is left on continuously to enable open mic functionality, then speech recognition can be initiated instantly, but battery life is reduced and privacy concerns arise
Solution Approach 1:
The system transitions from continuous microphone monitoring to periodic activation based on gesture detection. The microphone remains off by default and is activated only when a long-touch gesture is detected, creating a periodic on-demand activation pattern rather than continuous operation. This significantly reduces battery consumption while maintaining the ability to instantly respond to user intent when the gesture is performed.
3Ease of operation
If the microphone is left on continuously to enable open mic functionality, then speech recognition can be initiated instantly, but ambient noise and privacy concerns increase
Solution Approach 1:
The system activates the microphone only during specific periods when a long-touch gesture is detected, rather than continuously monitoring the environment. This periodic activation pattern ensures the microphone captures audio only when the user explicitly intends to speak, preventing the continuous capture of ambient noise and addressing privacy concerns while maintaining instant response capability.
4Reliability
If multiple separate actions are required to issue a multi-modal command, then explicit activation is achieved, but user confusion and errors increase
Solution Approach 1:
The patent merges object selection and speech recognition activation into a single long-touch gesture action. Instead of requiring separate actions for object selection and speech activation, the system interprets a long-touch gesture as both selecting the touched object and simultaneously activating speech recognition. This unified approach reduces the number of interaction steps, decreases user confusion, and maintains reliable command issuance by clearly indicating user intent through the gesture duration.
Data Source
AI summary
A system, method and computer-readable storage devices are disclosed for multi-modal interactions with a system via a long-touch gesture on a touch-sensitive display. A system operating per this disclosure can receive a multi-modal input comprising speech and a touch on a display, wherein the speech comprises a pronoun. When the touch on the display has a duration longer than a threshold duration, the system can identify an object within a threshold distance of the touch, associate the object with the pronoun in the speech, to yield an association, and perform an action based on the speech and the association.


