Listening Controls for Touchscreen Voice Activation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing voice recognition applications in mobile devices often result in nonsensical messages due to changes in displayed text after the send button is tapped, and there is a need for improved multimodal interaction with graphical user interface controls using spoken commands.
Innovation Solution
The implementation of listening controls that integrate with graphical user interfaces, allowing users to interact with controls using spoken words, transitioning between touch and listening modes through gestures, and displaying visual indicators for speech activation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If speech recognition is used to control mobile device applications, then voice input capability is improved, but text recognition accuracy deteriorates due to nonsensical messages
Solution Approach 1:
The patent implements dynamic mode switching between touch mode and listening mode. The system transitions from static button pressing to dynamic voice-activated control, where the interface adapts its operational state based on detected user intent. This allows the system to optimize for either tactile precision or voice convenience depending on the context.
Solution Approach 2:
The patent introduces an intermediary processing layer that captures speech input, processes it through speech-to-text conversion, and then routes it to the appropriate application function. This intermediary layer acts as a buffer between raw speech and the application logic, enabling accurate text recognition by processing voice commands through a dedicated speech recognition pipeline rather than treating them as regular text input.
2Measurement precision
If traditional touch interface controls are used, then text input accuracy is improved, but user interaction efficiency deteriorates
Solution Approach 1:
The patent replaces the mechanical touch-based interaction system with an acoustic voice-based system for specific control functions. By substituting the mechanical pressing action with voice commands, the system maintains text input accuracy through speech-to-text conversion while significantly improving interaction efficiency for tasks that benefit from hands-free operation.
3Ease of operation
If voice commands are implemented for UI controls, then ease of operation is improved, but device complexity increases
Solution Approach 1:
The patent implements universal listening controls that can be applied across multiple applications and interface elements. Rather than creating separate voice control systems for each application, the patent develops a reusable listening control component that can be instantiated throughout the system, reducing overall complexity through code reuse and standardized interfaces.
Solution Approach 2:
The patent segments the voice control functionality into distinct, modular components: the listening control interface element, the speech recognition processing layer, and the command execution layer. This segmentation allows each component to be developed and tested independently, simplifying the overall system architecture despite the added voice capability.
Data Source
AI summary
Traditional buttons or other controls on a touchscreen are limited to activation by touch, and require the touch to be at a specific location on the touchscreen. Herein, these controls are extended so that they can be additionally activated by voice, without the need to touch them. These extended, listening controls work in either a touch mode or a listening mode to fulfill the same functions irrespectively of how they are activated. The user may control when the device enters and exits the listening mode with a gesture. When in the listening mode, the display of the controls may change to indicate that they have become speech activated, and captured speech may be displayed as text. A visual indicator may show microphone activity. The device may also accept spoken commands in the listening mode.


